
Staff/Senior Software Engineer, Machine Learning Platform (Ad Cloud) - Tokyo
Most applications go out cold — see where you stand first. No sign-up to start.
Don't just apply. Show up ready.
Olive works from this exact posting — no sign-up to start.
At a glance
Requirements
Credentials this posting asks for.
Job overview
Appier is hiring a Staff/Senior Software Engineer, Machine Learning Platform (Ad Cloud) - Tokyo. Appier seeks a Staff/Senior Software Engineer to lead the Machine Learning Platform team in Tokyo. The role will shape the architecture of large‑scale batch and streaming pipelines, build robust job execution frameworks, and develop internal APIs and developer tools on Kubernetes. The engineer will champion observability, adopt LLM‑based assistants, and mentor junior staff to advance the platform.
Key focus areas include Architect and scale batch (Spark) and streaming (Flink) pipelines processing billions of records daily., Design and operate robust ML job execution frameworks for training, inference, and post‑processing., and Build and maintain API servers and developer tools to orchestrate ML jobs on Kubernetes with Argo, Helm, Terraform..
Successful candidates bring Bachelor's Degree In Computer Science Engineering Or Related Field and 4+ Years Hands-On Experience In Data Systems Machine Learning Infrastructure Or Platform Engineering. Important skills include Apache Spark, Apache Flink, Argo, Kubernetes, Helm, and Terraform. Preferred (not required): Prometheus, Grafana, GitHub Copilot, and ChatGPT.
Skills & qualifications
Skills
Qualifications
Full job description
About Appier
Appier is an AI-native Agentic AI as a Service (AaaS) company that uses artificial intelligence (AI) to power business decision-making. Founded in 2012 with a vision of democratizing AI, Appier’s mission is turning AI into ROI by making software intelligent. Appier now has 17 offices across APAC, Europe and U.S., and is listed on the Tokyo Stock Exchange (Ticker number: 4180). Visit www.appier.com for more information.
The Impact You’ll Make at Appier
We’re looking for a Staff/Senior Machine Learning Platform Engineer to join our Machine Learning Platform Team, which powers end-to-end infrastructure for model training, evaluation, deployment, and monitoring at scale. Our platform supports daily execution of hundreds of ML models and processes billions of data records across batch and streaming pipelines.
In this role, you’ll shape the architecture and core components of our ML platform—covering batch (Spark), streaming (Flink), job orchestration (Argo on Kubernetes), and infrastructure tools—while ensuring the platform remains robust, scalable, and developer-friendly. You’ll also champion best practices and modern development tools including LLM-based programming assistants.
What You’ll Work On
- Architect, implement, and scale batch (Spark) and streaming (Flink) pipelines that process billions of records daily for ML training and evaluation.
- Design and operate robust ML job execution frameworks for training, inference, and post-processing.
- Build and maintain internal API servers and developer tools to orchestrate ML jobs on Kubernetes (via Argo Workflows, Helm, Terraform).
- Design and monitor data infrastructure using ClickHouse and PostgreSQL.
- Ensure high availability and observability through monitoring tools like Prometheus and Grafana.
- Collaborate with data scientists, product managers, and engineers to deliver reliable and efficient ML platform capabilities.
- Actively adopt and promote the use of LLM-based tools (e.g., GitHub Copilot, ChatGPT) to accelerate development, documentation, and debugging.
- Mentor junior engineers and help evolve team engineering culture and standards.
What We’re Looking For
- Bachelor’s degree in Computer Science, Engineering, or a related field; Master’s preferred.
- 4+ years of hands-on experience in data systems, machine learning infrastructure, or platform engineering.
- Strong coding proficiency in Python and/or Java, with experience building large-scale production systems.
- Practical experience with Spark, Flink, Kubernetes (GKE), and infrastructure-as-code tools such as Terraform and Helm.
- Experience managing high-throughput data infrastructure using ClickHouse, PostgreSQL, or similar systems.
- Deep understanding of ML pipelines and distributed job execution in production environments.
- Proven ability to apply LLM-based tools (claude code, codex) to boost engineering productivity.
- Strong ownership, architectural thinking, and ability to lead cross-functional platform projects.
#LI-AK1
You've read the whole posting — now see how you match it.