Jobs in Japan
Explore hand-picked jobs in Japan for English speakers across tech, education, marketing, and more.
Vetted companies only. Apply from overseas.
Explore hand-picked jobs in Japan for English speakers across tech, education, marketing, and more.
Vetted companies only. Apply from overseas.
For more details, see the Overview of Our Positions section on our Careers site. As an MLOps engineer on the AI/LLM team, you will own how our machine learning and LLM models reach production and stay healthy there in our cloud-native environment. Your focus is the production serving, deployment, and operations that turn models into reliable, cost-efficient services, seamlessly integrating with our machine learning operations to serve tens of millions of users. Build and operate the production ML serving and Data Orchestration platform behind Mercari Group's AI and LLM features, serving tens of millions of users. Drive the strategy for model inference at scale, bridging the gap between complex data retrieval and fast-moving ML model inference to ensure high-performance, cost-effective service delivery. Shape Mercari's next-generation LLM serving stack, from inference optimization (quantization, dynamic batching, KV caching) to the evaluation and execution infrastructure needed for emerging agentic AI workloads.
- Shared belief in the mission and values of Mercari Group. - Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience. - 5+ years of software engineering experience, including proven experience in production MLOps: end-to-end model deployment, serving, and CI/CD in cloud environments. - Experience designing and operating large-scale, high-availability distributed systems, including observability, SLO definition, and incident response. - Strong experience in cloud-native infrastructure (Kubernetes, Docker). - Proficiency in Python and infrastructure-as-code (Terraform). - Excellent written and verbal communication.
- Experience integrating ML serving with large-scale distributed data layers (e.g., data warehouses, wide-column stores, in-memory caches). - Expertise in model inference optimization (TensorRT-LLM, quantization, JAX). - Experience operating large-scale model inference gateways and orchestrators. - 2+ years of hands-on experience operating GenAI/LLM workloads in production (e.g., LLM serving frameworks, token throughput and cost optimization). - Experience building LLM evaluation, guardrail, or quality-monitoring pipelines (e.g., LLM-as-judge, golden datasets, drift detection). - Experience with serving infrastructure for RAG or agentic AI workloads (vector search, tool-calling execution environments). - Experience partnering closely with research or data science teams to bring research innovations into production. - Master's or Ph.D. in a related technical field.