ML Operations Engineer (AI/LLM) - Mercari
Salary not provided
KubernetesDockerPython
English only
English: Fluent
Minimum year of experience: 5
MercariML Operations Engineer (AI/LLM) - Mercari
Work Responsibilities
- Data and Model Orchestration: Own the end-to-end orchestration of model inference, including integrating with DataServ for retrieval (BigQuery, BigTable, Valkey) and managing the Model Inference Gateway and Console.
- Model Serving and Deployment: Own production model serving on the cloud-native NVIDIA and TPU stacks (Triton Inference Server, TensorRT-LLM, JAX/TPU Gateways). Manage model repositories, dynamic batching, and concurrent model execution. Build CI/CD, rollout, and rollback paths for safe, scalable model deployment, and automate provisioning and lifecycle management (Terraform, Kubernetes) so the platform scales seamlessly across teams.
- Inference Performance Optimization: Profile and optimize deployments across LLM and non-LLM workloads, including model compilation, quantization, and batching strategies to hit latency and throughput targets while managing cost. Maintain performance baselines and regression detection.
- Monitoring and Reliability: Build robust monitoring and alerting for model health and latency, including service-level metrics for the data retrieval and inference gateway layers. Define SLOs and own on-call and incident response for the serving layer.
- Model Quality and Evaluation: Build automated evaluation and quality monitoring into the deployment path, regression and drift detection, offline/online evaluation, and LLM output quality checks, so models stay healthy long after launch.
- Research-to-Production Enablement: Partner with ML engineers and researchers to turn experimental models into production-ready services, providing self-service workflows and abstractions that let teams deploy safely and quickly without deep infrastructure expertise.
Unique Challenges
- Build and operate the production ML serving and Data Orchestration platform behind Mercari Group's AI and LLM features, serving tens of millions of users.
- Drive the strategy for model inference at scale, bridging the gap between complex data retrieval and fast-moving ML model inference to ensure high-performance, cost-effective service delivery.
- Shape Mercari's next-generation LLM serving stack, from inference optimization (quantization, dynamic batching, KV caching) to the evaluation and execution infrastructure needed for emerging agentic AI workloads.
Qualifications
Required Experience/Skills
- Shared belief in the mission and values of Mercari Group.
- Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
- 5+ years of software engineering experience, including proven experience in production MLOps: end-to-end model deployment, serving, and CI/CD in cloud environments.
- Experience designing and operating large-scale, high-availability distributed systems, including observability, SLO definition, and incident response.
- Strong experience in cloud-native infrastructure (Kubernetes, Docker).
- Proficiency in Python and infrastructure-as-code (Terraform).
- Excellent written and verbal communication.
Preferred Experience/Skills
- Experience integrating ML serving with large-scale distributed data layers (e.g., data warehouses, wide-column stores, in-memory caches).
- Expertise in model inference optimization (TensorRT-LLM, quantization, JAX).
- Experience operating large-scale model inference gateways and orchestrators.
- 2+ years of hands-on experience operating GenAI/LLM workloads in production (e.g., LLM serving frameworks, token throughput and cost optimization).
- Experience building LLM evaluation, guardrail, or quality-monitoring pipelines (e.g., LLM-as-judge, golden datasets, drift detection).
- Experience with serving infrastructure for RAG or agentic AI workloads (vector search, tool-calling execution environments).
- Experience partnering closely with research or data science teams to bring research innovations into production.
- Master's or Ph.D. in a related technical field.
Language
- English: Proficient (CEFR - B2)
- Japanese: Independent (CEFR - B2) optional