ML Operations Engineer (AI/LLM) - Mercari

Salary not provided

KubernetesDockerPython
English only
English: Fluent

Minimum year of experience: 5

Mercari

ML Operations Engineer (AI/LLM) - Mercari

Work Responsibilities

  • Data and Model Orchestration: Own the end-to-end orchestration of model inference, including integrating with DataServ for retrieval (BigQuery, BigTable, Valkey) and managing the Model Inference Gateway and Console.
  • Model Serving and Deployment: Own production model serving on the cloud-native NVIDIA and TPU stacks (Triton Inference Server, TensorRT-LLM, JAX/TPU Gateways). Manage model repositories, dynamic batching, and concurrent model execution. Build CI/CD, rollout, and rollback paths for safe, scalable model deployment, and automate provisioning and lifecycle management (Terraform, Kubernetes) so the platform scales seamlessly across teams.
  • Inference Performance Optimization: Profile and optimize deployments across LLM and non-LLM workloads, including model compilation, quantization, and batching strategies to hit latency and throughput targets while managing cost. Maintain performance baselines and regression detection.
  • Monitoring and Reliability: Build robust monitoring and alerting for model health and latency, including service-level metrics for the data retrieval and inference gateway layers. Define SLOs and own on-call and incident response for the serving layer.
  • Model Quality and Evaluation: Build automated evaluation and quality monitoring into the deployment path, regression and drift detection, offline/online evaluation, and LLM output quality checks, so models stay healthy long after launch.
  • Research-to-Production Enablement: Partner with ML engineers and researchers to turn experimental models into production-ready services, providing self-service workflows and abstractions that let teams deploy safely and quickly without deep infrastructure expertise.

Unique Challenges

  • Build and operate the production ML serving and Data Orchestration platform behind Mercari Group's AI and LLM features, serving tens of millions of users.
  • Drive the strategy for model inference at scale, bridging the gap between complex data retrieval and fast-moving ML model inference to ensure high-performance, cost-effective service delivery.
  • Shape Mercari's next-generation LLM serving stack, from inference optimization (quantization, dynamic batching, KV caching) to the evaluation and execution infrastructure needed for emerging agentic AI workloads.

Qualifications

Required Experience/Skills

  • Shared belief in the mission and values of Mercari Group.
  • Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
  • 5+ years of software engineering experience, including proven experience in production MLOps: end-to-end model deployment, serving, and CI/CD in cloud environments.
  • Experience designing and operating large-scale, high-availability distributed systems, including observability, SLO definition, and incident response.
  • Strong experience in cloud-native infrastructure (Kubernetes, Docker).
  • Proficiency in Python and infrastructure-as-code (Terraform).
  • Excellent written and verbal communication.

Preferred Experience/Skills

  • Experience integrating ML serving with large-scale distributed data layers (e.g., data warehouses, wide-column stores, in-memory caches).
  • Expertise in model inference optimization (TensorRT-LLM, quantization, JAX).
  • Experience operating large-scale model inference gateways and orchestrators.
  • 2+ years of hands-on experience operating GenAI/LLM workloads in production (e.g., LLM serving frameworks, token throughput and cost optimization).
  • Experience building LLM evaluation, guardrail, or quality-monitoring pipelines (e.g., LLM-as-judge, golden datasets, drift detection).
  • Experience with serving infrastructure for RAG or agentic AI workloads (vector search, tool-calling execution environments).
  • Experience partnering closely with research or data science teams to bring research innovations into production.
  • Master's or Ph.D. in a related technical field.

Language

  • English: Proficient (CEFR - B2)
  • Japanese: Independent (CEFR - B2) optional