Senior Platform Infrastructure Engineer - Mercari
Salary not provided
KubernetesGoElasticSearch
English only
English: Fluent
MercariSenior Platform Infrastructure Engineer - Mercari
Work Responsibilities
- Co-own the GKE cluster lifecycle: upgrades, node-pool strategy, add-on management (Cert Manager, autoscaling, workload identity), capacity planning, and incident response.
- Own and evolve the Terraform IaC: custom modules, large-blast-radius migrations, and the Terraform CI/CD that keeps changes safe at scale - drift detection, staged rollout, revert planning, and pipeline reliability.
- Build and maintain "kits" - reusable Terraform modules and self-service tooling that let product teams provision repos (with CODEOWNERS), PagerDuty, GitHub Actions workflows, and deployment pipelines without infra hand-holding.
- Own autoscaling and capacity at scale: Cluster Autoscaler / HPA / VPA, cost-aware right-sizing, and multi-region and node-pool strategies for large workloads.
- Ship internal developer-platform components in Go: k8s controllers, CLIs and internal APIs, and AI agent skills that cut friction on the most-used developer paths.
- Join the Platform on-call rotation after ramp-up; lead root-cause analysis and follow-through improvements.
- Drive reliability, security, and developer-experience initiatives in partnership with FinOps, Network, Observability, DBRE, and Enablement teams.
Unique Challenges
- Kubernetes at multi-tenant scale - We run GKE clusters that host hundreds of microservices deployed by teams across the company. Node-pool strategy (autoscaling, spot, ARM diversification), control-plane and add-on upgrade cadence, and workload isolation are everyday work. Every change to a shared Terraform module, a GKE add-on, or a kit touches all of those services at once, so getting the design and the rollout right up-front is a core part of the job.
- Custom controllers, admission policy, and add-on ownership - A meaningful part of the platform is our own Kubernetes controllers, admission webhooks, and policy layer (RBAC, network policies, OPA / Gatekeeper) that keep a multi-tenant cluster safe and predictable. You design and run these, not just consume upstream.
- Multi-region and multi-architecture roadmap - We are actively expanding to a multi-region topology and adopting different k8s node-types for cost and capacity headroom. You will help shape this rollout.
- Developer-experience as an SLI - The self-service kits are used every day by product engineers. A regression in a kit shows up as friction across the whole engineering org, so we take DX seriously as a reliability signal, not a nice-to-have.
- Cross-team ownership - Platform Infra sits between application teams, cloud providers, and peer infra teams (Search, Network, Observability, FinOps, DBRE, Enablement). Success in this role requires investing substantial time in cross-team architecture, shared platform capabilities, and organizational alignment not just work within your own team.
Qualifications
Required Experience/Skills
- Shared belief in Mercari's mission and values.
- Senior-level engineering experience operating distributed systems in production at scale.
- Deep understanding of Kubernetes internals - scheduler, controllers, kubelet, admission webhooks, autoscaler, network policy, workload identity - and production ownership of a managed Kubernetes environment (GKE, EKS, or AKS): cluster upgrades, incident debugging, add-on management.
- Deep Terraform experience at scale - authoring custom modules, running blast-radius-aware migrations, refactoring shared IaC without breaking downstream consumers.
- Strong system design and distributed systems fundamentals: consistency, replication, capacity planning, failure modes.
- Track record of owning an infrastructure domain end-to-end: design, build, on-call, incident follow-through.
- Professional working proficiency in English.
Preferred Experience/Skills
- Go for building k8s controllers, CLIs, and internal APIs.
- Networking fundamentals - TCP/IP, Kubernetes networking (CNI, Services, ingress, network policy), DNS, HTTP/2, mTLS - enough to reason about incidents that cross the network boundary.
- CI/CD platform experience (GitHub Actions, Jenkins, or similar) - authoring workflows, custom actions, self-hosted runners.
- Autoscaling / capacity-planning depth: HPA / VPA / Cluster Autoscaler internals, cost-aware autoscaling, right-sizing.
- Experience building internal developer platforms - self-service tooling, golden paths, CLI / API ergonomics.
- Multi-region or multi-cluster operational experience.
- Exposure to Elasticsearch or other search infrastructure (helpful for cross-team collaboration, not required).
- FinOps sensibility - understanding cloud spend drivers and how infra choices show up on the bill.
Language
- Japanese: Not required
- English: Business level (CEFR B2)