Summary
What you’ll impact
This role is a Staff or Principal-level Platform Engineer responsible for end-to-end ownership of the cloud infrastructure powering the organization's realtime text-to-speech and LLM routing products. The engineer will design, scale, and secure a Kubernetes-based platform across multiple cloud providers to support low-latency, high-availability AI inference at consumer scale. They will also lead SRE practices, build tooling, and drive platform improvements across the organization.
Responsibilities
What you'll do
- You will design, scale, and secure a Kubernetes-based internal platform across multiple cloud providers, supporting low-latency, high-availability AI inference at consumer scale.
- Working closely with engineers across the organization, you will shape how the company builds, deploys, monitors, and scales its services, and champion a “you build it, you run it” culture.
- Design, deploy, and maintain reliable, high-performance, and secure cloud infrastructure for the company’s realtime TTS and LLM routing products.
- Manage and scale production Kubernetes clusters, authoring Helm charts and Kustomize manifests for application deployments.
- Own CI/CD pipelines and infrastructure deployments using Terraform, Terragrunt, ArgoCD, GitHub Actions, and related GitOps tooling.
- Partner with engineering teams to deploy and evolve services across Google Cloud Platform, Microsoft Azure, and Oracle Cloud.
- Build the tooling and processes that let teams monitor the reliability, availability, and performance of their own services.
- Lead root cause analysis for critical incidents and deliver automated solutions that prevent recurrence.
- Identify and build AI-powered developer tooling and workflows that increase engineering velocity across the organization.
- Influence org-wide engineering practices and tooling decisions, optimizing the platform for realtime, low-latency AI inference.
Requirements
What you’ll bring
- Staff or Principal-level Platform Engineer