Summary
What you’ll impact
The Site Reliability Engineer will own the reliability of the platform, working across CI/CD, observability, and incident response. They will collaborate with engineering and security teams to design monitoring, set service level objectives, and ensure robust multi-tenant resource scheduling.
Responsibilities
What you'll do
- Define and own end-to-end reliability, from CI/CD flows to production observability and incident response.
- Develop appropriate Service Level Objectives for distributed training systems, balancing job completion reliability and scheduling latency with development velocity.
- Design and implement monitoring and observability across the full training path.
- Drive incident response for Tinker platform issues, ensuring rapid recovery, thorough incident reviews, and systematic improvements that prevent recurrence.
- Harden multi-tenant isolation and resource scheduling so that LoRA-based workload co-scheduling maximizes utilization without compromising reliability or data separation
- Collaborate with security teams to address production vulnerabilities
Requirements
What you’ll bring
- Bachelor's degree or equivalent experience in computer science, engineering, or similar.
- Experience in distributed systems, cloud infrastructure, or site reliability engineering.
- Proficiency writing software to solve reliability problems, including building tooling and automation.
- Experience with production incident response, postmortems, and systematic reliability improvement.
- Strong communication skills and track record of coordination across engineering and research teams.
- Deep experience operating production cloud services at scale (e.g., public cloud platforms, internal cloud services)
- Background in distributed training frameworks and how infrastructure failures surface in training behavior.
- Track record building checkpoint and recovery systems for long-running distributed jobs.
- Expertise in Kubernetes at scale: deploying, operating, debugging, and tuning clusters handling heterogeneous GPU workloads.