Summary
What you’ll impact
The Staff DevOps/SRE will own reliability for a commercial multi-tenant cloud product built on Kubernetes, Terraform, and AWS, handling availability, incident response, and the overall reliability story. They will collaborate with engineering leadership to set SLOs, reduce toil, and ensure observability and deployments meet staff-level expectations while working US Pacific business hours.
Responsibilities
What you'll do
- Own reliability for the commercial cloud offering.
- Own availability, incident response, and the reliability story for a commercial multi-tenant cloud product built on Kubernetes, Terraform, and AWS.
- Work with engineering leadership to set SLOs.
- Reduce toil.
- Make observability and deploys reliable for staff-level expectations.
Requirements
What you’ll bring
- 8+ years in SRE, platform, or systems engineering with a strong public cloud and Kubernetes background.
- Proficiency with Terraform, Prometheus-style observability, and on-call for production systems in a B2B context.
- US Pacific overlap for leadership meetings and on-call handoffs; must communicate clearly in English across distributed teams.