Summary
What you’ll impact
HappyRobot is seeking a Site Reliability Engineer to lead the scaling of operational resilience, owning stability, observability, and debugging workflows. The role focuses on reducing incident load, building internal tooling, and shifting operations from reactive to proactive to improve system uptime and developer focus.
Responsibilities
What you'll do
- Take the lead on scaling operational resilience as the company grows.
- Own the stability, observability, and debugging workflows that keep systems running smoothly.
- Be the go-to person for untangling complex failures in real time.
- Design tools that turn chaos into clarity.
- Help shift from reactive to proactive operations.
- Reduce incident load.
- Build internal tooling.
- Directly improve developer focus and system uptime.
Requirements
What you’ll bring
- 3+ years of hands-on experience debugging production systems (logs, traces, incidents, etc.)
- Strong problem-solving skills and ability to dive into unfamiliar backend codebases
- Strong Go and Kubernetes experience.
- Familiarity with observability and monitoring tools (e.g., Grafana, Prometheus, Sentry)
- Clear, calm communication under pressure — especially during live incidents