Featured Job

Site Reliability Engineer

San Francisco, CA Full-time On-site 08/26/2026 Job ID: 000049
Apply Now
Debugging production systems Logs Traces Incidents Problem-solving

Summary

What you’ll impact

HappyRobot is seeking a Site Reliability Engineer to lead the scaling of operational resilience, owning stability, observability, and debugging workflows. The role focuses on reducing incident load, building internal tooling, and shifting operations from reactive to proactive to improve system uptime and developer focus.

Responsibilities

What you'll do

  • Take the lead on scaling operational resilience as the company grows.
  • Own the stability, observability, and debugging workflows that keep systems running smoothly.
  • Be the go-to person for untangling complex failures in real time.
  • Design tools that turn chaos into clarity.
  • Help shift from reactive to proactive operations.
  • Reduce incident load.
  • Build internal tooling.
  • Directly improve developer focus and system uptime.

Requirements

What you’ll bring

  • 3+ years of hands-on experience debugging production systems (logs, traces, incidents, etc.)
  • Strong problem-solving skills and ability to dive into unfamiliar backend codebases
  • Strong Go and Kubernetes experience.
  • Familiarity with observability and monitoring tools (e.g., Grafana, Prometheus, Sentry)
  • Clear, calm communication under pressure — especially during live incidents

Ready to Move Forward?

Apply now and our recruiting team will reach out with next steps, interview guidance, and client insights tailored to this role.