Summary
What you’ll impact
The role focuses on designing, building, and managing enterprise monitoring and observability solutions across cloud and container environments. The engineer will lead reliability initiatives, develop automated remediation, and mentor platform teams.
Responsibilities
What you'll do
- Design and manage enterprise monitoring and observability solutions.
- Build telemetry pipelines, dashboards, alerts, metrics, logs, and distributed tracing.
- Define and monitor SLIs, SLOs, SLAs, and error budgets.
- Implement observability across AWS/Azure, Kubernetes, Docker, and microservices.
- Develop automated remediation and self-healing/reliability solutions.
- Lead reliability initiatives, create runbooks, and mentor engineers across platform teams.
Requirements
What you’ll bring
- Experience: 10+ years in SRE, Platform, Systems, or Reliability Engineering.
- Strong hands-on experience with OpenTelemetry, Elastic Observability, Grafana, OpsRamp, and BigPanda.
- Strong experience with AWS CloudWatch and/or Azure Monitor.
- Strong scripting skills in Python, Bash, PowerShell, JavaScript, or C-family.
- Experience with Jenkins, GitHub/GitLab CI, Terraform, and Ansible is preferred.
- Strong troubleshooting skills across distributed systems, networking, databases, and performance.
- Experience with REST APIs, JSON, ServiceNow, and DevSecOps is a plus.