Featured Job

Senior Reliability Engineer

Lake Mary, FL Full-time On-site 09/23/2026 Job ID: 000266
Apply Now
Site Reliability Engineering Software Engineering Java AppDynamics Dynatrace

Summary

What you’ll impact

The Sr. Site Reliability Automation Engineer will design and implement observability solutions, drive automation to reduce operational toil, support production incidents, and improve system reliability and performance. The role involves working with distributed systems, monitoring tools, and collaborating with engineering and support teams to enhance stability and scalability.

Responsibilities

What you'll do

  • Design and implement end-to-end observability (logs, metrics, traces) across distributed systems
  • Build Observability & Monitoring
  • Integrate and optimize tools such as AppDynamics, Dynatrace, Grafana, and Splunk
  • Develop dashboards, alerts, and telemetry frameworks to provide real-time visibility
  • Identify gaps in monitoring and drive adoption of best practices
  • Identify repetitive operational work and automate it using code and tooling
  • Build self-healing and auto-remediation solutions
  • Enable scalable, reliable processes through automation and engineering rigor
  • Improve operational efficiency across production environments
  • Troubleshoot and resolve complex production issues across distributed systems
  • Participate in incident management, triage, and root cause analysis
  • Improve monitoring and automation based on recurring incident patterns
  • Collaborate with support and engineering teams to improve system stability
  • Define and measure service health using SLIs/SLOs and key performance metrics
  • Identify system bottlenecks and reliability risks
  • Contribute to performance optimization and capacity planning
  • Provide input into system architecture to improve resilience and scalability

Requirements

What you’ll bring

  • 3–6 years of experience in Site Reliability Engineering, Software Engineering
  • Strong programming background in Java (preferred) or another modern language
  • Experience with at least one observability platform: AppDynamics, Dynatrace, Grafana, or Splunk
  • Hands-on experience supporting and troubleshooting production systems
  • Strong analytical and problem-solving skills
  • Ability to identify inefficiencies and drive automation

Ready to Move Forward?

Apply now and our recruiting team will reach out with next steps, interview guidance, and client insights tailored to this role.