Featured Job

Senior Reliability Engineer

Lake Mary, FL Full-time On-site 09/23/2026 Job ID: 000256
Apply Now
Site Reliability Engineering Software Engineering Java Observability Monitoring

Summary

What you’ll impact

The SVP Site Reliability Engineer will lead the design and implementation of observability, automation, and reliability initiatives across distributed systems. The role involves building monitoring solutions, automating operational tasks, supporting production incidents, and improving system performance and scalability.

Responsibilities

What you'll do

  • Design and implement end-to-end observability (logs, metrics, traces) across distributed systemsBuild Observability & Monitoring
  • Integrate and optimize tools such as AppDynamics, Dynatrace, Grafana, and Splunk
  • Develop dashboards, alerts, and telemetry frameworks to provide real-time visibility
  • Identify gaps in monitoring and drive adoption of best practices
  • Identify repetitive operational work and automate it using code and tooling
  • Build self-healing and auto-remediation solutions
  • Enable scalable, reliable processes through automation and engineering rigor
  • Improve operational efficiency across production environments
  • Troubleshoot and resolve complex production issues across distributed systems
  • Participate in incident management, triage, and root cause analysis
  • Improve monitoring and automation based on recurring incident patterns
  • Collaborate with support and engineering teams to improve system stability
  • Define and measure service health using SLIs/SLOs and key performance metrics
  • Identify system bottlenecks and reliability risks
  • Contribute to performance optimization and capacity planning
  • Provide input into system architecture to improve resilience and scalability

Requirements

What you’ll bring

  • 9+ years of experience in Site Reliability Engineering, Software Engineering
  • Strong programming background in Java (preferred) or another modern language
  • Experience with at least one observability platform: AppDynamics, Dynatrace, Grafana, or Splunk
  • Hands-on experience supporting and troubleshooting production systems
  • Strong analytical and problem-solving skills
  • Ability to identify inefficiencies and drive automation

Ready to Move Forward?

Apply now and our recruiting team will reach out with next steps, interview guidance, and client insights tailored to this role.