Summary
What you’ll impact
The Senior Staff DevOps Engineer (Site Reliability) will lead reliability engineering for the organization, designing and operating scalable, secure cloud infrastructure on AWS. The role involves hands‑on work with CI/CD, Kubernetes, infrastructure‑as‑code, AI‑assisted tools, and incident response while mentoring teams and shaping engineering standards.
Responsibilities
What you'll do
- Lead reliability engineering efforts for the Vantor Hub platform and related infrastructure services.
- Design, implement, and maintain scalable CI/CD pipelines for software and infrastructure delivery.
- Build, operate, and improve cloud infrastructure automation using tools such as Terraform, CloudFormation, Kubernetes, Docker, and AWS-native services.
- Troubleshoot complex infrastructure, deployment, networking, performance, reliability, and production issues.
- Improve service availability, scalability, maintainability, reliability, security, and overall operational readiness.
- Design and operate highly available, resilient, observable, and secure infrastructure across commercial and government AWS environments.
- Partner with engineering teams to improve deployment safety, rollback capabilities, observability, production readiness, and operational support.
- Participate in a team on-call rotation, currently approximately one week every twelve weeks, supporting production reliability and incident response.
- Leverage AI development tools as force multipliers for software design, implementation, testing, and documentation while maintaining reliability, security, and engineering quality.
- Contribute to shared engineering standards for responsible AI-assisted development, including validation practices, documentation expectations, and review patterns.
- Use AI-assisted engineering tools to accelerate infrastructure automation, documentation, runbook development, test scaffolding, incident analysis, and repeatable operational workflows; validate outputs through peer review, testing, and security practices.
Requirements
What you’ll bring
- Bachelor's Degree in Software Engineering, Computer Science, a related engineering field, or equivalent experience.
- 8+ years of experience in DevOps, Platform Engineering, Site Reliability Engineering, or cloud operations, with demonstrated ownership of production systems.
- Demonstrated experience owning, operating, and improving production systems in cloud-based environments.
- Strong experience with Site Reliability Engineering practices, including service-level indicators, service-level objectives, error budgets, incident response, post-incident reviews, reliability metrics, and reliability-focused automation.
- Strong proficiency with Amazon Web Services, including experience operating production workloads in multi-account AWS environments.
- Strong proficiency with containerization technologies such as Docker and Kubernetes.
- Strong proficiency with infrastructure-as-code (IaC) tools like Terraform or CloudFormation.
- Proficiency in scripting languages such as Python, Bash, or PowerShell.
- Solid understanding of networking concepts and protocols.
- The ability to communicate and collaborate with team members and other colleagues.
- Must be a U.S. citizen and be willing and able to obtain a U.S. Government security clearance.