Featured Job

Lead Site Reliability Engineer

Washington, DC Full-time Remote $145k — $185k per year 09/17/2026 Job ID: 000245
Apply Now
AWS EKS ECS Terraform Packer

Summary

What you’ll impact

Our organization seeks a Senior AWS Site Reliability Engineer to provide cloud platform engineering, Kubernetes operations, and automation support for CRM, Integration, and Platform Services in a remote role serving a Washington, DC government customer. The role focuses on improving reliability, observability, incident response, and operational readiness while collaborating with stakeholders across cybersecurity, architecture, and product teams.

Responsibilities

What you'll do

  • Design, configure, develop, integrate, test, document, and sustain AWS capabilities using AWS, EKS, ECS, Terraform, Packer, Vault, Consul, GitHub Actions, CI/CD pipelines, CloudWatch, Open Telemetry, container platforms, and security automation.
  • Implement and improve CI/CD pipelines, infrastructure as code, container platform operations, monitoring, alerting, and secure deployment automation.
  • Translate business, mission, security, accessibility, and operational requirements into practical technical solutions.
  • Support platform architecture, backlog refinement, implementation planning, release readiness, and production transition activities.
  • Develop reusable patterns, configuration standards, automation, documentation, and support procedures that reduce delivery and sustainment risk.
  • Troubleshoot complex issues across platform configuration, code, data, APIs, identity, security, performance, and user experience.
  • Collaborate with cybersecurity, privacy, data, infrastructure, QA, and change management teams to align delivery with federal operating expectations.
  • Maintain clear technical documentation, design decisions, implementation notes, test evidence, and operational runbooks.

Requirements

What you’ll bring

  • Bachelor's degree in Computer Science, Information Systems, Software Engineering, Data Analytics, Cybersecurity, or a related discipline, or equivalent work experience.
  • 7+ years of experience in site reliability, production operations, and cloud platform engineering, preferably in complex enterprise or government environments.
  • Hands-on experience with AWS implementation, configuration, development, integration, testing, or operations.
  • Experience working with Agile delivery teams and translating stakeholder needs into maintainable technical outcomes.
  • Strong hands-on knowledge of AWS capabilities, implementation patterns, administration, development, integration, and lifecycle management.
  • Ability to design and implement secure, supportable, upgrade-aware solutions that avoid unnecessary customization and reduce technical debt.
  • Experience with APIs, identity and access controls, data management, testing, monitoring, troubleshooting, and release coordination.
  • Ability to document technical designs, configuration decisions, operational procedures, test results, and risks in clear business and technical language.
  • Strong communication skills with the ability to work across government stakeholders, product owners, engineers, QA, cybersecurity, and operations teams.
  • Ability to obtain a Public Trust clearance.

Ready to Move Forward?

Apply now and our recruiting team will reach out with next steps, interview guidance, and client insights tailored to this role.