Summary
What you’ll impact
Our organization seeks a Senior AWS Site Reliability Engineer to provide cloud platform engineering, Kubernetes operations, and automation support for CRM, Integration, and Platform Services in a remote role serving a Washington, DC government customer. The role focuses on improving reliability, observability, incident response, and operational readiness while collaborating with stakeholders across cybersecurity, architecture, and product teams.
Responsibilities
What you'll do
- Design, configure, develop, integrate, test, document, and sustain AWS capabilities using AWS, EKS, ECS, Terraform, Packer, Vault, Consul, GitHub Actions, CI/CD pipelines, CloudWatch, Open Telemetry, container platforms, and security automation.
- Implement and improve CI/CD pipelines, infrastructure as code, container platform operations, monitoring, alerting, and secure deployment automation.
- Translate business, mission, security, accessibility, and operational requirements into practical technical solutions.
- Support platform architecture, backlog refinement, implementation planning, release readiness, and production transition activities.
- Develop reusable patterns, configuration standards, automation, documentation, and support procedures that reduce delivery and sustainment risk.
- Troubleshoot complex issues across platform configuration, code, data, APIs, identity, security, performance, and user experience.
- Collaborate with cybersecurity, privacy, data, infrastructure, QA, and change management teams to align delivery with federal operating expectations.
- Maintain clear technical documentation, design decisions, implementation notes, test evidence, and operational runbooks.
Requirements
What you’ll bring
- Bachelor's degree in Computer Science, Information Systems, Software Engineering, Data Analytics, Cybersecurity, or a related discipline, or equivalent work experience.
- 7+ years of experience in site reliability, production operations, and cloud platform engineering, preferably in complex enterprise or government environments.
- Hands-on experience with AWS implementation, configuration, development, integration, testing, or operations.
- Experience working with Agile delivery teams and translating stakeholder needs into maintainable technical outcomes.
- Strong hands-on knowledge of AWS capabilities, implementation patterns, administration, development, integration, and lifecycle management.
- Ability to design and implement secure, supportable, upgrade-aware solutions that avoid unnecessary customization and reduce technical debt.
- Experience with APIs, identity and access controls, data management, testing, monitoring, troubleshooting, and release coordination.
- Ability to document technical designs, configuration decisions, operational procedures, test results, and risks in clear business and technical language.
- Strong communication skills with the ability to work across government stakeholders, product owners, engineers, QA, cybersecurity, and operations teams.
- Ability to obtain a Public Trust clearance.