Summary
What you’ll impact
Our organization is seeking a Senior Platform Engineer to own engineering enablement, improving the reliability, deployment, and observability of production services on AWS. The role focuses on building shared infrastructure, automating workflows, and supporting coding agents while collaborating with product, AI, data, and security teams.
Responsibilities
What you'll do
- Improve the reliability and operation of production services on AWS ECS.
- Build shared infrastructure, service templates, and delivery workflows that reduce the time from code to production.
- Improve local development, CI, test, and deployment workflows so engineers can work without platform support for routine tasks.
- Build agent-ready repositories, documentation, and interfaces so coding agents can complete useful tasks with less supervision.
- Automate recurring work such as service setup, dependency changes, release checks, and incident triage.
- Create guardrails and feedback loops that make agent-produced changes fast to validate and safe to ship.
- Give engineers useful logs, metrics, alerts, and service health information.
- Define practical standards for incident response, recovery, capacity, and service reliability.
- Build security, privacy, and compliance controls into infrastructure and delivery workflows.
- Track infrastructure cost and improve efficiency as traffic, data, and AI workloads grow.
- Work with product, AI, data, and security engineers to solve production problems.
Requirements
What you’ll bring
- Experience building and operating production systems on AWS.
- Strong knowledge of containers, infrastructure as code, CI/CD, networking, and observability.
- Experience with container orchestration and automated delivery.
- Experience improving developer environments, CI performance, service templates, or internal engineering tools.
- The ability to write software and automation instead of relying on manual procedures.
- Daily experience with coding agents and evidence that you have built tools, context, or workflows that make them more reliable.
- Experience measuring or improving engineering throughput, CI time, test quality, or self-service adoption.
- Experience diagnosing incidents and making follow-up improvements.
- Good judgment about reliability, security, cost, and delivery speed.
- Clear written and verbal communication across engineering teams.
- Many candidates will have 5+ years of relevant experience.