Summary
What you’ll impact
Our company is seeking a Site Reliability Engineer to join the Infrastructure team, focusing on reliability, observability, and security of systems handling PHI. The role involves rolling out on-call tooling, implementing monitoring instrumentation, and managing credentials, certificates, and encryption keys. The position offers growth into deeper SRE ownership within a regulated healthcare environment.
Responsibilities
What you'll do
- Continue the PagerDuty rollout across engineering teams, helping them define escalation policies, schedules, and runbooks so on-call is calm and clear.
- Build in-app Datadog instrumentation — metrics, traces, and alerts — working from the observability standards Infrastructure sets.
- Manage credential storage and rotation across our environments.
- Manage certificate issuance, renewal, and rotation.
- Manage encryption key lifecycle and rotation, keeping PHI protected end to end.
Requirements
What you’ll bring
- 2–4 years in SRE, infrastructure, systems, or a strong adjacent engineering background.
- Familiarity with PagerDuty or a comparable on-call and incident-response tool.
- Hands-on with Datadog or another observability platform, or eager to go deep on it.
- Working understanding of secrets, certificate, and key management (e.g. AWS Secrets Manager / KMS / ACM, HashiCorp Vault).
- Careful and methodical with anything touching security and rotation — you sweat the details.
- Eager to grow into broader SRE ownership, and energized by working in a regulated healthcare (HIPAA / PHI) environment.