Summary
What you’ll impact
The role is a senior Site Reliability Engineer focused on building and scaling a centralized observability platform for satellite and aerospace networks. The engineer will design metrics, logging, and tracing systems, define SLO/SLI frameworks, and lead monitoring and incident response strategies while collaborating across infrastructure and software teams.
Responsibilities
What you'll do
- Help design and build Aalyria's centralized observability platform, integrating and scaling tools for metrics (e.g. Prometheus), logging (e.g. Loki), and distributed tracing (e.g. Tempo/OpenTelemetry).
- Define, implement, and manage a robust framework of Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets for our core products, ensuring we are launch-ready.
- Partner with SWEs to implement observability best practices, develop standard templates and documentation, and configure tooling (e.g., OpenTelemetry libraries).
- Automate the deployment, scaling, and management of the entire observability stack using Infrastructure as Code (e.g. Terraform) and GitOps principles (e.g. ArgoCD).
- Partner closely with the core infrastructure team to ensure deep visibility into our Kubernetes clusters and underlying GCP and AWS environments.
- Develop and lead the company's monitoring, alerting, and incident response strategy, driving a culture of proactive reliability and blameless post-mortems.
- On-call responsibilities.
Requirements
What you’ll bring
- 4+ years of experience in an SRE or platform engineering role, with a focus on observability for large-scale, distributed compute or network systems.
- Deep, hands-on expertise building, scaling, and managing observability platforms (e.g., Prometheus, Grafana, Loki/ELK, OpenTelemetry, Tempo/Jaeger, Honeycomb, etc.). You have proven experience using these tools to support performance analysis and debugging of complex distributed systems.
- Strong production-level experience with Google Cloud Platform (GCP) and Kubernetes.
- Experience using Infrastructure as Code (IaC) and GitOps principles (e.g., ArgoCD).
- Proficiency in a systems programming language, with a strong preference for Go and Python for debugging and writing tooling.
- Demonstrable experience defining, implementing, and managing SLOs, SLIs, and error budgets for production services for high availability distributed systems.
- Experience operating a multi-cloud environment, specifically GCP and AWS.
- Hands-on experience with GitLab CI for CI/CD pipelines.
- Working knowledge of service mesh technologies such as Istio or Linkerd.
- Familiarity with instrumenting applications written in Go and C++.
- An active Secret clearance, or higher, is preferred for this position.
- Experience with JVM observability (tuning, monitoring) for Java-based applications.