Summary
What you’ll impact
The Site Reliability Engineer is responsible for keeping business‑critical systems running optimally and ensuring continuity of service while improving operational efficiency through automation. The role involves designing, building, and maintaining infrastructure across production, QA, and development environments, primarily on cloud infrastructure, and writing code to automate operations.
Responsibilities
What you'll do
- Ensure the reliability, performance, and availability of development and production systems in line with business expectations and customer SLAs.
- Automate systems, procedures, and processes wherever practical — and document the rest.
- Build, improve, and maintain CI/CD pipelines and software build, packaging, and deployment processes.
- Implement and improve service- and host-level monitoring, alerting, and observability.
- Respond to incidents, troubleshoot issues, and perform root cause analysis to prevent recurrence.
- Use development and QA environments as a proving ground for new processes and an early warning system for production issues.
- Test system security and implement or advise on improvements, including backups and security tooling.
- Support the migration of our remaining on-premises systems to Oracle Cloud Infrastructure (OCI).
- Stay current with industry trends, recommend improvements worth adopting, and take on other duties as required.
Requirements
What you’ll bring
- 4–5 years of experience in a Site Reliability Engineering, DevOps, systems engineering, or similar role.
- Strong background in operations, software development, or both.
- Substantial experience administering Linux-based infrastructure (Red Hat Enterprise Linux preferred), with working knowledge of Windows.
- Experience with cloud platforms — we run on Oracle Cloud Infrastructure (OCI), and experience with any major cloud provider applies.
- Proficiency with configuration management and automation tools — we use Ansible and Puppet.
- Experience running Docker containers in production.
- Strong scripting and programming skills (e.g., Python, Bash, PowerShell) and working familiarity with Java.
- Experience with Git-based source control (we use Bitbucket) and CI/CD pipelines.
- Experience with databases such as Oracle, Elasticsearch, Redis, or MongoDB.
- Experience with monitoring, alerting, and incident response tools, and performing root cause analysis.
- Familiarity with Agile practices (Scrum/Kanban) and tools such as Jira and Confluence.
- Excellent troubleshooting and analytical skills, with the ability to spot issues before they become problems.
- Clear written and verbal communication with technical and non-technical audiences, strong documentation habits, and the ability to prioritize multiple tasks in a fast-paced team environment.
- Bachelor’s Degree in Computer Science, Information Systems, Software Engineering, or the equivalent combination of education, training, or work experience.