About this role
The team is seeking a Senior Engineer specializing in Site Reliability Engineering to ensure high availability, performance, scalability, and resilience of cloud and infrastructure platforms. The ideal candidate will apply SRE engineering principles, automation-first practices, and observability to drive continual reliability improvements across services and platforms.
Key Responsibilities:
- Implement SRE frameworks, including SLIs, SLOs, SLAs, and error budgets.
- Conduct performance engineering and establish reliability guardrails across cloud platforms and services.
- Drive automation for provisioning, deployment, and monitoring of infrastructure.
- Collaborate with development teams to enhance system performance and reliability.
- Troubleshoot and resolve incidents, ensuring minimal downtime and impact on services.
- Develop and maintain documentation related to SRE practices and procedures.
Required Skills & Qualifications:
- Strong experience with cloud platforms (AWS, Azure, GCP).
- Proficiency in scripting and automation tools (Python, Bash, Terraform).
- Familiarity with container orchestration (Kubernetes, Docker).
- Knowledge of monitoring and observability tools (Prometheus, Grafana, ELK stack).
- Excellent problem-solving skills and ability to work in a fast-paced environment.
Experience:
- Minimum of 5 to 8 years in Site Reliability Engineering or a related field.
What we offer:
- Opportunity to work with cutting-edge technologies in a dynamic environment.
- A collaborative team culture that values innovation and growth.
- Professional development opportunities to enhance your skills and career.
Applications are read by our talent team, usually within two working days.
If you look like a fit we will call you, and you will hear from us either way.