About this role
As an SRE-2 at the team, you'll be a critical member of our Site Reliability Engineering team, responsible for the health and performance of key services. You will contribute directly to the evolution of our infrastructure at a scale that few engineers get to experience. This role offers you the chance to deepen your technical expertise, take on more ownership, and mentor emerging talent while working on a platform that operates at the cutting edge.
Key Responsibilities:
- Take ownership of the reliability and performance of critical services.
- Implement and maintain monitoring, alerting, and incident response processes.
- Collaborate with development teams to improve the reliability and scalability of applications.
- Design and implement infrastructure improvements to enhance service performance.
- Mentor junior engineers and promote best practices in reliability engineering.
- Participate in on-call rotations to ensure service availability.
Required Skills & Qualifications:
- Strong experience in SRE or DevOps roles, with a focus on reliability engineering.
- Proficiency in cloud platforms such as AWS, Azure, or Google Cloud.
- Experience with containerization and orchestration tools like Docker and Kubernetes.
- Familiarity with programming languages such as Python, Go, or Java.
- Knowledge of monitoring and logging tools like Prometheus, Grafana, or ELK stack.
- Excellent problem-solving skills and the ability to work in a fast-paced environment.
Experience: 5-8 years in Site Reliability Engineering or related fields.
What we offer:
- An opportunity to work on cutting-edge technologies in a dynamic environment.
- A culture of learning and professional growth with mentorship opportunities.
- A collaborative team atmosphere where your contributions are valued.
Applications are read by our talent team, usually within two working days.
If you look like a fit we will call you, and you will hear from us either way.