About this role
The team is seeking an experienced Engineering Manager to lead their Site Reliability Engineering (SRE) team. The ideal candidate will possess a strong background in SRE principles and practices, along with experience in managing and mentoring engineering teams. The SRE Manager will play a crucial role in ensuring the overall success of the SRE team, focusing on the reliability, scalability, and security of systems.
Key Responsibilities:
- Lead and mentor a team of SRE engineers, fostering a culture of continuous improvement.
- Oversee the monitoring of system stability and availability, particularly for mission-critical applications.
- Develop and implement SRE best practices and processes to enhance system performance.
- Collaborate with cross-functional teams to ensure alignment on reliability goals and initiatives.
- Manage incident response and post-mortem processes to drive learning and improvement.
- Evaluate and implement tools and technologies that enhance operational efficiency.
Required Skills & Qualifications:
- Strong understanding of SRE principles and practices.
- Proven experience in managing engineering teams, with a focus on mentoring and development.
- Proficiency in cloud platforms (e.g., AWS, GCP, Azure) and container orchestration (e.g., Kubernetes).
- Experience with monitoring and observability tools (e.g., Prometheus, Grafana).
- Strong problem-solving skills and the ability to work under pressure.
- Excellent communication and collaboration skills.
Experience:
- 5-8 years of experience in Site Reliability Engineering or related fields, with at least 2 years in a managerial role.
What we offer:
- An opportunity to lead a talented team in a dynamic environment.
- A culture that promotes innovation and professional growth.
- The chance to work on cutting-edge technologies and impactful projects.
Applications are read by our talent team, usually within two working days.
If you look like a fit we will call you, and you will hear from us either way.