About this role
The team is seeking a Senior Site Reliability Engineer (SRE) to join their team and lead production reliability, observability, and operational excellence initiatives. In this role, you will be instrumental in ensuring the stability and performance of critical systems while collaborating with various teams to enhance the overall infrastructure.
Key Responsibilities:
- Lead efforts to improve system reliability and performance through proactive monitoring and incident response.
- Design and implement observability solutions to gain insights into system behavior and performance metrics.
- Collaborate with development teams to ensure best practices in deployment and operational processes.
- Automate repetitive tasks and streamline workflows to enhance operational efficiency.
- Conduct post-incident reviews and implement recommendations to prevent future incidents.
- Mentor and guide junior engineers in best practices for reliability and operational excellence.
Required Skills & Qualifications:
- Strong experience in Site Reliability Engineering or DevOps roles, with a focus on production systems.
- Proficiency in cloud platforms such as AWS, Azure, or Google Cloud.
- Experience with container orchestration technologies like Kubernetes and Docker.
- Familiarity with monitoring tools such as Prometheus, Grafana, or similar.
- Strong scripting skills in languages like Python, Bash, or Go.
- Excellent problem-solving skills and a proactive approach to system reliability.
Experience:
- 5-8 years of relevant experience in Site Reliability Engineering or related fields.
What we offer:
The team provides a dynamic work environment that fosters innovation and collaboration. You will have the opportunity to work with cutting-edge technologies and contribute to meaningful projects that impact the healthcare industry.
Applications are read by our talent team, usually within two working days.
If you look like a fit we will call you, and you will hear from us either way.