About this role
The team is seeking a Senior Site Reliability Engineer to join their dynamic team. In this role, you will be responsible for ensuring the reliability, availability, and performance of the team’s platform that supports advanced clinical AI solutions. You will work closely with development and operations teams to enhance system reliability and automate processes.
Key Responsibilities:
- Design and implement scalable and reliable systems to support the Rapid Enterprise™ Platform.
- Monitor system performance and troubleshoot issues to ensure optimal uptime.
- Collaborate with engineering teams to improve deployment processes and system architecture.
- Develop automation tools for system maintenance and monitoring.
- Participate in on-call rotations to respond to incidents and outages.
- Conduct post-mortem analyses to identify root causes and implement preventative measures.
Required Skills & Qualifications:
- Strong experience in Site Reliability Engineering or DevOps roles.
- Proficiency in cloud services such as AWS or Azure.
- Experience with containerization technologies like Docker and orchestration tools like Kubernetes.
- Solid understanding of networking, system architecture, and security practices.
- Familiarity with scripting languages such as Python, Bash, or Go.
- Excellent problem-solving skills and ability to work under pressure.
Experience:
- Minimum of 5-8 years of relevant experience in Site Reliability Engineering or a similar field.
What we offer:
- Opportunity to work in a cutting-edge technology environment.
- Collaborative and innovative work culture.
- Professional development and growth opportunities.
Applications are read by our talent team, usually within two working days.
If you look like a fit we will call you, and you will hear from us either way.