About this role
The team is seeking a Staff Site Reliability Engineer (SRE) to join their Core SRE team. The primary mission of the SRE organization is to uphold uptime promises to customers by ensuring that Service Level Objectives (SLOs) and Service Level Agreements (SLAs) are consistently met. In this role, you will have the opportunity to drive initiatives that enhance the reliability, stability, and cost efficiency of the team Singularity Platform.
Key Responsibilities:
- Collaborate with engineering teams to facilitate the rapid and quality delivery of software to customers.
- Monitor system performance and implement improvements to maintain high availability.
- Develop and maintain tools and processes for incident response and resolution.
- Analyze and optimize system architecture for scalability and performance.
- Drive reliability improvements through automation and proactive monitoring.
- Participate in on-call rotations and respond to incidents as they arise.
Required Skills & Qualifications:
- Strong experience in site reliability engineering or a related field.
- Proficiency in cloud services such as AWS or Azure.
- Knowledge of containerization technologies like Docker and orchestration tools like Kubernetes.
- Experience with monitoring tools and practices (e.g., Prometheus, Grafana).
- Strong scripting skills in languages such as Python, Go, or Bash.
- Excellent problem-solving skills and a proactive mindset.
Experience:
- Minimum of 5-8 years in a site reliability engineering role or similar.
What we offer:
The team provides a dynamic work environment with opportunities for professional growth, a collaborative team culture, and the chance to make a significant impact on the reliability and efficiency of their services.
Applications are read by our talent team, usually within two working days.
If you look like a fit we will call you, and you will hear from us either way.