About this role
The team is seeking a Staff Site Reliability Engineer to join their Site Reliability & Infrastructure Engineering team. This role is crucial in ensuring the reliability and health of applications running on a fully cloud-based infrastructure. You will be responsible for designing monitoring signals, diagnosing issues, and implementing systems that enhance performance as the company scales.
Key Responsibilities:
- Design and implement monitoring and alerting systems to ensure application reliability.
- Troubleshoot and resolve issues in production environments.
- Collaborate with development teams to improve application performance and reliability.
- Develop automation tools to streamline operations and reduce manual intervention.
- Analyze system performance and capacity, making recommendations for improvements.
Required Skills & Qualifications:
- Strong experience with cloud platforms (AWS, Azure, or Google Cloud).
- Proficiency in scripting and programming languages (Python, Go, or similar).
- Experience with container orchestration tools (Kubernetes, Docker).
- Familiarity with infrastructure as code tools (Terraform, Ansible).
- Solid understanding of networking, security, and system architecture.
Experience:
- Minimum of 5-8 years in Site Reliability Engineering or related fields.
What we offer:
- Opportunity to work in a dynamic and innovative environment.
- The chance to make a significant impact on the reliability of applications used by thousands of companies.
- Professional development and growth opportunities.
Applications are read by our talent team, usually within two working days.
If you look like a fit we will call you, and you will hear from us either way.