About this role
The team is seeking a Senior Staff Site Reliability Engineer to join their Site Reliability & Infrastructure Engineering team. In this role, you will be responsible for ensuring the reliability and health of cloud-based applications. You will design monitoring signals to detect issues and build systems that enhance performance as the company scales. Your work will significantly impact thousands of companies across the U.S. and internationally.
Key Responsibilities:
- Design and implement monitoring and alerting systems to ensure application reliability.
- Collaborate with development teams to identify and resolve reliability issues.
- Build and maintain infrastructure that supports scalable applications.
- Optimize system performance for better efficiency and cost-effectiveness.
- Participate in incident response and post-mortem analysis to improve system resilience.
Required Skills & Qualifications:
- Proven experience in Site Reliability Engineering or a similar role.
- Strong knowledge of cloud platforms (e.g., AWS, Azure, Google Cloud).
- Proficiency in programming languages such as Python, Go, or Java.
- Experience with containerization and orchestration tools (e.g., Docker, Kubernetes).
- Familiarity with monitoring tools (e.g., Prometheus, Grafana) and CI/CD practices.
Experience:
- Minimum of 5 to 8 years in Site Reliability Engineering or related fields.
What we offer:
- Opportunity to work with cutting-edge technology in a collaborative environment.
- A chance to make a significant impact on the reliability of applications used by thousands.
- Professional growth and development opportunities.
Applications are read by our talent team, usually within two working days.
If you look like a fit we will call you, and you will hear from us either way.