About this role
The team is seeking a Senior Site Reliability Engineer to join their dynamic team. In this role, you will be responsible for ensuring the reliability and performance of cloud-based applications while optimizing infrastructure and processes.
Key Responsibilities:
- Design, deploy, and manage secure and scalable cloud infrastructure on AWS.
- Ensure high availability, fault tolerance, and cost optimization for business-critical applications.
- Implement and enhance observability and monitoring solutions using tools like DataDog.
- Proactively detect issues and drive continuous performance improvements across AWS environments.
- Lead the development and optimization of CI/CD pipelines to streamline deployment processes.
- Collaborate with cross-functional teams to improve system architecture and operational efficiency.
Required Skills & Qualifications:
- Strong experience with AWS services and cloud infrastructure management.
- Proficiency in monitoring and observability tools, particularly DataDog.
- Solid understanding of CI/CD practices and tools.
- Experience with scripting and automation tools.
- Excellent problem-solving skills and ability to work in a fast-paced environment.
Experience:
- 5-8 years of relevant experience in Site Reliability Engineering or a similar role.
What we offer:
- Opportunity to work in a collaborative and innovative environment.
- Professional development and growth opportunities.
- A chance to make a significant impact on the reliability and efficiency of cloud services.
Applications are read by our talent team, usually within two working days.
If you look like a fit we will call you, and you will hear from us either way.