About this role
The team is seeking a skilled Site Reliability Engineer (SRE) to join their team. The successful candidate will play a crucial role in ensuring the reliability, availability, and performance of their systems and services. You will collaborate with development and operations teams to build and maintain scalable, efficient infrastructure that supports their data and analytics solutions.
Key Responsibilities:
- Design, implement, and manage infrastructure solutions to ensure high availability and performance.
- Monitor system performance and troubleshoot issues to maintain service reliability.
- Automate operational processes to improve efficiency and reduce manual intervention.
- Collaborate with development teams to ensure that applications are designed for reliability and scalability.
- Participate in on-call rotations to provide support for production systems.
- Implement and maintain CI/CD pipelines to streamline deployment processes.
Required Skills & Qualifications:
- Strong experience with cloud platforms such as AWS, Azure, or Google Cloud.
- Proficiency in scripting and automation using languages like Python, Bash, or similar.
- Experience with containerization technologies such as Docker and orchestration tools like Kubernetes.
- Familiarity with monitoring and logging tools (e.g., Prometheus, Grafana, ELK stack).
- Solid understanding of networking, security, and database management.
- Excellent problem-solving skills and a proactive mindset.
Experience:
- Minimum of 5-8 years of experience in Site Reliability Engineering or related fields.
What we offer:
- An opportunity to work in a dynamic and innovative environment.
- A chance to contribute to cutting-edge technology solutions.
- Professional development and growth opportunities.
Applications are read by our talent team, usually within two working days.
If you look like a fit we will call you, and you will hear from us either way.