About this role
The team is seeking a Senior Site Reliability Engineer (SRE) to join their Production Support team. In this role, you will be responsible for ensuring the reliability, availability, and performance of critical production systems. You will work closely with development teams to design and implement scalable solutions that enhance system performance and reliability.
Key Responsibilities:
- Monitor and maintain production systems to ensure high availability and performance.
- Collaborate with development teams to design and implement robust and scalable infrastructure solutions.
- Troubleshoot and resolve production issues in a timely manner.
- Automate repetitive tasks to improve efficiency and reduce downtime.
- Develop and maintain documentation related to system architecture and operational procedures.
- Participate in on-call rotation to provide support for production incidents.
Required Skills & Qualifications:
- Bachelor's degree in Computer Science, Engineering, or a related field.
- 5-8 years of experience in Site Reliability Engineering or a similar role.
- Strong knowledge of cloud environments (AWS, Azure, etc.) and containerization technologies (Docker, Kubernetes).
- Proficiency in scripting languages such as Python, Bash, or similar.
- Experience with monitoring and logging tools (Prometheus, Grafana, ELK stack, etc.).
- Excellent problem-solving skills and ability to work under pressure.
What we offer:
The team provides a dynamic work environment that fosters innovation and collaboration. You will have the opportunity to work with cutting-edge technologies and contribute to meaningful projects that impact communities worldwide.
Applications are read by our talent team, usually within two working days.
If you look like a fit we will call you, and you will hear from us either way.