About this role
The team is seeking a skilled Site Reliability Engineer (SRE) to join their Engineering Productivity (EngProd) team. In this role, you will be instrumental in maintaining and supporting a rapidly expanding infrastructure and internal user base. The ideal candidate will be versatile, enthusiastic about learning new technologies, and capable of wearing many hats as part of the software engineering team.
Key Responsibilities:
- Design, build, and administer secure, scalable, and fault-tolerant systems.
- Collaborate with team members to enhance engineering productivity and streamline processes.
- Monitor system performance and troubleshoot issues to ensure high availability.
- Implement automation tools and frameworks to improve operational efficiency.
- Develop and maintain documentation for infrastructure and processes.
- Participate in on-call rotations and incident response activities.
Required Skills & Qualifications:
- Strong experience in site reliability engineering or DevOps practices.
- Proficiency in cloud platforms such as AWS, Azure, or Google Cloud.
- Experience with containerization and orchestration tools like Docker and Kubernetes.
- Familiarity with scripting languages such as Python, Bash, or similar.
- Knowledge of infrastructure as code (IaC) tools like Terraform or Ansible.
- Excellent problem-solving skills and ability to work in a fast-paced environment.
Experience: 5-8 years in a related field.
What we offer:
- A dynamic work environment that encourages innovation and professional growth.
- Opportunities to learn and work with cutting-edge technologies.
- A collaborative team culture that values diverse perspectives.
Applications are read by our talent team, usually within two working days.
If you look like a fit we will call you, and you will hear from us either way.