About this role
The team is seeking an experienced Engineer III specializing in AI Powered Site Reliability Engineering (SRE) to join their team during UK shift timings. In this role, you will be instrumental in enhancing the reliability and performance of large-scale distributed systems that process nearly 3 trillion events daily.
Key Responsibilities:
- Develop and implement AI-driven solutions to improve system reliability and performance.
- Monitor system performance and proactively address potential issues before they impact customers.
- Collaborate with cross-functional teams to design and deploy scalable infrastructure.
- Automate operational processes to enhance efficiency and reduce manual intervention.
- Participate in on-call rotations to provide support for production systems.
Required Skills & Qualifications:
- Strong experience with AI and machine learning concepts as applied to systems engineering.
- Proficiency in programming languages such as Python, Go, or Java.
- Familiarity with cloud platforms like AWS or Azure.
- Experience with containerization technologies like Docker and orchestration tools like Kubernetes.
- Solid understanding of networking, security, and distributed systems.
Experience:
- A minimum of 5-8 years of experience in Site Reliability Engineering or a related field, with a focus on AI technologies.
What we offer:
- Opportunity to work with cutting-edge technology in a dynamic environment.
- A collaborative team culture that values innovation and continuous improvement.
- Professional development opportunities to expand your skills and knowledge.
Applications are read by our talent team, usually within two working days.
If you look like a fit we will call you, and you will hear from us either way.