About this role
The team is seeking a skilled Platform Engineer III with expertise in EKS and Python/Go to join their Site Reliability Engineering (SRE) team. In this role, you will be instrumental in designing, building, and operating the foundational infrastructure that supports their SaaS platform. Your focus will be on automation, reliability, scalability, and operability, ensuring that the multi-tenant cloud platform consistently meets performance and availability standards.
Key Responsibilities:
- Design and implement robust infrastructure solutions using EKS.
- Develop and maintain automation scripts in Python and Go.
- Monitor system performance and troubleshoot issues to ensure high availability.
- Collaborate with cross-functional teams to enhance system reliability and scalability.
- Implement best practices for system security and compliance.
- Participate in on-call rotations and incident response efforts.
Required Skills & Qualifications:
- Strong experience with Kubernetes, specifically EKS.
- Proficiency in programming languages such as Python and Go.
- Solid understanding of cloud infrastructure and services.
- Experience with monitoring and logging tools.
- Familiarity with CI/CD processes and tools.
- Excellent problem-solving skills and attention to detail.
Experience:
- A minimum of 5-8 years of experience in site reliability engineering or a related field is required.
What we offer:
The team provides a dynamic work environment with opportunities for professional growth, a collaborative team culture, and the chance to work on cutting-edge technologies.
Applications are read by our talent team, usually within two working days.
If you look like a fit we will call you, and you will hear from us either way.