About this role
The team is seeking an experienced Infrastructure Site Reliability Engineer (SRE) specializing in High-Performance Computing (HPC) to join their innovative team. The ideal candidate will play a crucial role in ensuring the reliability, availability, and performance of the infrastructure that supports their cutting-edge AI platform.
Key Responsibilities:
- Design, implement, and maintain scalable HPC infrastructure solutions.
- Monitor system performance and troubleshoot issues to ensure optimal operation.
- Collaborate with development teams to integrate HPC resources into AI applications.
- Automate deployment and management processes to enhance efficiency.
- Develop and maintain documentation for system configurations and procedures.
- Participate in on-call rotations to support critical infrastructure.
Required Skills & Qualifications:
- Strong experience in HPC environments and technologies.
- Proficiency in scripting languages such as Python, Bash, or similar.
- Familiarity with cloud platforms (AWS, Azure, etc.) and container orchestration (Kubernetes, Docker).
- Solid understanding of networking, storage, and server architectures.
- Experience with monitoring tools and performance tuning.
- Excellent problem-solving skills and ability to work in a fast-paced environment.
Experience:
- 5-8 years of relevant experience in infrastructure engineering or site reliability engineering, with a focus on HPC.
What we offer:
- Opportunity to work on groundbreaking AI technology.
- Collaborative and innovative work environment.
- Professional development and growth opportunities.
Applications are read by our talent team, usually within two working days.
If you look like a fit we will call you, and you will hear from us either way.