Home·Find Jobs·Infrastructure SRE - HPC
V

Infrastructure SRE - HPC

VettedBench
Posted 13 days ago
Chennai5-8 yearsOn-site

About this role

The team is seeking an experienced Infrastructure Site Reliability Engineer (SRE) specializing in High-Performance Computing (HPC) to join their innovative team. The ideal candidate will play a crucial role in ensuring the reliability, availability, and performance of the infrastructure that supports their cutting-edge AI platform.

Key Responsibilities:

  • Design, implement, and maintain scalable HPC infrastructure solutions.
  • Monitor system performance and troubleshoot issues to ensure optimal operation.
  • Collaborate with development teams to integrate HPC resources into AI applications.
  • Automate deployment and management processes to enhance efficiency.
  • Develop and maintain documentation for system configurations and procedures.
  • Participate in on-call rotations to support critical infrastructure.

Required Skills & Qualifications:

  • Strong experience in HPC environments and technologies.
  • Proficiency in scripting languages such as Python, Bash, or similar.
  • Familiarity with cloud platforms (AWS, Azure, etc.) and container orchestration (Kubernetes, Docker).
  • Solid understanding of networking, storage, and server architectures.
  • Experience with monitoring tools and performance tuning.
  • Excellent problem-solving skills and ability to work in a fast-paced environment.

Experience:

  • 5-8 years of relevant experience in infrastructure engineering or site reliability engineering, with a focus on HPC.

What we offer:

  • Opportunity to work on groundbreaking AI technology.
  • Collaborative and innovative work environment.
  • Professional development and growth opportunities.
Applications are read by our talent team, usually within two working days. If you look like a fit we will call you, and you will hear from us either way.

More open roles

View all roles →
Apply for this role →