About this role
The team is seeking a skilled Machine Learning Ops Engineer to join their dynamic team. The ideal candidate will be responsible for automating the deployment of machine learning models and managing the infrastructure required for training and inference.
Key Responsibilities:
- Develop, deploy, and manage ML models on Databricks using MLflow for tracking experiments and managing models.
- Set up scalable and fault-tolerant infrastructure to support model training and inference in cloud environments such as AWS, GCP, or Azure.
- Implement monitoring systems to track model performance, accuracy, and drift.
- Collaborate with data scientists and software engineers to integrate ML models into production systems.
- Optimize model performance and ensure reliability in production environments.
Required Skills & Qualifications:
- Proficiency in machine learning frameworks and libraries (e.g., TensorFlow, PyTorch).
- Strong experience with cloud platforms (AWS, GCP, Azure) for deploying ML solutions.
- Familiarity with CI/CD practices for machine learning workflows.
- Knowledge of containerization technologies (e.g., Docker, Kubernetes).
- Excellent problem-solving skills and ability to work in a fast-paced environment.
Experience:
- 3-5 years of relevant experience in machine learning operations or a related field.
What we offer:
- Opportunity to work with cutting-edge technologies in a collaborative environment.
- Professional development and growth opportunities.
- A supportive team that values innovation and creativity.
Applications are read by our talent team, usually within two working days.
If you look like a fit we will call you, and you will hear from us either way.