About this role
The team is seeking an AI Evaluation Engineer with a strong background in Python, QA, or Security to join their project-based team. This role focuses on testing, evaluating, and improving AI systems, specifically in the context of developing datasets for AI coding agents.
Key Responsibilities:
- Design and create challenging tasks for AI coding agents to evaluate their performance in real-world developer scenarios.
- Develop evaluation criteria to assess the effectiveness of AI models.
- Collaborate with cross-functional teams to ensure the quality and reliability of AI systems.
- Conduct thorough testing and analysis of AI outputs to identify areas for improvement.
- Document findings and provide actionable insights to enhance AI capabilities.
Required Skills & Qualifications:
- Proficiency in Python and experience with AI evaluation methodologies.
- Strong background in quality assurance or security testing.
- Familiarity with AI and machine learning concepts.
- Excellent problem-solving skills and attention to detail.
- Ability to work independently and manage multiple projects simultaneously.
Experience:
- Minimum of 3-5 years in a relevant field, with a focus on AI evaluation or software testing.
What we offer:
- Opportunity to work on innovative AI projects with leading tech companies.
- Flexible, project-based work environment.
- Collaboration with a team of experts in the AI field.
Applications are read by our talent team, usually within two working days.
If you look like a fit we will call you, and you will hear from us either way.