Senior Data Scientist – AI/LLM Evaluation
We are looking for a Senior Data Scientist to join an international project focused on Artificial Intelligence, Large Language Models (LLMs), and AI evaluation . The project focuses on developing and improving methodologies for evaluating the quality and performance of AI models and agentic workflows . The role involves designing evaluation metrics, building ground-truth datasets, analyzing model performance, and developing advanced approaches to automated and human-feedback-based evaluation.
Daily tasks
- Design and validate evaluation metrics and ground-truth datasets.
- Analyze evaluation quality, including accuracy, false positives, and false negatives.
- Optimize metric normalization and scoring methodologies.
- Develop and improve LLM-based judges for automated evaluation.
- Design and implement human-feedback-based evaluation approaches.
- Explore and develop new evaluation metrics, including:
- complexity,
- prompt adherence,
- output quality and correctness.
- Analyze experiment results and identify opportunities to improve model performance and evaluation methodologies.
- Collaborate with Data Science, AI/ML, and Engineering teams to develop evaluation solutions for AI models and agentic workflows.
Requirements
Proven experience as a Senior Data Scientist or in a similar role.
Strong data science, statistical analysis, and experimentation skills .
Hands-on experience with AI/LLM evaluation .
Strong understanding of LLMs and agentic workflows .
Experience designing, validating, and optimizing evaluation metrics for AI/ML models .
Strong analytical skills and experience investigating model errors, including false positives and false negatives .
Experience working with ground-truth datasets and defining quality criteria.
Ability to independently analyze experimental results and translate findings into actionable recommendations.
Strong problem-solving and analytical mindset.
Nice to Have Basic knowledge of 3D, geometry, or rendering .
Experience with LLM-as-a-Judge / LLM-based evaluation systems .
Experience with human-in-the-loop evaluation methodologies.
Previous experience evaluating AI agents or agentic systems .
Familiarity with advanced approaches to evaluating LLM-generated content and outputs .
Must have: Data science, Statistical analysis, Experimantation skills, AI, LLM, Agentic workflow
Nice to have: 3D, Geometry, Rendering, LLM-as-a-Judge, LLM-based evaluation