Senior Data Scientist – AI/LLM Evaluation

Square One Resourcesnofluffjobspublished 08/28/2026
Must-have:DataAISenior

We are looking for a Senior Data Scientist to join an international project focused on Artificial Intelligence, Large Language Models (LLMs), and AI evaluation . The project focuses on developing and improving methodologies for evaluating the quality and performance of AI models and agentic workflows . The role involves designing evaluation metrics, building ground-truth datasets, analyzing model performance, and developing advanced approaches to automated and human-feedback-based evaluation.

Daily tasks

  • Design and validate evaluation metrics and ground-truth datasets.
  • Analyze evaluation quality, including accuracy, false positives, and false negatives.
  • Optimize metric normalization and scoring methodologies.
  • Develop and improve LLM-based judges for automated evaluation.
  • Design and implement human-feedback-based evaluation approaches.
  • Explore and develop new evaluation metrics, including:
  • complexity,
  • prompt adherence,
  • output quality and correctness.
  • Analyze experiment results and identify opportunities to improve model performance and evaluation methodologies.
  • Collaborate with Data Science, AI/ML, and Engineering teams to develop evaluation solutions for AI models and agentic workflows.

Requirements

Proven experience as a Senior Data Scientist or in a similar role.

Strong data science, statistical analysis, and experimentation skills .

Hands-on experience with AI/LLM evaluation .

Strong understanding of LLMs and agentic workflows .

Experience designing, validating, and optimizing evaluation metrics for AI/ML models .

Strong analytical skills and experience investigating model errors, including false positives and false negatives .

Experience working with ground-truth datasets and defining quality criteria.

Ability to independently analyze experimental results and translate findings into actionable recommendations.

Strong problem-solving and analytical mindset.

Nice to Have Basic knowledge of 3D, geometry, or rendering .

Experience with LLM-as-a-Judge / LLM-based evaluation systems .

Experience with human-in-the-loop evaluation methodologies.

Previous experience evaluating AI agents or agentic systems .

Familiarity with advanced approaches to evaluating LLM-generated content and outputs .

Must have: Data science, Statistical analysis, Experimantation skills, AI, LLM, Agentic workflow

Nice to have: 3D, Geometry, Rendering, LLM-as-a-Judge, LLM-based evaluation