Large Language Model (LLM) Evaluation Engineer

KUAILU SOFTWARE (SINGAPORE) PTE. LTD.Singaporemycareersfuturepublished 09/24/2026
Must-have:PythonCI/CDAI

Job Responsibilities

  • Build and maintain an automated LLM evaluation pipeline covering multiple dimensions, including general capabilities, Agent capabilities, and persona/role-playing. The pipeline should support one-click evaluation, historical result comparison, and regression testing.
  • Conduct general capability evaluations using benchmarks such as MMLU, C-Eval, HumanEval, GSM8K, MATH, and IFEval, including benchmark deployment, execution, and results analysis.
  • Conduct Agent capability evaluations, including setting up evaluation environments and tracking metrics for benchmarks such as BFCL, τ-bench, and GAIA.
  • Design and execute persona/role-playing evaluation frameworks, covering metrics such as identity recognition, role compatibility, multi-turn stability, and style consistency.
  • Record and analyze evaluation results from training runs, conduct comparative analysis and anomaly detection, and produce checkpoint evaluation reports.
  • Conduct regular intermediate evaluations during the pre-training stage to track the evolution and improvement of model capabilities.

Job Requirements

  • Bachelor's degree or above in Computer Science, Artificial Intelligence, or a related field.
  • Familiarity with mainstream LLM evaluation benchmarks and frameworks, such as lm-eval-harness, OpenCompass, and EvalPlus.
  • Strong proficiency in Python, with the ability to independently build evaluation pipelines covering model inference/deployment, batch evaluation, and results analysis.
  • Familiarity with LLM inference frameworks such as vLLM and SGLang, with the ability to deploy models for batch inference and evaluation.
  • Experience in evaluation data analysis and visualization.
  • Detail-oriented and rigorous, with a strong focus on ensuring the reproducibility and reliability of evaluation results.

Preferred Qualifications

  • Experience with Agent evaluation, particularly BFCL, τ-bench, GAIA, or SWE-bench.
  • Experience with persona or role-playing evaluation, such as CharacterBench or RMTBench.
  • Experience with evaluation automation and CI/CD integration.
  • Understanding of model training workflows, with the ability to understand the relationship between training checkpoints and evaluation results.