Research Scientist

Jobgether· Brazil· lever· publicada el 07/07/2026
Imprescindible:AILeadRemote

Accountabilities Design and develop original evaluation methodologies that accurately measure the capabilities of frontier AI models across reasoning, coding, agentic workflows, tool use, and multimodal tasks.

Create comprehensive benchmark packages supported by expert-verified ground truth, rigorous quality control processes, deterministic verification methods, and well-calibrated scoring rubrics.

Evaluate benchmark quality by ensuring strong construct validity, appropriate task discrimination, sufficient performance headroom, and minimal benchmark contamination.

Recruit, assess, and coordinate subject matter experts to develop, validate, and review high-quality evaluation datasets and benchmark content.

Serve as the final technical reviewer for evaluation quality, benchmark correctness, and task complexity before external delivery.

Collaborate with AI research organizations to understand evaluation objectives and translate research needs into effective benchmarking strategies.

Lead evaluation pilots from initial concept through final delivery, ensuring all outputs meet demanding scientific and technical standards.

Continuously refine evaluation methodologies to support emerging capabilities and evolving best practices in frontier AI research.

Requirements

Demonstrated research experience in machine learning evaluation, benchmarking, AI assessment methodologies, or a closely related field.

Proven expertise designing or contributing to AI benchmarks, evaluation research, published datasets, or similar research initiatives relied upon by AI organizations.

Strong understanding of large language model evaluation, including benchmarking methodologies, scoring rubrics, pass rates, contamination risks, construct validity, and discriminative task design.

Deep knowledge of code model evaluation and technical assessment frameworks for advanced AI systems.

Experience working with subject matter experts and maintaining rigorous quality standards across technical research projects.

Excellent analytical, research, and problem-solving skills with a strong attention to scientific rigor.

Outstanding written and verbal communication skills in English.

Ability to work independently in a remote, research-driven environment while collaborating effectively with multidisciplinary teams.

Spanish language proficiency is considered a plus.

Benefits

Fully remote contract position based in Brazil .

Opportunity to work on frontier AI research and evaluation projects with global impact.

High level of ownership and influence over research methodology and benchmark design.

Collaboration with leading AI experts, researchers, and technical specialists.

Flexible remote working environment that supports autonomy and innovation.

Exposure to cutting-edge advancements in large language models and AI evaluation.

How Jobgether works: We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team. We appreciate your interest and wish you the best!  Why Apply Through Jobgether? 

Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

#LI-CL1