AI Researcher – Multilingual Data

Jobgether· Brussels (Firmensitz, recherchiert)· lever· offentliggjort 27.07.2026
Skal:PythonQA/TestAI

Accountabilities Design and conduct research focused on multilingual datasets, including data collection, filtering, deduplication, quality assessment, and optimization.

Develop innovative strategies for low-resource and long-tail languages through advanced sampling, data augmentation, and curriculum learning techniques.

Research and improve multilingual large language models by enhancing cross-lingual transfer, alignment, robustness, and representation learning.

Build, maintain, and refine multilingual evaluation benchmarks to measure model quality and performance across languages.

Collaborate closely with machine learning engineers and researchers to influence training pipelines, model architectures, and production deployment strategies.

Publish research findings at leading AI and NLP conferences while contributing to open-source initiatives when appropriate.

Translate research outcomes into practical improvements that enhance production-ready AI systems.

Requirements

Advanced background in Natural Language Processing, Machine Learning, Artificial Intelligence, or a closely related field.

Proven research experience in multilingual or cross-lingual language modeling with publications at recognized conferences or journals such as ACL, EMNLP, NeurIPS, ICML, or ICLR.

Hands-on experience working with large-scale multilingual text datasets and modern machine learning workflows.

Strong understanding of multilingual tokenization, vocabulary design, transfer learning, multilingual representation learning, dataset quality assessment, filtering techniques, and bias mitigation.

Proficiency in Python and modern deep learning frameworks such as PyTorch or JAX.

Ability to work independently, manage research initiatives, and deliver high-quality results in a fast-moving startup environment.

Experience with low-resource languages, non-Latin scripts, multilingual evaluation benchmarks (XTREME, FLORES, TyDi QA), open-source NLP projects, or large language model training is considered a strong advantage.

Benefits

Competitive compensation package.

Meaningful equity opportunity within an early-stage, high-growth company.

Significant ownership over research direction and technical decision-making.

Opportunity to balance academic research with real-world production impact.

Access to large-scale multilingual datasets, modern AI infrastructure, and rapid experimentation cycles.

Collaborative environment that values innovation, research excellence, and continuous learning.

Opportunity to publish at leading international AI and NLP conferences while contributing to impactful open-source initiatives.

How Jobgether works: We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team. We appreciate your interest and wish you the best!  Why Apply Through Jobgether? 

Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

#LI-CL1