About the Institute of Foundation Models
The Institute of Foundation Models (IFM) at MBZUAI is a research lab dedicated to meaningful foundation model research — building models from scratch, understanding them deeply, and publishing work that shapes the field. You’ll work alongside world-class researchers and engineers on problems that directly define the models we ship.
The Role
Join the PAN world model project — our effort to build world models: foundation models that simulate, predict, and interact with the physical world. As a Research Scientist, you’ll drive the core research behind PAN — large-scale video generation, interactive and action-conditioned world models, and their applications in robotics and embodied AI — and publish at top venues while turning breakthroughs into working systems.
What You'll Do
Conduct original research on video world models, video diffusion models, and action-conditioned generation — from idea to publication and deployment.
Design pre-training and post-training recipes for large-scale diffusion transformers, including scaling-law studies for video pre-training.
Advance world action models / video action models and their applications in robotics and embodied agents.
Develop rigorous evaluation benchmarks for physical accuracy, controllability, and interactivity.
Collaborate with engineering and data teams on large-scale training, data curation, and simulation-based data generation.
What We're Looking For
PhD in Machine Learning, Computer Science, Computer Vision, Robotics, or a related field, with first-author publications at top-tier venues (e.g., NeurIPS, ICML, ICLR, CVPR, ICCV, ECCV, RSS, CoRL).
Research experience with state-of-the-art video generative models and world models (e.g., Cosmos-3, LTX 2.3, Self-Forcing, Lingbot-World, or comparable systems).
Deep expertise in at least one of the following areas:
Full-stack data pipelines — large-scale video data pipelines and/or simulation data collection; annotation and filtering workflows for video / world model training.
Model training & infrastructure — training large-scale diffusion transformers on large GPU clusters.
Rendering engines & simulation — Unreal Engine and Blueprint-based gym environments, game-engine integration, building interactive simulated environments.
World action models & robotics — world action models / video action models, action-conditioned video generation, world-model applications in robotics.
Strong systems and engineering expertise in deep learning frameworks such as PyTorch.
Highly proficient with modern AI coding agents and web-based coding tools (e.g., Claude Code, Codex, Cursor), and skilled at leveraging them to dramatically accelerate research workflows.
Exceptional problem-solving skills and the ability to navigate ambiguity in rapidly evolving research areas.
Nice To Have
Experience accelerating diffusion model inference (distillation, few-step generation, real-time interactive generation).
Experience with visual tokenization and multimodal foundation models.
Experience deploying world models in robotics or embodied-AI settings.