Senior Localization Engineer / Research Engineer, Robotics AI (Visual Navigation)

GrabSingaporeJob.boveröffentlicht 11.09.2026
Muss:PythonAILeadHybrid

About Grab and Our Workplace Grab is Southeast Asia's leading superapp. From getting your favourite meals delivered to helping you manage your finances and getting around town hassle-free, we've got your back with everything. In Grab, purpose gives us joy and habits build excellence, while harnessing the power of Technology and AI to deliver the mission of driving Southeast Asia forward by economically empowering everyone, with heart, hunger, honour, and humility.

Get to Know Our Team The Robotics Technology team is a core part of Grab's long-term vision to build urban embodied AI. We take full ownership of the product lifecycle — from perception and navigation research to real-world deployment on our autonomous delivery fleet across Southeast Asian cities. This is a fast-moving, multidisciplinary environment where robotics researchers, ML engineers, and hardware specialists collaborate to solve practical challenges at scale. Based in Singapore, you will have the opportunity to work on frontier autonomy research, deploy solutions in complex real-world environments, and directly shape the future of last-mile logistics.

Get to Know the Role As a Research Engineer on the Spatial Intelligence team, you will advance the core visual navigation and VLN capabilities of our autonomous platforms. Your research will focus on enabling robots to understand, reason about, and navigate complex urban environments using vision-based and multimodal learning approaches. You will bridge cutting-edge academic research with production-grade deployment, working closely with systems engineers to bring novel algorithms into our operating fleet. You will report to the Head of Engineering and work onsite at Grab's One North office.

The Critical Tasks You Will Perform

  1. Visual Navigation & VLN Research (50%)

Design and implement novel deep learning models for visual navigation, Vision-Language Navigation (VLN), and visual place recognition in real-world outdoor environments

Develop and evaluate multimodal perception pipelines combining RGB, depth, semantic segmentation, and language-conditioned navigation signals

Research robust scene understanding techniques for dynamic urban environments: pedestrian-dense streets, repetitive urban facades, and GNSS-degraded corridors

Investigate end-to-end learning approaches and modular architectures for navigation policy learning, with a focus on generalisation across unseen environments

Publish and engage with the research community; contribute findings back to internal and open-source repositories

  1. Algorithm Integration & Validation (30%)

Implement and adapt state-of-the-art VLN and visual navigation algorithms (e.g., R2R, REVERIE, EmbodiedBERT derivatives) for deployment on real robot hardware

Build performance evaluation frameworks and benchmarks tailored to outdoor, last-mile delivery scenarios

Collaborate with Perception and Planning/Control teams to integrate visual navigation outputs into the broader autonomy stack

  1. Research Iteration & Collaboration (20%)

Drive closed-loop dataset collection and annotation pipelines to support continuous model improvement

Partner with hardware and systems teams on sensor selection, calibration, and data quality for vision-based pipelines

Present research findings to internal stakeholders and contribute to the team's research roadmap

What Essential Skills You Will Need Education: PhD in Computer Science, Robotics, AI, Computer Vision, or a closely related field

Research depth: Strong publication record or demonstrable research contributions in one or more of: Visual Navigation, Vision-Language Navigation (VLN), Embodied AI, or Robot Perception

Technical expertise: Proficiency in at least two of the following: Vision-Language Navigation: familiarity with VLN benchmarks (R2R, REVERIE, SOON), instruction-following models, and cross-modal grounding

Visual Place Recognition / Localisation: experience with image-based retrieval, learned descriptors, or topological navigation

Deep learning for perception: CNNs, Transformers, multimodal architectures (ViT, CLIP, LLaVA variants) applied to robotics or embodied AI

Reinforcement Learning for navigation: policy learning, reward shaping, sim-to-real transfer

Engineering: Proficient in Python and PyTorch (or JAX/TensorFlow); experience with ROS/ROS 2 for robot integration

Research rigour: Ability to design, execute, and critically evaluate experiments; strong written communication for internal research reports and publications

Good to Have Experience deploying learned navigation models on physical robot platforms (not simulation only)

Familiarity with LiDAR-visual fusion or hybrid metric-topological navigation approaches

Experience with large-scale vision-language models (LLMs/VLMs) applied to robotic instruction following

Prior internship or research collaboration with an autonomous vehicle or robotics company

Demonstrated proficiency in leveraging AI tools to accelerate research workflows

Life at Grab We care about your well-being at Grab, here are some of the global benefits we offer: We have your back with Term Life Insurance and comprehensive Medical Insurance. With GrabFlex, create a benefits package that suits your needs and aspirations. Celebrate moments that matter in life with loved ones through Parental and Birthday leave , and give back to your communities through Love-all-Serve-all (LASA) volunteering leave We have a confidential Grabber Assistance Programme to guide and uplift you and your loved ones through life's challenges. What We Stand For at Grab We are committed to building an inclusive and equitable workplace that enables diverse Grabbers to grow and perform at their best. As an equal opportunity employer, we consider all candidates fairly and equally regardless of nationality, ethnicity, religion, age, gender identity, sexual orientation, family commitments, physical and mental impairments or disabilities, and other attributes that make them unique.