Founding AI Scientist

PROVENA PTE. LTD.Singaporemycareersfuturefoilsithe 16/09/2026
Riachtanach:AI

Obsessed about training data? We're building the data infrastructure layer for AI training.

The role

As our founding research engineer, you'll own data curation as a measured discipline across every stage of training, from pre-training to post-training. Working alongside the founders, you'll turn what works into product features — and as one of the first researchers, the methods and data infrastructure you build become the foundation the company runs on.

Examples of work you might do

  • Measure what's in a dataset — quality, provenance, safety, and coverage metrics that predict downstream performance
  • Diagnose where data will hurt a model — decontamination, difficulty annotation, multilingual asymmetries, long-tail gaps
  • Design data curation interventions — pruning, filtering, synthetic augmentation, relabeling — and attribute the gains to specific changes
  • Turn the latest research into running experiments — training and evaluation loops that prove a data hypothesis
  • Ship the methods that work as product features, and share findings as technical reports or papers

What we’re looking for

  • Data obsession — you want to know exactly which properties of a dataset move a model, and you won't trust a result you can't measure
  • Bias toward iteration — you design the smallest experiment that answers the question and move fast on the answer
  • Research that ships — you optimize for models that measurably improve, not for leaderboard wins
  • Product engineer — you own problems end to end and turn research into features that ship, not just results in a notebook
  • Drawn to data-centric AI — you're excited by data-centric methods and the move toward autonomous AI research

Required qualifications

  • Strong machine learning and deep-learning fundamentals
  • Enough software engineering and PyTorch / Jax experience (or willingness to learn) to run ML experiments and build production prototypes
  • Hands-on experience in one or more stages of training and evaluating LLMs/vLLMs
  • Industry or research experience in one or more of: data curation, data pruning & selection, synthetic data generation, curriculum learning, dataset distillation, large-scale language/multimodal training
  • Comfortable reading ML research — sourcing, vetting, and implementing promising ideas from the literature
  • Able to drive applied research independently in a fast, ambiguous, early-stage environment

Nice to have

  • Post-training experience — SFT, preference optimization (DPO, RLVR), or reward modeling
  • Multilingual or multimodal data work
  • Multi-GPU / distributed training experience
  • Open-source, HuggingFace contributions
  • Public technical writing, published research papers