Founding AI Scientist
Indispensable :AI
Obsessed about training data? We're building the data infrastructure layer for AI training.
The role
As our founding research engineer, you'll own data curation as a measured discipline across every stage of training, from pre-training to post-training. Working alongside the founders, you'll turn what works into product features — and as one of the first researchers, the methods and data infrastructure you build become the foundation the company runs on.
Examples of work you might do
- Measure what's in a dataset — quality, provenance, safety, and coverage metrics that predict downstream performance
- Diagnose where data will hurt a model — decontamination, difficulty annotation, multilingual asymmetries, long-tail gaps
- Design data curation interventions — pruning, filtering, synthetic augmentation, relabeling — and attribute the gains to specific changes
- Turn the latest research into running experiments — training and evaluation loops that prove a data hypothesis
- Ship the methods that work as product features, and share findings as technical reports or papers
What we’re looking for
- Data obsession — you want to know exactly which properties of a dataset move a model, and you won't trust a result you can't measure
- Bias toward iteration — you design the smallest experiment that answers the question and move fast on the answer
- Research that ships — you optimize for models that measurably improve, not for leaderboard wins
- Product engineer — you own problems end to end and turn research into features that ship, not just results in a notebook
- Drawn to data-centric AI — you're excited by data-centric methods and the move toward autonomous AI research
Required qualifications
- Strong machine learning and deep-learning fundamentals
- Enough software engineering and PyTorch / Jax experience (or willingness to learn) to run ML experiments and build production prototypes
- Hands-on experience in one or more stages of training and evaluating LLMs/vLLMs
- Industry or research experience in one or more of: data curation, data pruning & selection, synthetic data generation, curriculum learning, dataset distillation, large-scale language/multimodal training
- Comfortable reading ML research — sourcing, vetting, and implementing promising ideas from the literature
- Able to drive applied research independently in a fast, ambiguous, early-stage environment
Nice to have
- Post-training experience — SFT, preference optimization (DPO, RLVR), or reward modeling
- Multilingual or multimodal data work
- Multi-GPU / distributed training experience
- Open-source, HuggingFace contributions
- Public technical writing, published research papers