Data Scientist / Applied ML Engineer
We are looking for a Data Scientist / Applied ML Engineer to join our team and build and improve the intelligence layer of our platform. You will design and orchestrate AI workflows that combine live web research, LLMs, and machine learning models to generate insights. This is a hands-on engineering role, you will design prompts, build async pipelines, train and deploy ML models, and ship production features end to end.
Design, build, and orchestrate AI workflows that combine live web search, LLM scoring, and structured reasoning pipelines
Build and tune prompt systems for structured LLM outputs - scoring rubrics, reasoning generation, hallucination prevention
Train, evaluate, and improve ML/deep learning models for prediction tasks using structured data
Deploy ML models to production as inference endpoints/services, manage versioning and rollback, and continuously monitor model performance (drift, accuracy decay, data quality), retraining and updating models as needed
Build async data collection pipelines that gather data from external sources in real time
Monitor LLM behavior in production using observability tooling and continuously improve prompt quality
Design scoring systems that are explainable and consistent
Improve data quality by building smart deduplication, classification, and validation logic
3+ years of professional experience in Machine Learning / AI, with proven experience delivering production-grade ML solutions
Strong Python skills - async/await, asyncio, production-quality code
Experience working with LLM APIs and prompt engineering
Understanding of ML/deep learning fundamentals - feature engineering, model training, evaluation, hyperparameter tuning
Experience with gradient boosting models (e.g., CatBoost, XGBoost, LightGBM) for prediction tasks
Experience with MLOps - deploying models to production, versioning, and monitoring model performance over time
Experience with web scraping and data extraction at scale
Comfortable working with relational databases and async database drivers
Strong debugging skills - you can trace a bug through a multi-step async pipeline
Product sense - you understand that AI outputs need to be explainable and trustworthy to end users
Experience with deep learning frameworks (e.g., PyTorch, TensorFlow)
Experience fine-tuning and serving transformer-based models
Familiarity with RAG (Retrieval-Augmented Generation) systems
Experience with AWS RDS, S3, and SageMaker
Experience with hyperparameter optimization frameworks