Tech lead Expert Data Retrieval RAG - Freelance (M/F)
THE CONTEXT
Join a data adventure at the heart of Legaltech: our client is building the search and indexing engine that will provide access, tomorrow, to a massive and heterogeneous legal heritage for several teams across Europe.
Your mission is not to write a simple script: it is to design robust data pipelines, capable of ingesting, cleaning, and indexing enormous legal corpora, and then exposing fast, relevant, and hybrid (semantic + lexical) search at scale.
YOUR DAILY MISSIONS
- Collection, ingestion & indexing
- Design ingestion pipelines from heterogeneous sources (publishing, open data, legal databases): parsing, cleaning, chunking, normalization.
- Choose and industrialize embedding models, feed and maintain vector databases.
- Guarantee index freshness, scalability on massive volumes, and traceability (lineage, versioning).
- Search & relevance (Retrieval)
- Combine lexical and semantic search (BM25 + embeddings), reranking, and metadata filtering.
- Implement relevance metrics (recall, precision, nDCG, MRR) and maintain them over time.
- Structure the retrieval phase that feeds the LLM (RAG: context, citations, de-duplication, context window management).
- Instrument pipelines (logs, metrics, quality traces) and guarantee data security/compliance (GDPR, country-based partitioning).
- Technical leadership & governance
- Promote a Clean Code, SOLID, and automated testing culture.
- Contribute to the group's technical influence (articles, meetups, open-source).
EDUCATION
Master's, PhD, or equivalent.
EXPECTED SKILLS
- Proven experience in data engineering / information retrieval on massive production volumes.
- Confirmed Retrieval expertise: indexing, embeddings, hybrid search - and the ability to measure then improve relevance.
- Strong Data culture: reproducible pipelines, data quality and lineage, dataset/index versioning.
- Solid AI / ML culture (embeddings, NLP, good understanding of LLM, without necessarily being the serving expert).
- A plus: appetite for other languages (notably Java) for cross-functional topics.
- Fluent English (daily European context).
SOFT SKILLS
- Product vision: your indexes and your search engine are managed as a product; other developers are your customers.
- Technical leadership & pedagogy: ability to evangelize best practices in data/retrieval/AI.
- Obsession with quality: absolute rigor regarding data security (Legaltech requirement) and service reliability.
- Innovation mindset: curiosity for advancements in AI, NLP, and data.
- Communication: ability to simplify complex concepts.
Python 3.11+ (mastery required), GoIndexingVector databases (pgvector, Weaviate, Qdrant, OpenSearch), hybrid search / BM25Embeddings & NLPsentence-transformers, embedding models, tokenization, document parsingRAG & LLMLangChain, RAG, reranking, OpenAI, Mistral, BedrockData & infraPostgreSQL, Airflow / Spark, Kafka, Redis, Docker, CI/CD, K8sTests & qualityPytest, relevance evaluation (recall, nDCG, MRR)