Data Engineer - Legal Data, Analysis, Retrieval & Automation
Data Engineer for RegTech startup - Legal Data, Analysis, Retrieval & Automation
We’re looking for an engineer with a background in engineering, mathematics, physics or another STEM field. You’ll build data ingestion pipelines and the analytical tools that verify data quality. Then you’ll put that data to work in RAG systems and automation flows for financial institutions across Europe. If you like hard data problems where getting it right matters more than getting it big, read on.
What we do
Enfx is a SaaS platform where risk & compliance teams from financial companies manage their work in one place, built on top of the regulations they follow. We track regulations such as DORA, MiCA and GDPR, the acts under them and the proposals that would change them, as well as Danish national law and enforcement decisions. Our product is only as good as the data under it.
Why join us
We’re a small tech startup from 2025 getting traction with 6 clients and lots of exciting ideas and plans that will change how financial companies manage regulations. You will be one of our early hires and will work directly with the founders, so you’ll help decide how things are built, not just build them. You’ll own the data layer from source to answer to action. As we grow, so will your role. Early hires will have the opportunity to shape the team, take on meaningful responsibility, and share in the upside as we build and grow the company together.
What the job involves
Building ingestion pipelines
- Querying and scraping a range of endpoints for regulatory data
- Parsing HTML and XML legal texts into their structure (articles, paragraphs, annexes, cross-references)
- Building LLM pipelines for structured extraction, e.g. “which articles of which regulation does this proposal change?”
- Making pipelines that run unattended, are safe to re-run, and fail loudly without corrupting anything
- Handling how public sources really behave: bot protection, timeouts, late translations and PDF-only documents
Proving the data is right
- Expanding our in-house tools for analysing what we ingest and what’s in the database
- Building everything from simple completeness checks to statistical checks on the connections between regulations
- Measuring how well our AI-based parsing and extraction performs
- Estimating the likely outcomes of EU legislative proposals
Building retrieval, RAG and automation
- Building retrieval systems that answer questions about a vast and hard-to-interpret regulatory landscape
- Helping compliance teams understand what a regulatory change means for them
- Automating workflows that are highly manual today
- Scanning clients’ systems and processes to assess how compliant they are
Must-haves
- A degree in engineering, physics, mathematics, computer science or another quantitative field
- 2–5 years writing production Python, or PHP where you shipped real, tested code
- Solid SQL and relational data modelling
- A scientific approach to data: form a hypothesis, measure, look for the counter-example
- Experience parsing semi-structured data (HTML, XML or JSON)
- The judgement to critically review AI-generated code rather than just accept it
- Care with production systems
- Fluent written and spoken English
Good-to-haves
- Danish, Swedish or Norwegian. Much of our data is Danish, and checking it means reading it.
- Information retrieval experience: embeddings, BM25, ranking and evaluation metrics
- TypeScript
- RDF, SPARQL or other linked-data technologies
- Experience with LLM-based extraction or evaluation
- Curiosity about how legislation is structured (no legal background needed)
General information
- Salary: DKK 35.000–45.000 per month, depending on experience
- Location: Copenhagen, CPH Fintech Labs. Hybrid: you can work from home up to two days a week.
- Start: As soon as possible