Senior AI Engineer
Riachtanach:PythonAWSAzureCloudDataAISeniorJunior
Job Description: What will you do:
- Experimentation & Evaluation
- Understand the business problem, POC objectives, and evaluation metrics.
- Design experiments to test different model configurations, prompts, or retrieval strategies.
- Analyse Gen AI outputs for quality, accuracy, and alignment with requirements; identify common failure modes (hallucination, bias, irrelevant answers, factual errors).
- Support SMEs in defining ground truth benchmarks for evaluation.
- Data Preparation & Pipelines
- Profile and clean sample datasets for experimentation (lightweight data prep).
- Build and test simple pipelines for data ingestion, prompt construction, and output evaluation.
- Gen AI & Agentic Techniques
- Work with foundation models via AWS Bedrock, Google Vertex AI, or Azure AI Foundry depending on engagement cloud posture.
- Apply working knowledge of China-origin models (DeepSeek, Qwen, GLM) as increasingly relevant, cost-effective alternatives.
- Apply agentic orchestration frameworks such as AWS Strands, LangGraph, or equivalent, for designing and testing multi-step agent workflows.
- Apply prompt strategies, prompt engineering patterns, and RAG design (chunking, embeddings, retrieval evaluation); support ingesting/vectorising content to knowledge bases.
- Provide insights and recommendations to improve model performance in quick iterations, including fine-tuning approaches where applicable.
- FDE & Development/Maintenance Coverage
- During FDE engagements: rapidly test candidate models, prompts, and retrieval strategies, giving the team fast, evidence-based go/no-go signals.
- During system development & maintenance engagements: support ongoing model/prompt tuning and monitoring as applications move toward production.
- Collaboration
- Collaborate with developers on integrating models into the POC workflow, and work closely with PM, devs, and SMEs to refine data and prompts.
- Partner with the Data Scientist on evaluation methodology where classical statistical baselines are in play, and with the AI/LLM Specialist when an engagement moves toward production-grade evaluation.
- Document experiments briefly but clearly (hypothesis → result → conclusion).
Role Levels We Are Hiring For
We are hiring at two levels for this role. All responsibilities above apply to both; the distinction is in scope of ownership, years of experience, and seniority of judgement expected.
AI Engineer
- 4–5 years of hands-on experience in AI/ML or Gen AI engineering. Runs experiments and prototypes independently within a defined POC/POV scope, under guidance from a Senior AI Engineer or AI/LLM Specialist.
- Executes rapid experimentation cycles for one engagement at a time; escalates ambiguous evaluation calls to senior team members.
Senior AI Engineer
- 6+ years of hands-on experience, including prior ownership of experimentation strategy for complex or ambiguous problem statements. Sets the experimentation approach across multiple engagements and mentors junior AI Engineers.
- Advises PMs and stakeholders directly on feasibility and experimentation trade-offs; represents technical experimentation findings in client conversations.
Qualifications
The ideal candidate should possess:
- 4+ years hands-on experience in AI/ML or Gen AI engineering (see Role Levels for the split between AI Engineer and Senior AI Engineer).
- Understanding of Gen AI concepts (tokenization, embeddings, RAG, prompting, evaluation).
- Familiar with at least one major cloud AI service (AWS Bedrock, Google Vertex AI, or Azure AI Foundry); working knowledge of others a plus.
- Working knowledge of the China AI model landscape (DeepSeek, Qwen, GLM) a strong plus.
- Familiarity with agentic orchestration frameworks (AWS Strands, LangGraph, or equivalent), for designing and testing multi-step agent workflows.
- Ability to do rapid experimentation rather than perfect models.
- Basic proficiency in Python and Gen AI tools (e.g., model SDKs, vector DBs).
- Analytical mindset: can quantify subjective output (accuracy, relevance, readability).
- Good data wrangling skills to prepare small datasets quickly.
Preferred Qualifications
- Generative AI Leader or Machine Learning Engineer certification, or equivalent.
- Exposure to LLMOps practices (model monitoring, versioning) for production transition.
- Familiarity with model fine-tuning techniques.
- Exposure to regulated government cloud environments.
Tech Stack (Illustrative)
- Languages: Python
- LLM Runtime: AWS Bedrock, Google Vertex AI, Azure AI Foundry; DeepSeek/Qwen/GLM (China stack)
- Agentic Frameworks: AWS Strands, LangGraph
- Data & Eval: Model SDKs, vector DBs, pandas/Jupyter-style tooling for experimentation