AI Engineer – LLM Algorithm Engineer (Agentic Commerce)
Job Description: Lead continuous pre-training and post-training of large language models (Dense and MoE architectures) for vertical business domains, including domain-specific data synthesis, multi-stage data curation, and monthly iteration to support downstream business integration.
Design and build autonomous agents (e.g., Plan-and-Execute agents) that automatically retrieve, extract, and synthesize domain knowledge from large-scale multilingual/unstructured data sources, producing high-confidence corpora for continued pre-training.
Develop and iterate RL-based post-training methods (e.g., CPT, SFT, DAPO) to improve model performance on domain-specific QA, reasoning, and knowledge-modeling tasks.
Build and optimize multimodal LLM applications for content generation and understanding, including product copywriting, semantic search recall, and business potential/CTR prediction.
Design end-to-end agent workflows integrating query understanding, content generation, validation, and downstream product/search systems.
Build generative representation-learning pipelines (e.g., token-based generative pre-training on behavioral sequences) to produce reusable embeddings for downstream ranking and recommendation models.
Collaborate cross-functionally to translate business requirements into scalable model training and agent system design; continuously monitor production metrics to guide model and agent optimization.
Requirements:
Master's degree or above in Computer Science, Natural Language Processing, Artificial Intelligence, Electrical and Electronics Engineering, Signal Processing or a related field.
Minimum 5 years of full-time industry experience in LLM/ML algorithm engineering, including hands-on experience with large-scale model continuous pre-training (Dense and/or MoE architecture) for vertical business domains.
Hands-on experience designing and building autonomous agent systems (e.g., using LangGraph or similar frameworks) for automated knowledge retrieval, extraction, and synthesis pipelines, combined with hands-on experience in LLM post-training techniques (SFT, RL-based methods such as DAPO) and data-mixture optimization techniques for pre-training data curation.
Experience fine-tuning and deploying multimodal large models for content generation, semantic recall, or business potential prediction, with demonstrated production impact on key business metrics (e.g., CTR, conversion).
Good programming and engineering skills; solid foundation in classic ML techniques (e.g., LightGBM, XGBoost, Bayesian modeling) is a plus.
Good problem-solving skills; able to independently drive projects from research through to large-scale production deployment.
Prior experience in multimodal retrieval/recommendation systems (e.g., cross-modal contrastive learning for content matching) will be a strong plus.