Large Model Quantization Algorithm Engineer

XG TECH PTE. LTD.Singaporemycareersfuturepublished 09/24/2026
Must-have:PythonMobileAI

About Company

Founded in 2022, XG Tech is driving the future of smart vehicles. Its mission is to empower the digital transformation of automobiles, moving from distributed computing to a centralized, cross-domain platform.

XG Tech focuses on the intelligent cockpit—the next frontier of differentiation—while seamlessly integrating advanced driving systems. By reimagining cars as mobile living spaces, XG Tech aligns with the evolving trend of vehicles becoming the “third living space.”

Role Summary

As a Large Model Quantization Algorithm Engineer, you will develop quantization and model compression algorithms for LLMs, VLMs, and video generation models. You will optimize model accuracy, memory efficiency, and inference performance across NPUs, GPUs, and CPUs, bridging the gap between model algorithms and on-device deployment. You will work closely with algorithm, compiler, and hardware teams to bring efficient AI inference technologies into production.

Key Responsibilities

  • Develop and optimize quantization algorithms for LLMs, VLMs, and video generation models, covering PTQ, QAT, and related model compression techniques.
  • Design and evaluate quantization schemes to balance model accuracy, inference performance, and memory efficiency.
  • Perform quantization calibration and error analysis, identifying sources of accuracy degradation and driving optimization solutions.
  • Optimize edge and on-device inference, including Prefill/Decode acceleration, KV Cache management, operator fusion, weight compression, and memory optimization.
  • Adapt and deploy models across heterogeneous hardware, including NPU and CPU platforms, working closely with compiler teams on model conversion, engine compilation, and performance tuning.
  • Develop quantization and model optimization toolkits, including automated quantization workflows, accuracy evaluation, and visualization/debugging tools.
  • Collaborate with model and architecture teams to develop quantization-friendly model architectures, training strategies, and inference optimization techniques.
  • Track and evaluate emerging research in model quantization, compression, sparsity, and efficient inference, and drive relevant techniques into production.

How will you stand out

  • Bachelor’s degree or above in Electronic Engineering, Computer Science, Automation, Operations Research, Statistics, Mathematics, or a related quantitative field.
  • Familiarity with LLM/VLM algorithms and deployment optimization techniques, including model quantization, sparsity/pruning, and inference acceleration frameworks.
  • Prior hands-on experience with PyTorch Quantization-Aware Training (QAT) development is advantageous.
  • Strong proficiency in Python and C++.