HPC Engineer

MOHAMED BIN ZAYED UNIVERSITY OF ARTIFICIAL INTELLIGENCEAbu Dhabigulftalentavaldatud 15.07.2026
Nõutav:PythonGitAWSAzureGoogle CloudDockerCloudAI

Application open:

Full-time

The Institute for Foundation Models (IFM) at MBZUAI operates some of the world’s largest AI supercomputing environments, supporting frontier AI research and foundation model development across thousands of GPUs.

We are seeking an HPC engineer to join our growing infrastructure team. This role is suitable for recent graduates and early-career engineers who are passionate about Linux systems, large-scale computing, distributed systems, and AI infrastructure.

Key responsibilities

Support operation and maintenance of large-scale GPU computing clusters.

Assist researchers with job submission, troubleshooting, and resource utilization.

Monitor cluster health, performance, and availability.

Troubleshoot Linux, hardware, storage, networking, and software issues.

Support Slurm administration and user management.

Assist with cluster deployment, upgrades, and validation.

Develop scripts and automation tools.

Maintain technical documentation and operational procedures.

Participate in incident response and operational support.

Collaborate with researchers, vendors, and internal teams.

Academic qualifications

Bachelor’s degree in computer science, computer engineering, electrical engineering, software engineering, information technology, mathematics, physics, or related disciplines.

Professional experience required

Preferred:

Linux administration experience.

Python, Bash, Go, or C/C++ programming.

Networking fundamentals.

Cloud platforms (Azure, AWS, GCP).

Containers (Docker, Apptainer, Enroot).

Git and software development workflows.

AI/ML infrastructure exposure.

HPC, distributed systems, or research computing experience.