HPC High Performance Computing IT Infra Engineer

PośrednikD L RESOURCES PTE LTDSingaporemycareersfutureopublikowano 25.08.2026
Wymagane:PythonAWSAzureDockerKubernetesCloudDevOpsCI/CDSecuritySeniorHybrid

Client : Research & Education Sector Key Responsibilities Support day-to-day operations of HPC clusters, including compute nodes, storage systems, and high-speed networking infrastructure Monitor system performance, workload execution, job scheduling, and resource utilization across distributed environments * Perform installation, configuration, patching, and maintenance of Linux/Unix operating systems (RHEL, CentOS, Ubuntu) * Administer physical and virtualized server environments (x86 architecture, VMware/KVM) * Support workload management systems and job schedulers (Slurm, PBS, LSF) for batch processing and resource allocation * Conduct system health checks, log analysis, troubleshooting, and incident management (root cause analysis) * Manage user accounts, access control, and authentication systems (LDAP, Active Directory integration) * Assist in provisioning and configuration of HPC environments, including software stack deployment and environment modules * Collaborate with senior engineers on cluster optimization, scaling, and performance tuning initiatives * Maintain system documentation, operational procedures, and runbooks Required Skills & Experience 1–3 years of experience in system administration, infrastructure support, or IT operations Hands-on experience with Linux/Unix systems administration and command-line environments * Basic exposure to HPC, distributed systems, or parallel computing environments * Strong understanding of networking fundamentals (TCP/IP, DNS, SSH, firewall concepts) * Knowledge of storage technologies (NAS, SAN, distributed/parallel file systems – basic awareness) * Scripting and automation using Bash/Shell and/or Python * Familiarity with monitoring and logging tools (e.g., Nagios, Zabbix, Prometheus, Grafana, ELK Stack) * Strong troubleshooting, analytical thinking, and incident resolution skills Preferred Skills Exposure to GPU computing environments (NVIDIA GPUs, CUDA – basic awareness) Familiarity with cloud platforms and HPC workloads on cloud (AWS, Azure) * Awareness of container technologies (Docker) and basic DevOps practices Project / Environment Tech Stack. (Mostly On-Prem) The project environment is primarily on-premise, with approximately 80% of the infrastructure hosted on Dell and servers, and approximately 20% involving AWS-based HPC/GPU server environments. This provides exposure to both traditional data centre infrastructure and cloud-based GPU/HPC workloads. of: 80% On-Premise HPC Infrastructure Dell servers, Huawei servers, physical data centre infrastructure

x86 server architecture, CPU-based compute infrastructure, bare-metal servers

Virtualized server environments using VMware / KVM where applicable

Linux/Unix-based systems including RHEL, CentOS, Ubuntu, AIX

HPC cluster components across compute, storage, and networking layers

High-performance storage / parallel file systems such as IBM Spectrum Scale (GPFS), Lustre, BeeGFS; with NAS / SAN / NFS exposure where applicable

Cluster networking and connectivity including TCP/IP, DNS, SSH, firewall concepts, high-throughput Ethernet; InfiniBand / RDMA exposure preferred

HPC workload scheduling and batch processing using Slurm, PBS, or LSF

Monitoring and logging tools such as Nagios, Zabbix, Prometheus, Grafana, ELK Stack, or Splunk

20% AWS HPC / GPU-Related Environment AWS cloud infrastructure supporting HPC / GPU-related workloads

AWS GPU server exposure, including NVIDIA GPU-based compute instances

GPU computing environment exposure, including NVIDIA CUDA, NVIDIA drivers, GPU utilization, and GPU workload monitoring where applicable

Cloud HPC workload support involving compute, storage, networking, and security configurations

Exposure to AWS HPC services or related tools such as AWS ParallelCluster, AWS Batch, EC2 GPU instances, EBS / FSx / S3 storage, VPC, IAM, and security groups

Container and DevOps exposure where applicable, including Docker, Kubernetes, and basic CI/CD or automation practices

Hybrid HPC environment exposure involving on-premise infrastructure integrated with AWS-based compute or GPU resources