Server Engineer - (GPU Server)

PT Aktualisasi Gratia Talenta IndonesiaJakarta Pusat, DKI Jakartaglintspublished 08/27/2026
This job is no longer listed
The source has removed this listing — applying via the original link is no longer possible.
Must-have:DevOpsAISecurity

Job Description:

  • Design, deploy, configure, and maintain server infrastructure, particularly GPU-based environments supporting AI and Machine Learning workloads.
  • Manage and maintain GPU Servers, including hardware installation, configuration, monitoring, troubleshooting, and performance optimization.
  • Support the implementation and operation of AI Infrastructure, including compute, storage, networking, and supporting infrastructure components.
  • Install, configure, and maintain server operating systems, drivers, firmware, and system-level software required for GPU and AI workloads.
  • Configure and optimize NVIDIA GPU environments, including GPU drivers, CUDA, and related NVIDIA software components.
  • Perform server health checks, capacity planning, performance monitoring, and preventive maintenance.
  • Troubleshoot hardware and software issues involving servers, GPUs, storage, networking, operating systems, and system components.
  • Monitor GPU utilization, temperature, memory usage, system performance, and overall infrastructure availability.
  • Support the deployment and configuration of AI/ML computing environments for development, testing, and production workloads.
  • Work closely with AI/ML Engineers, DevOps Engineers, Network Engineers, System Administrators, and other technical teams to ensure infrastructure readiness.
  • Implement infrastructure security, access control, backup, monitoring, and availability best practices.
  • Perform server installation, rack-and-stack activities, cabling, hardware replacement, and infrastructure upgrades when required.
  • Maintain technical documentation covering server configurations, infrastructure architecture, operational procedures, and troubleshooting activities.
  • Conduct incident investigation, root cause analysis, and corrective actions to maintain system reliability.
  • Support infrastructure migration, expansion, and technology refresh projects.
  • Ensure infrastructure complies with operational standards, security requirements, and project specifications.

Requirements:

  • Minimum D3 in Computer Science, Information Technology, Computer Engineering, Electrical Engineering, or a related field.
  • Hands-on experience as a Server Engineer, System Engineer, Infrastructure Engineer, Data Center Engineer, or similar role.
  • Proven experience working on GPU Server projects and AI Infrastructure.
  • Strong understanding of server hardware, including CPU, RAM, GPU, RAID, storage, power supply, and server components.
  • Experience with NVIDIA GPU Servers and NVIDIA GPU technologies.
  • Familiarity with CUDA, NVIDIA GPU Drivers, NVIDIA GPU monitoring, and GPU computing environments.
  • Good knowledge of Linux/Windows Server operating systems.
  • Understanding of server virtualization, storage, networking, backup, monitoring, and high-availability concepts.
  • Experience with data center infrastructure and server deployment is highly preferred.
  • Strong troubleshooting and problem-solving skills for both hardware and software infrastructure.

Skills: Nvidia CUDA, Artificial Intelligence, Server Testing, Gpu Server, Machine Learning, Cdcp, IT Infrastructure Management