Director of AI Data Center Operations

SoftConstructYerevanstaffampublished 10/08/2026
Must-have:Node.jsKubernetesBackendDevOpsAISecurityHybrid

We are seeking an experienced, forward-thinking Director of AI Data Center Operations to oversee our 5MW high-density compute footprint. Unlike traditional mega-scale facilities, this position is designed for an agile engineering leader who can manage a cutting-edge, ultra-dense deployment focused purely on Enterprise AI workloads (large language model fine-tuning, high-throughput inference, and advanced multi-node HPC clustering). This role is part of a landmark AI infrastructure project with a total investment of over USD 100 million, making it one of the largest AI data center initiatives in the region. In this role, you will hold ultimate ownership of the physical deployment, thermal management, and day-to-day operational efficiency of our liquid-cooled GPU clusters. You will act as the critical structural bridge between our physical colocation partners and our internal AI Platform Engineering teams to ensure maximum hardware yield, zero thermal throttling, and seamless power delivery.

Deploy & Manage Liquid Cooling: Oversee the operation and maintenance of direct-to-chip (DLC) liquid cooling loops, Coolant Distribution Units (CDUs), and localized liquid-to-air secondary heat exchangers in a dense environment

Rack-Level Power Strategy: Optimize active power distribution at the rack level, managing ultra-dense power envelopes ranging from 60kW to 100kW+ per rack 60-140across our 5MW footprint

PUE & Efficiency Optimization: Drive continuous physical optimization to lower Power Usage Effectiveness (PUE) and achieve carbon neutrality and environmental sustainability metrics

Colocation SLA Management: Act as primary liaison to wholesale colocation partners, holding suppliers strictly accountable to power, cooling, physical security, and infrastructure uptime SLAs

Hardware Lifecycle Operations: Direct the logistics, installation, staging, provisioning, and decommissioning of advanced AI hardware systems (including NVIDIA DGX/Blackwell/Hopper, Ver Rubin)

Network Fabric Readiness: Collaborate closely with Network Architecture teams to ensure high-bandwidth, ultra-low-latency backend topologies (InfiniBand, RoCEv2) are perfectly integrated and structurally protected

Proactive Telemetry & Monitoring: Implement and maintain centralized environmental telemetry pipelines that connect physical variables (coolant flow rates, pressure, ambient temperature) with software schedulers (Slurm, Kubernetes)

Professional Experience:  7+ years of direct experience in data center operations, high-performance computing (HPC), or critical facilities engineering, with at least 3 years in a direct leadership or managerial capacity

Advanced Thermal Expertise:  Deep technical understanding of direct-to-chip liquid cooling loops, Coolant Distribution Units, rear-door heat exchangers (RDHx), and general facility MEP (Mechanical, Electrical, Plumbing) designs

High-Density Hardware Literacy:  Direct experience deploying and maintaining high-power density compute environments (NVIDIA HGX/DGX, advanced liquid-cooled chassis, or specialized OEM architectures)

Colocation Contract Proficiency:  Proven experience negotiating and managing SLAs, Master Service Agreements (MSAs), and joint operating procedures inside high-tier commercial colocation facilities

Education:  Bachelor's degree in Mechanical Engineering, Electrical Engineering, Computer Science, or equivalent technical discipline / equivalent practical experience

Transition Management:  Experience guiding active environments through transitioning from legacy air cooling to hybrid air-and-liquid or 100% direct-to-chip liquid-cooled architectures

Software Infrastructure Literacy:  Functional knowledge of container orchestrators (Kubernetes) and AI training scheduling architectures (Slurm)

Strategic Partnerships:  Established, active relationships with key component providers (CDU vendors, quick-disconnect suppliers) and server OEMs