Senior Lead Network Engineer

IntegrantCairogulftalentpublished 07/21/2026
Must-have:AISecuritySeniorLeadPrincipal

Overview

Integrant is seeking a senior lead network & infrastructure engineer with 14+ years of experience to provide technical leadership in a fast-paced, complex HPC and AI environment. This is a multi-disciplinary senior role: the core is deep network engineering across Ethernet and InfiniBand fabrics, ideally complemented by hands-on experience in high-performance storage and/or Linux systems operations (SysOps). The senior lead owns fabric architecture and performance, acts as the highest technical escalation point, and works across engineering, platform, storage, and client teams. Adaptability, ownership, and clear communication are key to success in an environment where network performance is critical.

Responsibilities

Develop network configurations and architectures

Operate, maintain, and support Ethernet and InfiniBand networks in a high-performance computing (HPC) and AI environment

Perform ongoing maintenance, upgrades, and lifecycle management of network equipment

Monitor network health, performance, and capacity to ensure reliable, low-latency data flow

Respond to and resolve network and server-related incidents in a timely manner

Run hardware diagnostics and coordinate replacement of failing network components

Support and maintain Linux-based HPC and AI platforms across a wide range of technologies

Collaborate with senior network engineers, software teams, and platform teams on network efficiency, reliability, and security

Assist with configuration, deployment, and operational support of InfiniBand and Ethernet fabrics

Develop and maintain operational documentation, including configuration examples, build guides, and best practices

Support on-site staff during hardware updates, card replacements, and infrastructure changes

Stay current with advancements in data center networking, HPC interconnects, and AI infrastructure technologies

Work within the client ticketing / IT service management system (e.g., TopDesk) to manage incidents and service requests to SLA

Build and maintain automation and tooling (scripting, monitoring integrations, infrastructure-as-code) to improve operational efficiency

Collaborate with software, platform, storage, and client teams on efficiency, reliability, and security

Own the quality of operational documentation: configuration examples, build guides, runbooks, and best practices

Lead design reviews and knowledge-sharing

Work Conditions

Participate in a weekly on-call rotation and respond to network and infrastructure issues after hours when required.

Requirements

14+ years of hands-on experience supporting enterprise or data center-scale networks

Experience working in HPC, AI/ML, or performance-sensitive environments

Practical experience administering InfiniBand (Mellanox/NVIDIA) and Ethernet (Cumulus, SONiC) networks

Strong understanding of data center networking concepts, including servers, storage, and high-speed interconnects

Solid knowledge of Layer 2 and Layer 3 networking, including routing and switching fundamentals

Installing, monitoring, and maintaining very large-scale data center networks

Low-latency, high-bandwidth fabric support and performance tuning for distributed compute and GPU workloads

VXLAN/EVPN architectures and routing protocols such as BGP and OSPF

Exposure to communication libraries such as NCCL, UCX, and MPI

Network management and monitoring tools: UFM, OpenSM, NetQ, or similar

Ability to troubleshoot and resolve network issues in complex, distributed environments

Strong documentation and communication skills

Proven ability to work effectively as part of a team and provide operational support

Preferred (Multi-Skill) Qualifications

Storage: Hands-on experience with high-performance / parallel storage environments (e.g., Lustre, GPFS/Spectrum Scale, BeeGFS, Ceph, NVMe-oF), including storage networking and I/O performance troubleshooting

SysOps / Linux systems: Production Linux systems administration at scale — provisioning, configuration management (Ansible/Salt), kernel/network stack tuning, schedulers (Slurm), containerization

Benefits

Salary paid in USD

Six-month career advancing opportunities

Supportive and friendly work environment

Premium medical insurance [employee + family]

English language development courses

Interest-free loans paid over 2.5 years

Technical development courses

Employment referral program

Premium location in Maadi

Social insurance