NRE/Network Architect
Job Summary We are seeking a seasoned Network Reliability Engineer (NRE) with deep expertise in designing, operating, and optimizing hybrid & cloud network environments. You will focus on ensuring high availability, scalability, performance, and resilience of enterprise networks spanning on-premises data centers, SD-WAN, ACI, and major cloud platforms (AWS, Azure). This role combines network engineering with Reliability Engineering principle, emphasizing automation, observability, proactive incident prevention, and rapid recovery to meet stringent SLAs for mission-critical applications. Responsibilities · Design and implement highly reliable network infra. · Define and track Service Level Objectives (SLOs), Service Level Indicators (SLIs) and error budgets. · Lead major incident response,root cause analysis (RCA), and post-mortem reviews. · Implement blamelesspost-mortems and drive corrective actions to prevent recurrence. · Build automation scripts,tools, and self-healing mechanisms for network provisioning, configurationmanagement, monitoring, and failover. · Develop and maintaincomprehensive dashboards, alerts, and logging using tools like Prometheus,Django, Grafana, Datadog, Splunk etc. · Conduct network capacityplanning, performance tuning, and chaos engineering to validate resilience. · Collaborate on network securityposture, zero-trust models, firewall policies, and compliance requirements. · Work with DevOps, SRE, Cloud,Security, and Application teams to embed reliability into the developmentlifecycle. · Guide junior engineers and contribute to knowledge sharing within the global team. Requirements · Bachelor Degree in Computer Science, Engineering, or related field (or equivalent experience). · Preferred CCIE and/or CCNPcertification Core Networking · Deep expertise in Routing (BGP,OSPF, EIGRP), Switching (STP, VLANs, VXLAN), Firewalls, Load Balancers, VPNs, and SD-WAN. · Strong troubleshooting of complex Layer 2/3/4 issues. · Experience applying SREprinciples (error budgets, toil reduction, automation). · Proficiency in scripting(Python) and Infrastructure as Code (Terraform, Ansible, etc.). · Prometheus, Grafana, Django,Datadog, etc. · CI/CD & Automation -Jenkins, GitOps, Ansible. · Packet analysis: Wireshark,tcpdump. · Excellent problem-solving,communication, and stakeholder management · Ability to work in a global,24x7 on-call rotation. Must Have Skills · Site Reliability engineering(SRE) · Network Design and Engineering · Cisco Application CentricInfrastructure (ACI) · Network Automation · Firewall · Infrastructure as Code (IaC)