NRE/Network Architect
Wymagane:PythonAWSAzureCloudDevOpsCI/CDSecurityLeadJuniorHybrid
Job Summary
We are seeking a seasoned Network Reliability Engineer (NRE) with deep expertise in designing, operating, and optimizing hybrid & cloud network environments. You will focus on ensuring high availability, scalability, performance, and resilience of enterprise networks spanning on-premises data centers, SD-WAN, ACI, and major cloud platforms (AWS, Azure). This role combines network engineering with Reliability Engineering principle, emphasizing automation, observability, proactive incident prevention, and rapid recovery to meet stringent SLAs for mission-critical applications.
Responsibilities
- Design and implement highly reliable network infra.
- Define and track Service Level Objectives (SLOs), Service Level Indicators (SLIs) and error budgets.
- Lead major incident response,root cause analysis (RCA), and post-mortem reviews.
- Implement blamelesspost-mortems and drive corrective actions to prevent recurrence.
- Build automation scripts,tools, and self-healing mechanisms for network provisioning, configurationmanagement, monitoring, and failover.
- Develop and maintaincomprehensive dashboards, alerts, and logging using tools like Prometheus,Django, Grafana, Datadog, Splunk etc.
- Conduct network capacityplanning, performance tuning, and chaos engineering to validate resilience.
- Collaborate on network securityposture, zero-trust models, firewall policies, and compliance requirements.
- Work with DevOps, SRE, Cloud,Security, and Application teams to embed reliability into the developmentlifecycle.
- Guide junior engineers and contribute to knowledge sharing within the global team.
Requirements
- Bachelor Degree in Computer Science, Engineering, or related field (or equivalent experience).
- Preferred CCIE and/or CCNPcertification Core Networking
- Deep expertise in Routing (BGP,OSPF, EIGRP), Switching (STP, VLANs, VXLAN), Firewalls, Load Balancers, VPNs, and SD-WAN.
- Strong troubleshooting of complex Layer 2/3/4 issues.
- Experience applying SREprinciples (error budgets, toil reduction, automation).
- Proficiency in scripting(Python) and Infrastructure as Code (Terraform, Ansible, etc.).
- Prometheus, Grafana, Django,Datadog, etc.
- CI/CD & Automation -Jenkins, GitOps, Ansible.
- Packet analysis: Wireshark,tcpdump.
- Excellent problem-solving,communication, and stakeholder management
- Ability to work in a global,24x7 on-call rotation.
Must Have Skills
- Site Reliability engineering(SRE)
- Network Design and Engineering
- Cisco Application CentricInfrastructure (ACI)
- Network Automation
- Firewall
- Infrastructure as Code (IaC)