Network Engineer
Yerevan, Armenia · Full time
We are seeking a highly skilled Network Engineer to design, implement, and operate high-performance network fabrics supporting large-scale GPU clusters and AI training workloads. This role focuses on ultra-low latency, high-throughput infrastructure built on InfiniBand and high-speed Ethernet (200/400GbE) technologies.
Key Responsibilities
Network Architecture & Implementation
- Design and deploy leaf–spine topologies for GPU clusters (InfiniBand or high-speed Ethernet)
- Configure and manage switches such as Mellanox SN5600/SN4600, QM9700/QM9790, or equivalent
- Build and optimize RoCEv2 / InfiniBand fabrics (PFC, ECN, DCBX, congestion control)
- Implement network segmentation and multi-tenancy (VLANs, VRFs, VPCs, SDN overlays)
- Integrate BlueField-3 DPUs, SR-IOV vNICs, and networking offload solutions
Operations & Reliability
- Monitor fabric health (performance, congestion, drops, routing, credit starvation)
- Troubleshoot high-performance AI training pipelines (collectives, all-reduce, RDMA)
- Maintain routing and security configurations: BGP (EVPN), OSPF, MLAG/VPC
- Manage spine upgrades and firmware lifecycles
- Develop automation for network configuration (Ansible, Terraform, GitOps)
Storage Networking
- Deploy and optimize NVMe-over-Fabrics
- Integrate high-performance storage platforms (WekaFS, VAST, DDN, Pure, Ceph NVMe tier)
- Tune networking for ultra-low latency and high-throughput workloads (AI training & inference)
Data Center Integration
- Collaborate with electrical, mechanical, and compute teams on rack layouts, airflow, and cabling
- Validate rack bring-up: switch burn-ins, link activation, cross-connect testing
Security & Compliance
- Implement secure multi-tenant isolation for GPU cloud environments
- Harden APIs, management networks, and out-of-band access
Automation & Observability
- Build dashboards and alerting using Grafana, Prometheus, InfiniBand Telemetry, WJH
- Contribute to internal tooling for fabric validation, auto-provisioning, and drift detection
Required Skills
Core Networking
- Deep expertise in InfiniBand (HDR/NDR) or high-speed Ethernet (200G/400G)
- Strong RDMA knowledge (RoCEv2, QP states, congestion management)
- Advanced routing and switching (BGP, EVPN, OSPF, MLAG/VPC)
- Experience with Mellanox/NVIDIA platforms (ONIE, Cumulus, MLNX-OS, UFM, NEO)
- Strong Linux networking fundamentals
AI Infrastructure
- Experience supporting AI training clusters (H100, B200, A100 or similar GPUs)
- Familiarity with BlueField DPUs, SR-IOV, VirtIO
- Exposure to OVN/Kube-OVN/Calico networking
Storage Networking
- Understanding of NVMe-oF (TCP or RDMA)
- Experience with WekaFS, DDN, VAST, or Pure (strong advantage)
Nice to Have skills
- Kubernetes networking (CNI, Cilium, Calico, Multus)
- OpenStack Neutron / OVN experience
- Understanding of GPU virtualization and MIG networking patterns