Network Engineer

Yerevan, Armenia · Full time

We are seeking a highly skilled Network Engineer to design, implement, and operate high-performance network fabrics supporting large-scale GPU clusters and AI training workloads. This role focuses on ultra-low latency, high-throughput infrastructure built on InfiniBand and high-speed Ethernet (200/400GbE) technologies.

Key Responsibilities
Network Architecture & Implementation
  • Design and deploy leaf–spine topologies for GPU clusters (InfiniBand or high-speed Ethernet)
  • Configure and manage switches such as Mellanox SN5600/SN4600, QM9700/QM9790, or equivalent
  • Build and optimize RoCEv2 / InfiniBand fabrics (PFC, ECN, DCBX, congestion control)
  • Implement network segmentation and multi-tenancy (VLANs, VRFs, VPCs, SDN overlays)
  • Integrate BlueField-3 DPUs, SR-IOV vNICs, and networking offload solutions
Operations & Reliability
  • Monitor fabric health (performance, congestion, drops, routing, credit starvation)
  • Troubleshoot high-performance AI training pipelines (collectives, all-reduce, RDMA)
  • Maintain routing and security configurations: BGP (EVPN), OSPF, MLAG/VPC
  • Manage spine upgrades and firmware lifecycles
  • Develop automation for network configuration (Ansible, Terraform, GitOps)
Storage Networking
  • Deploy and optimize NVMe-over-Fabrics
  • Integrate high-performance storage platforms (WekaFS, VAST, DDN, Pure, Ceph NVMe tier)
  • Tune networking for ultra-low latency and high-throughput workloads (AI training & inference)
Data Center Integration
  • Collaborate with electrical, mechanical, and compute teams on rack layouts, airflow, and cabling
  • Validate rack bring-up: switch burn-ins, link activation, cross-connect testing
Security & Compliance
  • Implement secure multi-tenant isolation for GPU cloud environments
  • Harden APIs, management networks, and out-of-band access
Automation & Observability
  • Build dashboards and alerting using Grafana, Prometheus, InfiniBand Telemetry, WJH
  • Contribute to internal tooling for fabric validation, auto-provisioning, and drift detection
Required Skills
Core Networking
  • Deep expertise in InfiniBand (HDR/NDR) or high-speed Ethernet (200G/400G)
  • Strong RDMA knowledge (RoCEv2, QP states, congestion management)
  • Advanced routing and switching (BGP, EVPN, OSPF, MLAG/VPC)
  • Experience with Mellanox/NVIDIA platforms (ONIE, Cumulus, MLNX-OS, UFM, NEO)
  • Strong Linux networking fundamentals
AI Infrastructure
  • Experience supporting AI training clusters (H100, B200, A100 or similar GPUs)
  • Familiarity with BlueField DPUs, SR-IOV, VirtIO
  • Exposure to OVN/Kube-OVN/Calico networking
Storage Networking
  • Understanding of NVMe-oF (TCP or RDMA)
  • Experience with WekaFS, DDN, VAST, or Pure (strong advantage)
Nice to Have skills
  • Kubernetes networking (CNI, Cilium, Calico, Multus)
  • OpenStack Neutron / OVN experience
  • Understanding of GPU virtualization and MIG networking patterns