
Senior Software Engineer Tech Lead, Network System Validation
Yokneam, North District, IsraelFull-timePosted 1w agoStill listed 1w ago
Most applications go out cold — see where you stand first. No sign-up to start.
Watch jobs like this. New roles like this one near Yokneam, North District, by email.
Don't just apply. Show up ready.
Olive works from this exact posting.
At a glance
Olive lists jobs from US employers, including remote roles you can work from the United States.
Requirements
Credentials this posting asks for.
Job overview
NVIDIA seeks a Technical Lead for Network System Validation to design validation methodologies, develop test automation, debug complex networking solutions, and lead technical alignment across hardware and software teams in large‑scale AI clusters.
Skills & qualifications
Skills
Qualifications
Full job description
NVIDIA is shaping the next era of computing, where AI, accelerated computing, and high-speed networking come together to power the world’s most advanced AI systems. Within NVIDIA, the Networking Business Unit (NBU) builds the high-speed interconnect technologies — Ethernet, InfiniBand, NVLink, and BlueField DPUs — that connect thousands of GPUs into a single AI supercomputer.
NVIDIA is looking for a Technical Lead to join our Network System Validation group and lead the validation of advanced networking solutions across complex AI cluster environments. This is a deeply hands-on technical leadership role, combining ownership of the validation roadmap with technical mentoring and engineering excellence. You will develop validation methodologies and automation frameworks, while working hands-on on debugging, performance analysis, and cutting-edge AI networking technologies at scale. Join us to help push NVIDIA’s networking technologies to their limits and shape how next-generation AI infrastructure is validated.
What you’ll be doing:
- Review system and product requirements, design validation methodologies, develop and implement comprehensive test plans, functional and performance, for networking technologies in large-scale AI cluster solutions
- Develop and maintain benchmarks, automation tools and scripts for test execution, environment setup, log collection, and data analysis.
- Lead end-to-end investigation of complex issues by reproducing real-world scenarios, analyzing logs, telemetry, packet captures, and system metrics to identify functional issues and performance bottlenecks, triaging problems across the hardware and software stack, and driving them to root cause and resolution
- Read and understand source code (C/C++/Python) to investigate defects, validate fixes, and improve logging, instrumentation, and debugging capabilities
- Collaborate deeply with software and hardware development teams to debug networking technologies, including NCCL, RoCE, RDMA, and related software components using targeted experiments and code inspection
- Profile and research AI training and inference workloads, correlating application behavior with network and system telemetry to identify scalability and performance limitations
- Document findings, communicate technical results, and continuously improve validation methodologies, automation environments, and engineering processes
What we need to see:
- B.Sc. / B.A. in Computer Science, Electrical Engineering, or equivalent experience
- 12+ years of experience in networking, system validation, or related domains
- Proven experience debugging complex production systems by forming hypotheses, designing experiments, and driving issues to root cause
- Ability to read, debug, and reason about C/C++ code (Rust or Go a plus)
- Strong scripting and automation experience using Python, Bash, and/or Ansible
- Deep understanding of distributed systems: concurrency, consistency models, fault tolerance, and large-scale system performance under stress
- Ability to drive technical alignment across teams, communicate tradeoffs clearly, and make high-quality architectural decisions at speed
- Advance AI-driven approaches to test automation: intelligent scenario generation, LLM-augmented root-cause analysis, and autonomous validation pipelines
Ways to stand out from the crowd:
- Experience with large-scale clusters or distributed systems
- Familiarity with NVIDIA networking solutions (ConnectX, SpecX, BlueField)
- Background in performance analysis, Kubernetes, or cloud environments
- Background in chaos testing, fault injection, or simulation systems
We have some of the most forward-thinking and hardworking people working for us. If you're creative and autonomous, we want to hear from you! NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, disability status or any other characteristic protected by law.
Similar jobs, posted recently
Open roles like this one, listed in the last 30 days.
Senior Software Engineer, Network System ValidationNVIDIA · Yokneam, Israel, IsraelPosted 4w agoPosted 4w ago
Software Architect, Network System ValidationNVIDIA · Yokneam, North District, IsraelPosted 4w agoPosted 4w ago
Engineering Manager, Network System ValidationNVIDIA · Yokneam, North District, IsraelPosted 2w agoPosted 2w ago
Senior System Validation EngineerNVIDIA · Yokneam, North District, IsraelPosted 2 days agoPosted 2 days ago
Wireless Connectivity System and Architecture EngineerIntel · Petah-Tikva, Center District, Israel (Hybrid)Posted 3w agoPosted 3w ago
You've read the whole posting — now see how you match it.