
Software Engineering Intern, DLFW Comms - 2027
Shanghai, Shanghai, ChinaFull-timePosted 6 days agoStill listed 6 days ago
Most applications go out cold — see where you stand first. No sign-up to start.
Watch jobs like this. New roles like this one near Shanghai, Shanghai, by email.
Don't just apply. Show up ready.
Olive works from this exact posting.
At a glance
Olive lists jobs from US employers, including remote roles you can work from the United States.
Job overview
NVIDIA seeks a motivated Deep Learning engineer intern to integrate advanced communication technologies into AI stacks such as PyTorch, vLLM, SGLang, TRT‑LLM, and veRL, working with communication libraries like NCCL and NVSHMEM to improve multi‑GPU performance for training and inference.
Skills & qualifications
Skills
Qualifications
Full job description
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.
We are looking for a motivated Deep Learning engineer to integrate advanced communication technologies into AI stacks like PyTorch, vLLM, SGLang, TRT-LLM, and veRL. You will be working with the team that developed communication libraries -- such as NCCL and NVSHMEM -- for scaling Deep Learning applications. Your customers will have diverse multi-GPU needs, ranging from training on scales up to 100K GPUs to inference at microsecond latency. Communication performance between GPUs directly affects AI applications. Your work in AI toolkits will simplify these challenges for the community. This is an excellent opportunity for someone with an AI background to push the state of the art in this field. Are you ready to contribute to innovative technologies and help realize NVIDIA's vision?
What you'll be doing:
-
Integrate new communication libraries features in AI frameworks: from PoC to performance analysis to production.
-
Perform deep analysis of AI workloads and frameworks to identify multi-GPU communication requirements and opportunities. Collaborate hands-on with teams working on the latest AI models.
-
Author custom communication or fused compute-communication kernels to showcase ultimate performance on NV platforms.
-
Conduct in-depth research to achieve SOL GPU performance.
-
Build fault-tolerant and elastic solutions for large-scale or dynamic AI workloads.
-
Collaborate with a very dynamic team across multiple time zones.
What we need to see:
-
You are pursuing a M.S. or Ph.D. in CE/CS/EE with a strong background in communication, kernel authoring, and/or AI training/inference.
-
Rapid prototyping and development with Python, C++, CUDA or related DSLs (Triton, cuTe).
-
Solid understanding of LLM models and parallelisms.
-
Adaptability and passion to learn new areas and tools.
-
Flexibility to work and communicate effectively.
Ways to stand out from the crowd:
-
Development experience with frameworks such as PyTorch, JAX, TRT-LLM, vLLM, SGLang, or veRL.
-
Experience with DL communication patterns such as Expert Parallelism (EP), TP, DP & PP.
-
Experience with CUDA kernel optimization and profiling.
-
Experience with large-scale training or production inference stack.
Widely considered to be one of the technology world’s most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As you plan your future, see what we can offer to you and your family www.nvidiabenefits.com/
Similar jobs, posted recently
Open roles like this one, listed in the last 30 days.
Software Engineering Intern, Test Development - 2027NVIDIA · Shanghai, Shanghai, ChinaPosted 3w agoPosted 3w ago
Architecture Energy Modeling Engineering Intern - 2027NVIDIA · Shanghai, Shanghai, ChinaPosted 1w agoPosted 1w ago
AI Developer Technology Engineering Intern - 2027NVIDIA · Beijing, China, ChinaPosted 3w agoPosted 3w ago
System Software Intern, Video Chips - Summer 2027NVIDIA · Shanghai, Shanghai, ChinaPosted 2w agoPosted 2w ago
AI Computing Software Intern, GPU Kernel Libraries - 2027NVIDIA · Shanghai, Shanghai, ChinaPosted 3 days agoPosted 3 days ago
You've read the whole posting — now see how you match it.