
Deep Learning Performance Software Intern - 2027
Shanghai, Shanghai, ChinaFull-timePosted 3 days agoStill listed 2 days ago
Most applications go out cold — see where you stand first. No sign-up to start.
Watch jobs like this. New roles like this one near Shanghai, Shanghai, by email.
Don't just apply. Show up ready.
Olive works from this exact posting.
At a glance
Olive lists jobs from US employers, including remote roles you can work from the United States.
Job overview
NVIDIA seeks a Deep Learning Performance Software Engineering Intern to join its fast‑paced, customer‑oriented research and development team, building GPU‑accelerated software and optimizing deep learning performance.
Skills & qualifications
Skills
Qualifications
Full job description
We are now looking for a Deep Learning Performance Software Engineering Intern!
We are expanding our research and development for deep learning. We seek excellent Software Engineers to join our team. We specialize in developing GPU-accelerated Deep learning software. Researchers around the world are using NVIDIA GPUs to power a revolution in deep learning, enabling breakthroughs in numerous areas. Join the team that builds software to enable new solutions. Your ability to work in a fast-paced customer-oriented team is required and excellent communication skills are necessary.
What you’ll be doing:
-
Creating and maintaining SKILL, Wiki, and agent harness
-
Develop TileGym, Triton CUDA TileIR backend and CUDA Tile
-
Develop highly optimized deep learning kernels through tile-based GPU programming model
-
End-to-end performance optimization through tile-based GPU programming model
-
Do performance optimization, analysis, and tuning
What we need to see:
-
Pursuing a degree from a university in an engineering or computer science related field. A masters or doctoral candidate is preferred.
-
Understands the core components of agentic systems, including LLM APIs, prompting, tool use/function calling, agent loops, planning, reasoning, memory, RAG, skills, MCP, sub-agents, multi-agent architectures, context engineering, and harness engineering.
-
Excellent C/C++ programming and software design skills
-
Python experience a plus
-
MLIR experience a plus
-
Performance modelling, profiling, debug, and code optimization or architectural knowledge of CPU and GPU
-
GPU programming experience (CUDA or OpenCL) desired
NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most brilliant and talented people on the planet working for us. If you're creative and autonomous, we want to hear from you!
Similar jobs, posted recently
Open roles like this one, listed in the last 30 days.
Performance Software Intern, Deep Learning Libraries - 2027NVIDIA · Shanghai, Shanghai, ChinaPosted 1w agoPosted 1w ago
Senior Performance Software Engineer, Deep Learning LibrariesNVIDIA · Shanghai, Shanghai, ChinaPosted 1w agoPosted 1w ago
Deep Learning Compiler Intern - 2027NVIDIA · Shanghai, Shanghai, ChinaPosted 3 days agoPosted 3 days ago
AI Computing Software Development Intern - 2027NVIDIA · Shanghai, Shanghai, ChinaPosted 1w agoPosted 1w ago
AI Computing Software Development Intern, LLM Inference - 2027NVIDIA · Shanghai, Shanghai, ChinaPosted 1w agoPosted 1w ago
You've read the whole posting — now see how you match it.