
Software Engineer, LLM Inference
Shanghai, Beijing, ChinaFull-timePosted 1w agoStill listed today
Most applications go out cold — see where you stand first. No sign-up to start.
Watch jobs like this. New roles like this one near Shanghai, Beijing, by email.
Don't just apply. Show up ready.
Olive works from this exact posting.
At a glance
Olive lists jobs from US employers, including remote roles you can work from the United States.
Requirements
Credentials this posting asks for.
Job overview
NVIDIA seeks a CPU computing engineer in Shanghai to craft and develop robust inferencing software, analyze and optimize performance, stay abreast of AI research, and collaborate across teams to guide machine‑learning inferencing solutions.
Skills & qualifications
Skills
Qualifications
Full job description
NVIDIA has continuously reinvented itself over two decades. NVIDIA’s invention of the GPU in 1999 sparked the growth of the PC gaming market, redefined modern computer graphics, and revolutionized parallel computing. More recently, GPU deep learning ignited modern AI — the next era of computing — with the GPU acting as the brain of computers, robots, and self-driving cars that can perceive and understand the world.
This is our life’s work — to amplify human imagination and intelligence. AI becomes more and more important in AI-City and self-driving car. NVIDIA is at the forefront of the AI-City and self-driving revolution and providing powerful solutions for them. All these solutions are based on GPU-accelerated libraries, such as CUDA, cuDNN and TensorRT, etc. Now, we are now looking for an CPU computing engineer based in Shanghai.
What you’ll be doing:
-
Craft and develop robust inferencing software that can be scaled to multiple platforms for functionality and performance
-
Performance analysis, optimization and tuning
-
Closely follow academic developments in the field of artificial intelligence and feature update TensorRT and TensorRT Edge LLM
-
Collaborate across the company to guide the direction of machine learning inferencing, working with software, research and product teams
What we need to see:
-
Masters or higher degree in Computer Engineering, Computer Science, Applied Mathematics or related computing focused degree (or equivalent experience)
-
4+ years of relevant software development experience.
-
Excellent C/C++ programming and software design skills, including debugging, performance analysis, and test design.
-
Strong curiosity about artificial intelligence, awareness of the latest developments in deep learning like LLMs, generative models
-
Experience working with deep learning frameworks like PyTorch
-
Proactive and able to work without supervision
-
Excellent written and oral communication skills in English
-
Strong customer communication skills, powerfully motivated to provide highly responsive support as needed
#deeplearning
Similar jobs, posted recently
Open roles like this one, listed in the last 30 days.
AI Computing Software Development Intern, LLM Inference - 2027NVIDIA · Shanghai, Shanghai, ChinaPosted 4 days agoPosted 4 days ago
Senior Performance Software Engineer, Deep Learning LibrariesNVIDIA · Shanghai, Shanghai, ChinaPosted 1w agoPosted 1w ago
Senior Software Engineer, Longitudinal Planning – Autonomous VehiclesNVIDIA · Beijing, Beijing, ChinaPosted 2w agoPosted 2w ago
Senior Software Engineer, Lateral Planning – Autonomous VehiclesNVIDIA · Shanghai, Shanghai, ChinaPosted 2w agoPosted 2w ago
Senior Autonomous Driving Software Engineer, L4 PlanningNVIDIA · Shanghai, Beijing, China (Hybrid)Posted 4w agoPosted 4w ago
You've read the whole posting — now see how you match it.