Mini SOM Brings LLMs to Robots and Drones
The latest mini Jetson computer module from NVIDIA provides enough computing power to run AI models with as many as tens of billions of parameters on drones, robots, and other physical AI systems.
The Jetson Orin Nano 2 ushers in improvements across the board, bringing 2X the performance of the previous Jetson Orin Nano Super in the same form factor and using as much as 40% less power to hit the same performance. It sits several levels below NVIDIA’s Jetson Thor computer, which brings more than 2,000 TOPS to areas such as autonomous driving. But the Nano is aimed at a different class of system: robots and other edge devices where compute, power, and physical space are all tightly constrained.
At the heart of the module is a NVIDIA Ampere-based GPU paired with eight Arm Cortex-A78 CPU cores and 8 GB of LPDDR5X memory with up to 120 GB/s of memory bandwidth. Deepu Talla, VP of robotics and edge AI at NVIDIA, said it can deliver up to 78 TOPS, significantly boosting performance for image, video, and AI processing while staying in the small, power-efficient form factor of a system-on-module (SOM).
Talla added, "We can put frontier AI models such as LLMs and VLMs on top of the other autonomous capabilities. The level of accuracy in these models — we’re now able to bring it all the way down to entry-level edge AI products.”
New Jetson Nano Brings Frontier AI Models to the Edge
The new Nano stands out for its ability to run increasingly advanced large language models (LLMs) and vision language models (VLMs) within the confines of a compact robotics computer, said Talla. These so-called frontier models are on the leading edge of AI development. Trained with huge amounts of data in hyperscale data centers, they’re usually very general-purpose models that can serve as the foundation for more specific models working with text, audio, and/or video.
The largest frontier models can contain hundreds of billions or even more than a trillion parameters each. Their computational demands make them far too memory- and power-intensive to run directly on small, energy-efficient computers at the edge.
However, times are changing with a new class of small language models (SLMs) that can run locally on devices and deliver inference in real-time, said Talla. These models squeeze the performance of advanced LLMs and VLMs down to hundreds of millions to tens of billions of parameters, allowing them to fit into significantly smaller form factors (Fig. 1).
Talla noted that “today’s small and medium frontier models have reached the accuracy of last year’s largest frontier models, unlocking real-time intelligence for edge devices."
According to NVIDIA, the new Jetson packs the performance and energy efficiency required to run these frontier-class models in real-time, physical AI systems such as drones and robots. The module can run LLMs and VLMs optimized for memory-efficient edge inference, including open models such as NVIDIA’s Cosmos and Nemotron as well as Gemma 4 and Qwen 3. Talla said that these compact frontier models usually have in the range of 5 to 30 billion parameters.
The Orin Nano 2 doubles the performance of the Jetson Orin Nano Super in terms of tokens per second, which is the metric used to measure how fast one of these models delivers its output during inference. The performance gains come primarily from improved tensor cores in the GPU and higher bandwidth in the LPDDR5X. While raw compute capacity is critical, memory bandwidth is what usually limits the rate of tokens per second because models take too long to move data between memory and GPU.
Talla said a large data-center GPU such as NVIDIA’s H100 may be able to produce 30 to 40 tokens per second when running a frontier-class model such as Nemotron. In contrast, he said, the Orin Nano 2 can deliver more than 20 tokens per second when running the much more compact Nemotron Nano model, which is still enough to enable real-time interaction (Fig. 2). He added, “What we are seeing is that what used to take a large data center with multiple NVIDIA GPUs, we can now run that in real time at the edge on a Jetson.”
The new Orin Nano comes with configurable power consumption ranging from 15 to 40 W. In its 15-W mode, NVIDIA said it consumes 40% less power to deliver the same performance as its predecessor, which requires 25 W of power to match it.
SOM Expands Real-Time AI for Drones and Robots
The new SOM is powered by NVIDIA’s open software stack and the broader ecosystem around it. Talla said the platform will primarily be used in robots and drones so that they can better understand their surroundings, interpret context, and react in real-time as conditions change.
Wing, the drone-delivery subsidiary of Alphabet, already uses the Jetson Orin Nano Super and NVIDIA’s software stack across its drone fleet (Fig. 3). The company plans to evaluate the Orin Nano 2 to improve real-time AI perception and reasoning in its heavy-duty, autonomous drones, potentially enabling faster and safer deliveries from local businesses to residences.
“Drone delivery depends on AI that can enable fast, reliable understanding of the real world,” said Dinuka Abeywardena, Wing's head of perception.
Matic Robots is also adopting Jetson Orin Nano 2 in its home-cleaning robots. The additional computing power will help them map out their surroundings more accurately and better understand the arrangement of objects and spaces. Thus, they’re able to navigate around obstacles and clean in changing environments. The robots will also be able to use state-of-the-art AI models to interact with users and get deeper context into what is happening around them, said Navneet Dalal, CEO of Matic.
The new Orin Nano can be used with NVIDIA's Omniverse, which provides a photorealistic simulation environment to train and test robots, while the Jetson serves as the compact computer to be mounted inside the physical robot to run those trained models in the real world.
About the Author
James Morra
Senior Editor
James Morra is the senior editor for Electronic Design, covering the semiconductor industry and new technology trends, with a focus on power electronics and power management. He also reports on the business behind electrical engineering, including the electronics supply chain. He joined Electronic Design in 2015 and is based in Chicago, Illinois.





