News

Nvidia shrinks Blackwell for mass robotics

New modules target mainstream edge AI

Shawnee Blackwood

Nvidia introduced the Jetson Thor T3000 and T2000 modules last Wednesday in Tokyo, as part of Jensen Huang’s visit to Japan. The announcement comes as robotics moves out of research labs and into warehouses, factories, and homes, creating real demand for AI compute that runs directly on the machine rather than in a data center.

The Jetson Thor T3000 combines a 1,536-core Nvidia Blackwell GPU with an eight-core Arm Neoverse V3AE CPU, 32GB of LPDDR5X memory running at 273GB/s of bandwidth, and 25 GbE connectivity. It delivers 865 FP4 teraflops of AI compute, about half the size and power draw of the flagship T5000, which delivers 2,070 FP4 teraflops. Quoting from their tests, Nvidia says the T3000 matches the T5000’s inference performance for multimodal workloads: large language models, vision language models, vision language action models, and world foundation models.

Figure 1. Despite its smaller footprint, Nvidia says the inference performance of the T3000 is comparable to the T5000 for multimodal workloads. (Source: Nvidia)

An industrial variant, the IGX T3000, adds integrated functional safety and runs Nvidia’s Halos for Robotics safety stack, built for robots working directly alongside people. The T3000 has one significant difference from the higher tiers: it drops (MIG) GPU partitioning, a feature on the T4000 and T5000 that splits a single GPU into hardware-isolated slices, each running a separate model with strict resource isolation. As a result, robotics pipelines running perception, motion planning, and human-interaction models on the T3000 have to rely on software scheduling, which adds latency. Nvidia has not published full power specs for the T3000, but outside estimates put its thermal design power at roughly half the T5000’s 40-130W range. Those estimates are unconfirmed as of this writing. The 865 FP4 teraflops figure itself is a peak sparse-compute number, and actual throughput on any given model depends on memory bandwidth, precision, and batch size.

The Jetson Thor T2000 drops the compute floor further. Built around a 1,024-core Blackwell GPU, it delivers 400 FP4 teraflops, 16GB of LPDDR5 memory at 137 GB/s, and 2×10 GbE connectivity. Nvidia positions T2000 for visual AI agents, autonomous mobile robots, and industrial manipulators, applications that don’t need T3000’s larger memory pool or GPU parallelism. Together, the two new modules extend the Jetson lineup from 70 TOPS at the low end to the T5000’s 2,000-teraflop ceiling.

Figure 2. Nvidia now offers a scalable edge AI platform spanning performance from 70 TOPS to 2,000 teraflops. (Source: Nvidia)

Alongside the hardware, Nvidia also introduced Cosmos 3 Edge, a 4-billion-parameter addition to its Cosmos 3 world foundation model family designed to run locally on edge hardware such as Jetson Thor. Cosmos 3 Edge uses the same Mixture-of-Transformers architecture as its larger siblings, a Reasoner tower for vision-language understanding paired with a Generator tower for world simulation and action prediction, compressed into a parameter count small enough for on-device inference. The parent Cosmos 3 family launched at GTC Taipei in June 2026 with Super (32 billion parameters) and Nano (8 billion parameter) variants; Edge is the new tier built specifically for Jetson deployment. Nvidia says developers can post-train Cosmos 3 Edge for a specific robot and sensor configuration in about a day, closing what the industry calls the sim-to-real gap: the performance drop that occurs when a policy, the function or algorithm that dictates the robot’s behavior, developed through simulation meets real friction, contact dynamics, and sensor noise. That one-day estimate is Nvidia’s own projection, unvalidated at production scale.

Nvidia also released Jetson agent skills, AI-driven tools that automate memory optimization, system configuration, and deployment tasks that previously took weeks of manual engineering. The Jetson set of skills cover the full Jetson portfolio, including Thor and Orin, enabling teams to run heavier workloads on cheaper, lower-memory hardware without rewriting their software stacks. Nvidia named several customers already using the tools: humanoid robotics firms UBTech and Agile Robots, along with industrial integrator Connect Tech. Nvidia says these customers were able to cut memory usage by up to 15GB and move from the 64GB Jetson AGX Orin to the 32GB configuration. Smart-retail vision company SandStar cut memory use by up to 4GB, enough to shift onto the Orin NX 8GB module instead of the 16GB version. Intelligent-transportation firm NoTraffic trimmed memory requirements by 30% on the Jetson TX2 NX, freeing room for more AI capability in its smart traffic platform. Companion robotics maker GROOVE X, creator of the LOVOT social robot, used Jetson’s mixed AI accelerators to redistribute workloads across memory tiers. (These results come from Nvidia and its partners, not independent audits.)

Companies already building on the Jetson AGX Thor platform include 1X, Agile Robots, Amazon Robotics, Boston Dynamics, Fanuc, Hitachi, and Techman Robot. Hardware partners shipping Thor-based systems include Adlink, Advantech, Aaeon, Aetina, Auvidea, AVerMedia, Connect Tech, ForeCR, JWIPC, Nexcom Robotic Solutions, Realtimes, Seeed Studio, Twowin, Tztek, and Yuan, with Antmicro, Neurealm, Rebotnix, and RidgeRun handling software migration and emulation. Developers can start on the existing Jetson AGX Thor developer kit; T3000 emulation arrives through JetPack 7.2.1 later this month, with T2000 emulation following in a future release. Nvidia expects both modules to ship in the first quarter of 2027.

This announcement did not happen in isolation. Last month, Nvidia signed multi-year agreements with SK Hynix, Naver, SK Telecom, Doosan Group, and LG Group covering memory supply and AI infrastructure to address rising memory prices. Nvidia cites increasing memory prices as a reason to migrate down to Jetson modules. The T3000 and T2000 exist because robotics companies need cheaper, lighter AI hardware right now, and Nvidia needs its own memory supply chain locked down to keep selling it to them.

What do we think?

The T3000 and T2000 read as inventory management as much as innovation. Robotics companies want Blackwell-class AI without data-center power budgets, and Nvidia wants every price tier locked to its own architecture before AMD or a startup accelerator gets a foothold in edge robotics. The missing MIG support and unconfirmed power specs suggest Nvidia shipped ahead of full validation.

Nvidia’s T3000 and T2000 mark an inflection point in edge AI. Their introduction marks the moment full multimodal inference, language, vision, and world models moved from a data-center rack into a module smaller than a paperback book. Cosmos 3 Edge matters as much as the silicon. It is a foundation model developers can adapt to a specific robot in a day, removing robotics’ biggest bottleneck since sensors got cheap. Whether AMD or a Chinese rival closes this gap before Nvidia ships the T3000 remains open.