Open Source

NVIDIA's Cosmos 3 Edge brings 4B-parameter world model to edge devices

New open model runs real-time robot control at 15 Hz on Jetson Thor

Deep Dive

NVIDIA has launched Cosmos 3 Edge, a compact 4-billion-parameter open world model designed to bring data center-level reasoning to edge devices. The model is available on Hugging Face and targets physical AI systems in factories, warehouses, and hospitals. It runs across NVIDIA's edge computing portfolio, including RTX PRO GPUs, DGX, GeForce RTX GPUs, and the newly announced Jetson T2000 and T3000 modules. On Jetson Thor, Cosmos 3 Edge achieves real-time control at 15 Hz, generating 32 actions per inference with robot-control resolution (640×360 observations). It currently ranks #1 on VANTAGE-Bench for vision analytics among similar-size models and sets a new state-of-the-art for robot policy learning, making it ideal for smart infrastructure and robotics.

Cosmos 3 Edge combines two transformer towers in a novel architecture. The autoregressive tower processes vision and text tokens for understanding and reasoning, while the diffusion tower handles vision, audio, and action tokens for prediction and generation. Both towers share multimodal attention layers that align information across language, video, audio, and action, enabling the model to reason about a scene before generating outputs. Crucially, Cosmos 3 maps different embodiments—robot arms, vehicles, cameras—into a common action representation using compact geometric vectors for translation, rotation, and manipulation state. This allows the model to associate visual changes with physical motion, enabling real-time simulation and action generation on memory-constrained edge hardware without compromising performance.

Key Points
  • 4-billion-parameter open world model for edge inference on NVIDIA RTX PRO, DGX, GeForce RTX, and Jetson modules
  • Delivers real-time robot control at 15 Hz, generating 32 actions per inference on Jetson Thor
  • Combines autoregressive and diffusion transformer towers with shared multimodal attention for reasoning and action generation

Why It Matters

Enables real-time physical AI at the edge, bringing data center-level reasoning to robots and vision agents in factories and hospitals.

📬 Get the top 10 AI stories daily