Embodied Foundation Models
-
OLA Dimensions UniWAM
Hugging Face: https://huggingface.co/chenpyyy/UniWAM-base, https://huggingface.co/collections/chenpyyy/uniwam
GitHub: https://github.com/UniWAM/UniWAM
Website: https://uniwam.github.io
Technical report: https://arxiv.org/abs/2610.02054Developer: OLA Dimensions, HKUST(GZ)
Released: October 2026
Variants: UniWAM-base / UniWAM-robotwin-clean / UniWAM-libero / UniWAM-libero-plus
Parameters: 8B
Architecture: Mixture-of-Transformers, three experts: a Qwen3-VL-2B-Instruct physical reasoner, a Wan2.2-TI2V-5B world generator, a flow-matching action predictor; joint multimodal attention
License: Apache-2.0
Modalities: Camera images + proprioceptive state + language instruction in / physical language + future video + action chunk out
Runs on: NVIDIA datacenter GPUs, post-training published on 8 x H100; no inference requirement published
Formats: PyTorch (DeepSpeed checkpoint)
On disk: UniWAM-base 16.05GB -
InternRobotics InternW0-Delta
Hugging Face: https://huggingface.co/InternRobotics/InternW0-Delta-Base, https://huggingface.co/InternRobotics/InternW0-Delta-Libero, https://huggingface.co/InternRobotics/InternW0-Delta-RoboTwin, https://huggingface.co/InternRobotics/InternW0-Delta-RoboDojo
GitHub: https://github.com/InternRobotics/InternW0-Delta
Website: https://internrobotics.github.io/InternW0-Delta/
Technical report: https://arxiv.org/abs/2609.31394Developer: Shanghai AI Laboratory (InternRobotics)
Released: September 2026
Variants: Base / Libero / RoboTwin / RoboDojo
Architecture: World-action model: a pretrained Wan2.2-TI2V-5B video expert and an ActionDiT action expert coupled through 30 directed Mixture-of-Transformers blocks, with a frozen RynnBrain1.1-2B VLM conditioning the action expert and training-only 4D distillation
License: MIT
Modalities: Multi-view RGB + language instruction + proprioceptive state in / action chunk out (canonical 80-D action space)
Runs on: NVIDIA GPU (Linux, CUDA); no inference requirement published
Formats: PyTorch (.pt)
On disk: 12.42GB per checkpoint (Base pretrain.pt) -
Alibaba DAMO RynnValue
Hugging Face: https://huggingface.co/Alibaba-DAMO-Academy/RynnValue-8B-Quantile, https://huggingface.co/Alibaba-DAMO-Academy/RynnValue-4B-Quantile, https://huggingface.co/Alibaba-DAMO-Academy/RynnValue-8B, https://huggingface.co/Alibaba-DAMO-Academy/RynnValue-4B
GitHub: https://github.com/alibaba-damo-academy/RynnValue
Website: https://alibaba-damo-academy.github.io/RynnValue.github.io/
Technical report: https://arxiv.org/abs/2608.09853Developer: Alibaba DAMO Academy
Released: August 2026
Variants: RynnValue-4B, RynnValue-8B, RynnValue-4B-Quantile, RynnValue-8B-Quantile
Parameters: 5.14B (4B checkpoints), 9.57B (8B checkpoints)
Architecture: RynnBrain backbone on the Qwen3-VL architecture; absolute and relative distributional value heads (fixed-bin or 256-bin quantile) plus a language head; predicts remaining time to task completion per frame
License: Apache License 2.0
Modalities: Image/video + language in / per-frame remaining-time value in seconds, video analysis text out
Runs on: Desktop GPU, single GPU with 24GB+ memory for the 8B model in bf16
Formats: safetensors
On disk: 10.29GB (4B checkpoints), 19.15GB (8B checkpoints) -
Alibaba DAMO RynnWorld-Latent
Hugging Face: https://huggingface.co/Alibaba-DAMO-Academy/RynnWorld-Latent
GitHub: https://github.com/alibaba-damo-academy/RynnWorld-Latent
Website: https://alibaba-damo-academy.github.io/RynnWorld-Latent.github.io/Developer: Alibaba DAMO Academy
Released: October 2026
Parameters: 3.46B
Architecture: Full fine-tune of NVIDIA Cosmos3-Edge (Mixture-of-Transformers, Nemotron-2B backbone, SigLIP2 vision tower) into a latent-action-conditioned video world model; rectified-flow video generation on a frozen Wan2.2 VAE; 608-dimension RynnLAM latent actions shared by human and robot video
Modalities: First-frame image + Latent action sequence in / Video out
Formats: safetensors
On disk: 20.75GB safetensors (4 shards, BF16 weights plus F32 EMA copy) -
Xiaomi Robotics Xiaomi-Robotics-U0
Hugging Face: https://huggingface.co/XiaomiRobotics/Xiaomi-Robotics-U0, https://huggingface.co/XiaomiRobotics/Xiaomi-Robotics-U0-FlashAR, https://huggingface.co/XiaomiRobotics/Xiaomi-Robotics-U0-4B, https://huggingface.co/XiaomiRobotics/Xiaomi-Robotics-U0-Sequence, https://huggingface.co/XiaomiRobotics/Xiaomi-Robotics-U0-4B-Sequence
GitHub: https://github.com/XiaomiRobotics/Xiaomi-Robotics-U0
Website: http://robotics.xiaomi.com/xiaomi-robotics-u0.html
Technical report: https://arxiv.org/abs/2607.11643Developer: Xiaomi Robotics
Released: July 2026
Variants: U0, U0-FlashAR, U0-4B, U0-Sequence, U0-4B-Sequence
Parameters: 34B (U0, U0-Sequence), 4B (U0-4B, U0-4B-Sequence)
Architecture: Autoregressive transformer initialized from EMU3.5; shared discrete visual tokenizer, single next-token objective across text and image sequences; FlashAR decodes visual tokens in anti-diagonal groups; Sequence checkpoints interleave subtask text and observations
License: Apache-2.0
Modalities: Text + Image + Video
Runs on: U0-4B variants: Desktop GPU / U0, U0-Sequence, U0-FlashAR: Datacenter GPU, vendor benchmark on one NVIDIA H20
Formats: safetensors
On disk: 68.21GB (U0, U0-Sequence), 75.02GB (U0-FlashAR), 10.16GB (U0-4B, U0-4B-Sequence) -
TeleAI SMART-VLA
Hugging Face: https://huggingface.co/TeleEmbodied/SMART-VLA
GitHub: https://github.com/TeleHuman/PRTS
Website: https://teamillusion-smart.github.io/
Technical report: https://arxiv.org/abs/2610.07652Developer: TeleAI (China Telecom)
Released: October 2026
Parameters: 4.44B
Architecture: PRTS-architecture VLA on a Qwen3-VL-4B-Instruct backbone with a flow-matching DiT action head, pretrained on synthetic articulated-object manipulation data
License: MIT
Modalities: Camera images + language instruction in / 50-step action chunk (up to 32 action dimensions) out
Formats: safetensors (BF16, 2 shards)
On disk: 9.67GB (two safetensors shards) -
Microsoft Rho
Hugging Face: https://huggingface.co/microsoft/rho-base, https://huggingface.co/collections/microsoft/rho
GitHub: https://github.com/microsoft/rhobotics
Website: https://microsoft.github.io/rhobotics/
Technical report: https://arxiv.org/abs/2609.38164Developer: Microsoft Research
Released: September 2026
Variants: rho-base / rho-yam-box / rho-ur-ai-trainer / rho-fr3-duo / rho-libero / rho-roboeval
Parameters: 5B (Phi-Phy VLM 4.68B / flow-matching action expert 542M)
Architecture: Vision-language-action model: a physically grounded Phi-family VLM (Phi-Phy) with a 12-block flow-matching action expert that cross-attends to a VLM decoder layer, midtrained per embodiment
License: MIT
Modalities: Multi-camera images + language instruction + robot state in / action chunk out (up to 50 steps in pretraining)
Runs on: NVIDIA GPU (Linux, CUDA; FlashAttention 2 by default); no VRAM requirement published
Formats: safetensors (BF16, 3 shards)
On disk: rho-base 10.50GB -
Meta FAIR RoboJEPA
GitHub: https://github.com/facebookresearch/robo_jepa
Website: https://robojepa.github.io
Technical report: https://arxiv.org/abs/2610.10515Developer: Meta FAIR
Released: October 2026
Variants: 22M / 50M / 100M / 300M / 1B / 2B / 4B / 8B / 8B DROID 720p (3 views)
Parameters: 22M to 8B predictors on a frozen V-JEPA 2.1 ViT-G/384 encoder
Architecture: Action-conditioned JEPA latent world model: a transformer predictor over frozen V-JEPA 2.1 features, with latent CEM/MPC planning toward a goal image and an optional diffusion video decoder
License: CC BY-NC-SA 4.0
Modalities: Camera video + action sequence in / predicted future latent states out (optional decoded video); no action head, actions come from latent MPC planning
Runs on: NVIDIA GPU (CUDA 12.6), one GPU per model size for evaluation; no VRAM requirement published
Formats: PyTorch (.pth.tar) from dl.fbaipublicfiles.com
On disk: 8B 29.99GB, 1B 4.07GB, 300M 1.24GB (300k-step checkpoints), plus V-JEPA 2.1 ViT-G/384 encoder 30.24GB -
DeepCybo PhysBrain 1.5
Hugging Face: https://huggingface.co/DeepCybo/PhysBrain1.5-8B, https://huggingface.co/DeepCybo/PhysBrain1.5-2B
GitHub: https://github.com/DeepCybo-PhysAI/PhysBrain-1.5
Website: https://deepcybo-physai.github.io/PhysBrain-1.5/
Technical report: https://arxiv.org/abs/2609.14973Developer: DeepCybo, Zhongguancun Academy, Zhongguancun Institute of Artificial Intelligence
Released: September 2026
Variants: 2B, 8B
Parameters: 2.16B (2B checkpoint), 8.90B (8B checkpoint)
Context: 262,144 tokens
Architecture: Qwen3-VL backbone extended with action and visual-state tokens; one shared autoregressive transformer, single next-token objective, no task-specific heads; ActionPiece action tokens with one action codebook across robot setups
Modalities: Text + Image + Video + Action history in / Text + Spatial grounding + Action chunk + Future-state image, depth map, robot mask out
Runs on: 2B: Laptop, Desktop GPU / 8B: Desktop GPU
Formats: safetensors
On disk: 4.32GB (2B checkpoint), 17.81GB (8B checkpoint) -
Metacognition NavGPT3
Hugging Face: https://huggingface.co/Metacognition-AI/NavGPT3-8B, https://huggingface.co/Metacognition-AI/NavGPT3-4B
GitHub: https://github.com/metacognitionai/NavGPT-3
Website: https://metacognitionai.github.io/NavGPT3/
Technical report: https://arxiv.org/abs/2610.10787Developer: Metacognition
Released: October 2026
Variants: 4B, 8B
Parameters: 4.44B (4B checkpoint), 8.77B (8B checkpoint)
Architecture: Qwen3-VL fine-tune with a two-layer MLP action head on the last prompt token’s hidden state; one 3,072-token visual budget shared across a four-view, up to 16-step image history
License: GNU Affero General Public License v3.0
Modalities: Text + Four-view RGB image history in / 8 waypoints (x, y, theta) out
Runs on: NVIDIA GPU (Linux)
Formats: safetensors
On disk: 9.66GB (4B checkpoint), 17.54GB (8B checkpoint)