Skip to content
  • Categories
  • Recent
  • Tags
  • Popular
  • Users
Menu
  1. Home
  2. AI & Software
  3. Embodied Foundation Models

Embodied Foundation Models

Scheduled Pinned Locked Moved AI & Software
embodiedfoundation modelsrobotics
19 Posts 1 Posters 270 Views 1 Watching
  • Oldest to Newest
  • Newest to Oldest
  • Most Votes
Reply
  • Reply as topic
Log in to reply
This topic has been deleted. Only users with topic management privileges can see it.
  • montezM Offline
    montezM Offline
    montez
    wrote on last edited by montez
    #1

    Megathread for embodied foundation models for perception, spatial reasoning, and robot action, from device-native models to large world-action models.

    1 Reply Last reply
    0
    • montezM Offline
      montezM Offline
      montez
      wrote on last edited by
      #2

      NVIDIA Cosmos 3 Edge

      Hugging Face: https://huggingface.co/nvidia/Cosmos3-Edge, https://huggingface.co/collections/nvidia/cosmos3
      GitHub: https://github.com/nvidia/cosmos
      Website: https://research.nvidia.com/labs/cosmos-lab/cosmos3/
      White paper: https://research.nvidia.com/labs/cosmos-lab/cosmos3/technical-report.pdf

      Developer: NVIDIA
      Released: July 2026
      Variants: Cosmos3-Edge, Cosmos3-Edge-Policy-DROID
      Parameters: 4B
      Context: Reasoner: 256,000 tokens / Generator text input: 4,096 tokens
      Architecture: Mixture-of-Transformers, two towers: an autoregressive transformer for text, a diffusion transformer for image/video/action generation
      License: OpenMDW 1.1
      Modalities: Text + Image + Video + Action trajectory in / Text + Image + Video + Action out
      Formats: safetensors via Hugging Face, PyTorch (NVIDIA-proprietary inference only, BF16)
      Runs on: NVIDIA consumer GPU (Linux), Jetson (Linux)

      1 Reply Last reply
      0
      • montezM Offline
        montezM Offline
        montez
        wrote on last edited by
        #3

        Tencent Hy-Embodied-RxBrain-1.0

        Hugging Face: https://huggingface.co/tencent/Hy-Embodied-RxBrain-1.0
        GitHub: https://github.com/Tencent-Hunyuan/Hy-Embodied-RxBrain-1.0
        Website: https://tairos.tencent.com/openSourceModels/hy-embodied-rxbrain-1.0
        Technical report: https://arxiv.org/abs/2607.14187

        Developer: Tencent Robotics X, Futian Laboratory, Tencent Hy Team
        Released: July 2026
        Parameters: ~6.2B
        Architecture: Unified Mixture-of-Transformers, modality-specific text, vision, and generation pathways
        License: Apache 2.0
        Modalities: Text + Image + Video
        Runs on: NVIDIA GPU, CUDA 12.x, Linux recommended
        Formats: safetensors
        On disk: 12.42GB safetensors

        1 Reply Last reply
        0
        • montezM Offline
          montezM Offline
          montez
          wrote on last edited by montez
          #4

          NVIDIA Isaac GR00T N1.7

          Hugging Face: https://huggingface.co/nvidia/GR00T-N1.7-3B, https://huggingface.co/collections/nvidia/gr00t-n17
          GitHub: https://github.com/NVIDIA/Isaac-GR00T
          Website: https://developer.nvidia.com/isaac/gr00t
          Docs: https://huggingface.co/blog/nvidia/gr00t-n1-7

          Developer: NVIDIA
          Released: April 2026
          Variants: N1, N1.5, N1.6, N1.7
          Parameters: 3B
          Architecture: Dual-system VLA pairing a Cosmos-Reason2-2B vision-language backbone for task/subtask reasoning with a flow-matching diffusion transformer conditioned on the reasoning output
          License: NVIDIA Open Model License Agreement
          Modalities: RGB camera frames + language instruction + robot proprioception in / continuous robot action vectors out
          Runs on: Desktop GPU, Datacenter GPU, or NVIDIA Jetson edge modules; NVIDIA only
          Formats: safetensors, 2 shards
          On disk: 6.91GB

          1 Reply Last reply
          0
          • montezM Offline
            montezM Offline
            montez
            wrote on last edited by montez
            #5

            NVIDIA Cosmos Reason 2

            Hugging Face: https://huggingface.co/nvidia/Cosmos-Reason2-8B, https://huggingface.co/nvidia/Cosmos-Reason2-2B, https://huggingface.co/nvidia/Cosmos-Reason2-32B
            Website: https://build.nvidia.com/nvidia/cosmos-reason2-8b
            Docs: https://huggingface.co/blog/nvidia/nvidia-cosmos-reason-2-brings-advanced-reasoning

            Developer: NVIDIA
            Released: January 2026
            Variants: 2B, 8B, 32B
            Parameters: 8.77B (8B tier)
            Architecture: Qwen3-VL fine-tune post-trained for physical-AI reasoning
            License: NVIDIA Open Model License Agreement
            Modalities: Image/video + text in / spatio-temporal reasoning, trajectory, point, and bounding-box predictions out
            Runs on: NVIDIA GPU
            Formats: safetensors, 4 shards (8B tier)
            On disk: 17.53GB (8B tier)

            1 Reply Last reply
            0
            • montezM Offline
              montezM Offline
              montez
              wrote on last edited by montez
              #6

              Unitree UnifoLM-VLA-0

              Hugging Face: https://huggingface.co/unitreerobotics/Unifolm-VLM-Base, https://huggingface.co/unitreerobotics/Unifolm-VLA-Base, https://huggingface.co/unitreerobotics/Unifolm-VLA-Libero
              GitHub: https://github.com/unitreerobotics/unifolm-vla
              Website: https://unigen-x.github.io/unifolm-vla.github.io

              Developer: Unitree Robotics
              Released: January 2026
              Variants: VLM-Base, VLA-Base, VLA-LIBERO
              Architecture: Qwen2.5-VL-7B backbone with a diffusion-transformer flow-matching action head; continued pretraining fuses 2D/3D spatial understanding with action-chunking prediction and forward/inverse dynamics constraints
              Modalities: Multi-view RGB + robot state + language instruction in / continuous manipulation actions out
              Formats: PyTorch checkpoint
              On disk: VLA-Base 18.98GB

              1 Reply Last reply
              0
              • montezM Offline
                montezM Offline
                montez
                wrote on last edited by montez
                #7

                Alibaba DAMO RynnBrain 1.1

                Hugging Face: https://huggingface.co/Alibaba-DAMO-Academy/RynnBrain1.1-2B, https://huggingface.co/Alibaba-DAMO-Academy/RynnBrain1.1-9B, https://huggingface.co/Alibaba-DAMO-Academy/RynnBrain1.1-122B-A10B
                GitHub: https://github.com/alibaba-damo-academy/RynnBrain
                Technical report: https://arxiv.org/abs/2602.14979

                Developer: Alibaba DAMO Academy
                Released: July 2026
                Variants: 2B, 9B, 122B-A10B
                Parameters: 9.41B (9B checkpoint)
                Architecture: Decoder-only vision-language transformer (dense 2B/9B, sparse-MoE 122B-A10B) on a Qwen3.5 base
                License: Apache License 2.0
                Modalities: Image/video + language in / spatial/3D grounding, contact-point and affordance predictions, task planning out
                Formats: safetensors
                On disk: 18.82GB (9B checkpoint)

                1 Reply Last reply
                0
                • montezM Offline
                  montezM Offline
                  montez
                  wrote last edited by
                  #8

                  Black Forest Labs FLUX 3 Action

                  Hugging Face: https://huggingface.co/black-forest-labs/flux-3-action-base, https://huggingface.co/black-forest-labs/flux-3-action-droid, https://huggingface.co/black-forest-labs/flux-3-action-so101
                  GitHub: https://github.com/black-forest-labs/flux-action
                  Website: https://bfl.ai/models/flux-3-action
                  Docs: https://docs.bfl.ai/flux_3/flux3_action_overview

                  Developer: Black Forest Labs
                  Released: September 2026
                  Variants: Base / DROID / SO-101
                  Parameters: 7B
                  Architecture: Diffusion-transformer world action model derived from the multimodal FLUX 3 backbone, with joint future-video and action prediction
                  License: FLUX Kommunity License v1.0
                  Modalities: Text + Multi-view RGB + Robot state in / Video + Action out
                  Runs on: NVIDIA GPU (Linux), about 32GB VRAM for BF16 or 24GB with FP8r and text-encoder offload
                  Formats: safetensors, LeRobot and FLUX Action inference

                  1 Reply Last reply
                  0
                  • montezM Offline
                    montezM Offline
                    montez
                    wrote last edited by
                    #9

                    OpenWAM-alpha

                    Hugging Face: https://huggingface.co/OpenWAM/OpenWAM-Alpha-Pretrain-Foundation-Model, https://huggingface.co/collections/OpenWAM/openwam-alpha
                    GitHub: https://github.com/OpenWAM-Official/OpenWAM
                    Website: https://openwam-official.github.io/
                    Technical report: https://arxiv.org/abs/2609.07398

                    Developer: OpenWAM-Official
                    Released: September 2026
                    Variants: Pretrain-Foundation-Model / Sim fine-tunes (LIBERO, RoboTwin, RoboCasa365, RoboCasa-GR1, RoboDojo, EBench, VLABench) / Real fine-tunes (Dexterous-Hand-Wuji, RoboDojo-ARX-X5, RoboDojo-Piper, RoboDojo-Piper-X, Single-Arm-Franka)
                    Parameters: Video DiT 5B / ActionDiT 1B
                    Architecture: Dual-system world-action model: a pretrained Wan2.2-TI2V-5B video DiT and a dedicated ActionDiT coupled through joint self-attention with a mutual attention mask, on a frozen Wan2.2-VAE
                    License: Apache-2.0
                    Modalities: Multi-view RGB + language instruction + proprioceptive state in / future video + action chunk out (80-D unified action space)
                    Runs on: NVIDIA datacenter GPUs, 8 x 80 GB recommended for training (CUDA 12.8); no inference requirement published
                    Formats: safetensors
                    On disk: Pretrain-Foundation-Model 24.81GB

                    1 Reply Last reply
                    0
                    • montezM Offline
                      montezM Offline
                      montez
                      wrote last edited by
                      #10

                      OLA Dimensions UniWAM

                      Hugging Face: https://huggingface.co/chenpyyy/UniWAM-base, https://huggingface.co/collections/chenpyyy/uniwam
                      GitHub: https://github.com/UniWAM/UniWAM
                      Website: https://uniwam.github.io
                      Technical report: https://arxiv.org/abs/2610.02054

                      Developer: OLA Dimensions, HKUST(GZ)
                      Released: October 2026
                      Variants: UniWAM-base / UniWAM-robotwin-clean / UniWAM-libero / UniWAM-libero-plus
                      Parameters: 8B
                      Architecture: Mixture-of-Transformers, three experts: a Qwen3-VL-2B-Instruct physical reasoner, a Wan2.2-TI2V-5B world generator, a flow-matching action predictor; joint multimodal attention
                      License: Apache-2.0
                      Modalities: Camera images + proprioceptive state + language instruction in / physical language + future video + action chunk out
                      Runs on: NVIDIA datacenter GPUs, post-training published on 8 x H100; no inference requirement published
                      Formats: PyTorch (DeepSpeed checkpoint)
                      On disk: UniWAM-base 16.05GB

                      1 Reply Last reply
                      0
                      • montezM Offline
                        montezM Offline
                        montez
                        wrote last edited by
                        #11

                        InternRobotics InternW0-Delta

                        Hugging Face: https://huggingface.co/InternRobotics/InternW0-Delta-Base, https://huggingface.co/InternRobotics/InternW0-Delta-Libero, https://huggingface.co/InternRobotics/InternW0-Delta-RoboTwin, https://huggingface.co/InternRobotics/InternW0-Delta-RoboDojo
                        GitHub: https://github.com/InternRobotics/InternW0-Delta
                        Website: https://internrobotics.github.io/InternW0-Delta/
                        Technical report: https://arxiv.org/abs/2609.31394

                        Developer: Shanghai AI Laboratory (InternRobotics)
                        Released: September 2026
                        Variants: Base / Libero / RoboTwin / RoboDojo
                        Architecture: World-action model: a pretrained Wan2.2-TI2V-5B video expert and an ActionDiT action expert coupled through 30 directed Mixture-of-Transformers blocks, with a frozen RynnBrain1.1-2B VLM conditioning the action expert and training-only 4D distillation
                        License: MIT
                        Modalities: Multi-view RGB + language instruction + proprioceptive state in / action chunk out (canonical 80-D action space)
                        Runs on: NVIDIA GPU (Linux, CUDA); no inference requirement published
                        Formats: PyTorch (.pt)
                        On disk: 12.42GB per checkpoint (Base pretrain.pt)

                        1 Reply Last reply
                        0
                        • montezM Offline
                          montezM Offline
                          montez
                          wrote last edited by
                          #12

                          Alibaba DAMO RynnValue

                          Hugging Face: https://huggingface.co/Alibaba-DAMO-Academy/RynnValue-8B-Quantile, https://huggingface.co/Alibaba-DAMO-Academy/RynnValue-4B-Quantile, https://huggingface.co/Alibaba-DAMO-Academy/RynnValue-8B, https://huggingface.co/Alibaba-DAMO-Academy/RynnValue-4B
                          GitHub: https://github.com/alibaba-damo-academy/RynnValue
                          Website: https://alibaba-damo-academy.github.io/RynnValue.github.io/
                          Technical report: https://arxiv.org/abs/2608.09853

                          Developer: Alibaba DAMO Academy
                          Released: August 2026
                          Variants: RynnValue-4B, RynnValue-8B, RynnValue-4B-Quantile, RynnValue-8B-Quantile
                          Parameters: 5.14B (4B checkpoints), 9.57B (8B checkpoints)
                          Architecture: RynnBrain backbone on the Qwen3-VL architecture; absolute and relative distributional value heads (fixed-bin or 256-bin quantile) plus a language head; predicts remaining time to task completion per frame
                          License: Apache License 2.0
                          Modalities: Image/video + language in / per-frame remaining-time value in seconds, video analysis text out
                          Runs on: Desktop GPU, single GPU with 24GB+ memory for the 8B model in bf16
                          Formats: safetensors
                          On disk: 10.29GB (4B checkpoints), 19.15GB (8B checkpoints)

                          1 Reply Last reply
                          0
                          • montezM Offline
                            montezM Offline
                            montez
                            wrote last edited by
                            #13

                            Alibaba DAMO RynnWorld-Latent

                            Hugging Face: https://huggingface.co/Alibaba-DAMO-Academy/RynnWorld-Latent
                            GitHub: https://github.com/alibaba-damo-academy/RynnWorld-Latent
                            Website: https://alibaba-damo-academy.github.io/RynnWorld-Latent.github.io/

                            Developer: Alibaba DAMO Academy
                            Released: October 2026
                            Parameters: 3.46B
                            Architecture: Full fine-tune of NVIDIA Cosmos3-Edge (Mixture-of-Transformers, Nemotron-2B backbone, SigLIP2 vision tower) into a latent-action-conditioned video world model; rectified-flow video generation on a frozen Wan2.2 VAE; 608-dimension RynnLAM latent actions shared by human and robot video
                            Modalities: First-frame image + Latent action sequence in / Video out
                            Formats: safetensors
                            On disk: 20.75GB safetensors (4 shards, BF16 weights plus F32 EMA copy)

                            1 Reply Last reply
                            0
                            • montezM Offline
                              montezM Offline
                              montez
                              wrote last edited by
                              #14

                              Xiaomi Robotics Xiaomi-Robotics-U0

                              Hugging Face: https://huggingface.co/XiaomiRobotics/Xiaomi-Robotics-U0, https://huggingface.co/XiaomiRobotics/Xiaomi-Robotics-U0-FlashAR, https://huggingface.co/XiaomiRobotics/Xiaomi-Robotics-U0-4B, https://huggingface.co/XiaomiRobotics/Xiaomi-Robotics-U0-Sequence, https://huggingface.co/XiaomiRobotics/Xiaomi-Robotics-U0-4B-Sequence
                              GitHub: https://github.com/XiaomiRobotics/Xiaomi-Robotics-U0
                              Website: http://robotics.xiaomi.com/xiaomi-robotics-u0.html
                              Technical report: https://arxiv.org/abs/2607.11643

                              Developer: Xiaomi Robotics
                              Released: July 2026
                              Variants: U0, U0-FlashAR, U0-4B, U0-Sequence, U0-4B-Sequence
                              Parameters: 34B (U0, U0-Sequence), 4B (U0-4B, U0-4B-Sequence)
                              Architecture: Autoregressive transformer initialized from EMU3.5; shared discrete visual tokenizer, single next-token objective across text and image sequences; FlashAR decodes visual tokens in anti-diagonal groups; Sequence checkpoints interleave subtask text and observations
                              License: Apache-2.0
                              Modalities: Text + Image + Video
                              Runs on: U0-4B variants: Desktop GPU / U0, U0-Sequence, U0-FlashAR: Datacenter GPU, vendor benchmark on one NVIDIA H20
                              Formats: safetensors
                              On disk: 68.21GB (U0, U0-Sequence), 75.02GB (U0-FlashAR), 10.16GB (U0-4B, U0-4B-Sequence)

                              1 Reply Last reply
                              0
                              • montezM Offline
                                montezM Offline
                                montez
                                wrote last edited by
                                #15

                                TeleAI SMART-VLA

                                Hugging Face: https://huggingface.co/TeleEmbodied/SMART-VLA
                                GitHub: https://github.com/TeleHuman/PRTS
                                Website: https://teamillusion-smart.github.io/
                                Technical report: https://arxiv.org/abs/2610.07652

                                Developer: TeleAI (China Telecom)
                                Released: October 2026
                                Parameters: 4.44B
                                Architecture: PRTS-architecture VLA on a Qwen3-VL-4B-Instruct backbone with a flow-matching DiT action head, pretrained on synthetic articulated-object manipulation data
                                License: MIT
                                Modalities: Camera images + language instruction in / 50-step action chunk (up to 32 action dimensions) out
                                Formats: safetensors (BF16, 2 shards)
                                On disk: 9.67GB (two safetensors shards)

                                1 Reply Last reply
                                0
                                • montezM Offline
                                  montezM Offline
                                  montez
                                  wrote last edited by
                                  #16

                                  Microsoft Rho

                                  Hugging Face: https://huggingface.co/microsoft/rho-base, https://huggingface.co/collections/microsoft/rho
                                  GitHub: https://github.com/microsoft/rhobotics
                                  Website: https://microsoft.github.io/rhobotics/
                                  Technical report: https://arxiv.org/abs/2609.38164

                                  Developer: Microsoft Research
                                  Released: September 2026
                                  Variants: rho-base / rho-yam-box / rho-ur-ai-trainer / rho-fr3-duo / rho-libero / rho-roboeval
                                  Parameters: 5B (Phi-Phy VLM 4.68B / flow-matching action expert 542M)
                                  Architecture: Vision-language-action model: a physically grounded Phi-family VLM (Phi-Phy) with a 12-block flow-matching action expert that cross-attends to a VLM decoder layer, midtrained per embodiment
                                  License: MIT
                                  Modalities: Multi-camera images + language instruction + robot state in / action chunk out (up to 50 steps in pretraining)
                                  Runs on: NVIDIA GPU (Linux, CUDA; FlashAttention 2 by default); no VRAM requirement published
                                  Formats: safetensors (BF16, 3 shards)
                                  On disk: rho-base 10.50GB

                                  1 Reply Last reply
                                  0
                                  • montezM Offline
                                    montezM Offline
                                    montez
                                    wrote last edited by
                                    #17

                                    Meta FAIR RoboJEPA

                                    GitHub: https://github.com/facebookresearch/robo_jepa
                                    Website: https://robojepa.github.io
                                    Technical report: https://arxiv.org/abs/2610.10515

                                    Developer: Meta FAIR
                                    Released: October 2026
                                    Variants: 22M / 50M / 100M / 300M / 1B / 2B / 4B / 8B / 8B DROID 720p (3 views)
                                    Parameters: 22M to 8B predictors on a frozen V-JEPA 2.1 ViT-G/384 encoder
                                    Architecture: Action-conditioned JEPA latent world model: a transformer predictor over frozen V-JEPA 2.1 features, with latent CEM/MPC planning toward a goal image and an optional diffusion video decoder
                                    License: CC BY-NC-SA 4.0
                                    Modalities: Camera video + action sequence in / predicted future latent states out (optional decoded video); no action head, actions come from latent MPC planning
                                    Runs on: NVIDIA GPU (CUDA 12.6), one GPU per model size for evaluation; no VRAM requirement published
                                    Formats: PyTorch (.pth.tar) from dl.fbaipublicfiles.com
                                    On disk: 8B 29.99GB, 1B 4.07GB, 300M 1.24GB (300k-step checkpoints), plus V-JEPA 2.1 ViT-G/384 encoder 30.24GB

                                    1 Reply Last reply
                                    0
                                    • montezM Offline
                                      montezM Offline
                                      montez
                                      wrote last edited by
                                      #18

                                      DeepCybo PhysBrain 1.5

                                      Hugging Face: https://huggingface.co/DeepCybo/PhysBrain1.5-8B, https://huggingface.co/DeepCybo/PhysBrain1.5-2B
                                      GitHub: https://github.com/DeepCybo-PhysAI/PhysBrain-1.5
                                      Website: https://deepcybo-physai.github.io/PhysBrain-1.5/
                                      Technical report: https://arxiv.org/abs/2609.14973

                                      Developer: DeepCybo, Zhongguancun Academy, Zhongguancun Institute of Artificial Intelligence
                                      Released: September 2026
                                      Variants: 2B, 8B
                                      Parameters: 2.16B (2B checkpoint), 8.90B (8B checkpoint)
                                      Context: 262,144 tokens
                                      Architecture: Qwen3-VL backbone extended with action and visual-state tokens; one shared autoregressive transformer, single next-token objective, no task-specific heads; ActionPiece action tokens with one action codebook across robot setups
                                      Modalities: Text + Image + Video + Action history in / Text + Spatial grounding + Action chunk + Future-state image, depth map, robot mask out
                                      Runs on: 2B: Laptop, Desktop GPU / 8B: Desktop GPU
                                      Formats: safetensors
                                      On disk: 4.32GB (2B checkpoint), 17.81GB (8B checkpoint)

                                      1 Reply Last reply
                                      0
                                      • montezM Offline
                                        montezM Offline
                                        montez
                                        wrote last edited by
                                        #19

                                        Metacognition NavGPT3

                                        Hugging Face: https://huggingface.co/Metacognition-AI/NavGPT3-8B, https://huggingface.co/Metacognition-AI/NavGPT3-4B
                                        GitHub: https://github.com/metacognitionai/NavGPT-3
                                        Website: https://metacognitionai.github.io/NavGPT3/
                                        Technical report: https://arxiv.org/abs/2610.10787

                                        Developer: Metacognition
                                        Released: October 2026
                                        Variants: 4B, 8B
                                        Parameters: 4.44B (4B checkpoint), 8.77B (8B checkpoint)
                                        Architecture: Qwen3-VL fine-tune with a two-layer MLP action head on the last prompt token’s hidden state; one 3,072-token visual budget shared across a four-view, up to 16-step image history
                                        License: GNU Affero General Public License v3.0
                                        Modalities: Text + Four-view RGB image history in / 8 waypoints (x, y, theta) out
                                        Runs on: NVIDIA GPU (Linux)
                                        Formats: safetensors
                                        On disk: 9.66GB (4B checkpoint), 17.54GB (8B checkpoint)

                                        1 Reply Last reply
                                        0
                                        ↳

                                        OBJECTS Forum

                                        Join the conversation

                                        Create an account to return to your place in the thread, follow new replies, bookmark useful posts, and upvote contributions you value.

                                        Have something to add? Your perspective can make this thread better.

                                        Register Login
                                        Reply
                                        • Reply as topic
                                        Log in to reply
                                        • Oldest to Newest
                                        • Newest to Oldest
                                        • Most Votes


                                        • Login

                                        • Don't have an account? Register

                                        • Login or register to search.
                                        • First post
                                          Last post
                                        • 0
                                          • Categories
                                          • Recent
                                          • Tags
                                          • Popular
                                          • Users