<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Embodied Foundation Models]]></title><description><![CDATA[<p dir="auto">Megathread for embodied foundation models for perception, spatial reasoning, and robot action, from device-native models to large world-action models.</p>
]]></description><link>https://forum.objects.foundation/topic/27/embodied-foundation-models</link><generator>RSS for Node</generator><lastBuildDate>Sun, 30 Aug 2026 06:03:28 GMT</lastBuildDate><atom:link href="https://forum.objects.foundation/topic/27.rss" rel="self" type="application/rss+xml"/><pubDate>Sat, 29 Aug 2026 23:50:35 GMT</pubDate><ttl>60</ttl><item><title><![CDATA[Reply to Embodied Foundation Models on Sun, 30 Aug 2026 04:02:45 GMT]]></title><description><![CDATA[<h1>Alibaba DAMO RynnBrain 1.1</h1>
<p dir="auto"><strong>Hugging Face</strong>: <a href="https://huggingface.co/Alibaba-DAMO-Academy/RynnBrain1.1-2B" target="_blank" rel="noopener noreferrer nofollow ugc">https://huggingface.co/Alibaba-DAMO-Academy/RynnBrain1.1-2B</a>, <a href="https://huggingface.co/Alibaba-DAMO-Academy/RynnBrain1.1-9B" target="_blank" rel="noopener noreferrer nofollow ugc">https://huggingface.co/Alibaba-DAMO-Academy/RynnBrain1.1-9B</a>, <a href="https://huggingface.co/Alibaba-DAMO-Academy/RynnBrain1.1-122B-A10B" target="_blank" rel="noopener noreferrer nofollow ugc">https://huggingface.co/Alibaba-DAMO-Academy/RynnBrain1.1-122B-A10B</a><br />
<strong>GitHub</strong>: <a href="https://github.com/alibaba-damo-academy/RynnBrain" target="_blank" rel="noopener noreferrer nofollow ugc">https://github.com/alibaba-damo-academy/RynnBrain</a><br />
<strong>Technical report</strong>: <a href="https://arxiv.org/abs/2602.14979" target="_blank" rel="noopener noreferrer nofollow ugc">https://arxiv.org/abs/2602.14979</a></p>
<p dir="auto"><strong>Developer</strong>: Alibaba DAMO Academy<br />
<strong>Released</strong>: July 2026<br />
<strong>Variants</strong>: 2B, 9B, 122B-A10B<br />
<strong>Parameters</strong>: 9.41B (9B checkpoint)<br />
<strong>Architecture</strong>: Decoder-only vision-language transformer (dense 2B/9B, sparse-MoE 122B-A10B) on a Qwen3.5 base<br />
<strong>License</strong>: Apache License 2.0<br />
<strong>Modalities</strong>: Image/video + language in / spatial/3D grounding, contact-point and affordance predictions, task planning out<br />
<strong>Formats</strong>: safetensors<br />
<strong>On disk</strong>: 18.82GB (9B checkpoint)</p>
]]></description><link>https://forum.objects.foundation/post/183</link><guid isPermaLink="true">https://forum.objects.foundation/post/183</guid><dc:creator><![CDATA[montez]]></dc:creator><pubDate>Sun, 30 Aug 2026 04:02:45 GMT</pubDate></item><item><title><![CDATA[Reply to Embodied Foundation Models on Sun, 30 Aug 2026 04:02:34 GMT]]></title><description><![CDATA[<h1>Unitree UnifoLM-VLA-0</h1>
<p dir="auto"><strong>Hugging Face</strong>: <a href="https://huggingface.co/unitreerobotics/Unifolm-VLM-Base" target="_blank" rel="noopener noreferrer nofollow ugc">https://huggingface.co/unitreerobotics/Unifolm-VLM-Base</a>, <a href="https://huggingface.co/unitreerobotics/Unifolm-VLA-Base" target="_blank" rel="noopener noreferrer nofollow ugc">https://huggingface.co/unitreerobotics/Unifolm-VLA-Base</a>, <a href="https://huggingface.co/unitreerobotics/Unifolm-VLA-Libero" target="_blank" rel="noopener noreferrer nofollow ugc">https://huggingface.co/unitreerobotics/Unifolm-VLA-Libero</a><br />
<strong>GitHub</strong>: <a href="https://github.com/unitreerobotics/unifolm-vla" target="_blank" rel="noopener noreferrer nofollow ugc">https://github.com/unitreerobotics/unifolm-vla</a><br />
<strong>Website</strong>: <a href="https://unigen-x.github.io/unifolm-vla.github.io" target="_blank" rel="noopener noreferrer nofollow ugc">https://unigen-x.github.io/unifolm-vla.github.io</a></p>
<p dir="auto"><strong>Developer</strong>: Unitree Robotics<br />
<strong>Released</strong>: January 2026<br />
<strong>Variants</strong>: VLM-Base, VLA-Base, VLA-LIBERO<br />
<strong>Architecture</strong>: Qwen2.5-VL-7B backbone with a diffusion-transformer flow-matching action head; continued pretraining fuses 2D/3D spatial understanding with action-chunking prediction and forward/inverse dynamics constraints<br />
<strong>Modalities</strong>: Multi-view RGB + robot state + language instruction in / continuous manipulation actions out<br />
<strong>Formats</strong>: PyTorch checkpoint<br />
<strong>On disk</strong>: VLA-Base 18.98GB</p>
]]></description><link>https://forum.objects.foundation/post/182</link><guid isPermaLink="true">https://forum.objects.foundation/post/182</guid><dc:creator><![CDATA[montez]]></dc:creator><pubDate>Sun, 30 Aug 2026 04:02:34 GMT</pubDate></item><item><title><![CDATA[Reply to Embodied Foundation Models on Sun, 30 Aug 2026 04:02:23 GMT]]></title><description><![CDATA[<h1>NVIDIA Cosmos Reason 2</h1>
<p dir="auto"><strong>Hugging Face</strong>: <a href="https://huggingface.co/nvidia/Cosmos-Reason2-8B" target="_blank" rel="noopener noreferrer nofollow ugc">https://huggingface.co/nvidia/Cosmos-Reason2-8B</a>, <a href="https://huggingface.co/nvidia/Cosmos-Reason2-2B" target="_blank" rel="noopener noreferrer nofollow ugc">https://huggingface.co/nvidia/Cosmos-Reason2-2B</a>, <a href="https://huggingface.co/nvidia/Cosmos-Reason2-32B" target="_blank" rel="noopener noreferrer nofollow ugc">https://huggingface.co/nvidia/Cosmos-Reason2-32B</a><br />
<strong>Website</strong>: <a href="https://build.nvidia.com/nvidia/cosmos-reason2-8b" target="_blank" rel="noopener noreferrer nofollow ugc">https://build.nvidia.com/nvidia/cosmos-reason2-8b</a><br />
<strong>Docs</strong>: <a href="https://huggingface.co/blog/nvidia/nvidia-cosmos-reason-2-brings-advanced-reasoning" target="_blank" rel="noopener noreferrer nofollow ugc">https://huggingface.co/blog/nvidia/nvidia-cosmos-reason-2-brings-advanced-reasoning</a></p>
<p dir="auto"><strong>Developer</strong>: NVIDIA<br />
<strong>Released</strong>: January 2026<br />
<strong>Variants</strong>: 2B, 8B, 32B<br />
<strong>Parameters</strong>: 8.77B (8B tier)<br />
<strong>Architecture</strong>: Qwen3-VL fine-tune post-trained for physical-AI reasoning<br />
<strong>License</strong>: NVIDIA Open Model License Agreement<br />
<strong>Modalities</strong>: Image/video + text in / spatio-temporal reasoning, trajectory, point, and bounding-box predictions out<br />
<strong>Runs on</strong>: NVIDIA GPU<br />
<strong>Formats</strong>: safetensors, 4 shards (8B tier)<br />
<strong>On disk</strong>: 17.53GB (8B tier)</p>
]]></description><link>https://forum.objects.foundation/post/181</link><guid isPermaLink="true">https://forum.objects.foundation/post/181</guid><dc:creator><![CDATA[montez]]></dc:creator><pubDate>Sun, 30 Aug 2026 04:02:23 GMT</pubDate></item><item><title><![CDATA[Reply to Embodied Foundation Models on Sun, 30 Aug 2026 04:00:29 GMT]]></title><description><![CDATA[<h1>NVIDIA Isaac GR00T N1.7</h1>
<p dir="auto"><strong>Hugging Face</strong>: <a href="https://huggingface.co/nvidia/GR00T-N1.7-3B" target="_blank" rel="noopener noreferrer nofollow ugc">https://huggingface.co/nvidia/GR00T-N1.7-3B</a>, <a href="https://huggingface.co/collections/nvidia/gr00t-n17" target="_blank" rel="noopener noreferrer nofollow ugc">https://huggingface.co/collections/nvidia/gr00t-n17</a><br />
<strong>GitHub</strong>: <a href="https://github.com/NVIDIA/Isaac-GR00T" target="_blank" rel="noopener noreferrer nofollow ugc">https://github.com/NVIDIA/Isaac-GR00T</a><br />
<strong>Website</strong>: <a href="https://developer.nvidia.com/isaac/gr00t" target="_blank" rel="noopener noreferrer nofollow ugc">https://developer.nvidia.com/isaac/gr00t</a><br />
<strong>Docs</strong>: <a href="https://huggingface.co/blog/nvidia/gr00t-n1-7" target="_blank" rel="noopener noreferrer nofollow ugc">https://huggingface.co/blog/nvidia/gr00t-n1-7</a></p>
<p dir="auto"><strong>Developer</strong>: NVIDIA<br />
<strong>Released</strong>: April 2026<br />
<strong>Variants</strong>: N1, N1.5, N1.6, N1.7<br />
<strong>Parameters</strong>: 3B<br />
<strong>Architecture</strong>: Dual-system VLA pairing a Cosmos-Reason2-2B vision-language backbone for task/subtask reasoning with a flow-matching diffusion transformer conditioned on the reasoning output<br />
<strong>License</strong>: NVIDIA Open Model License Agreement<br />
<strong>Modalities</strong>: RGB camera frames + language instruction + robot proprioception in / continuous robot action vectors out<br />
<strong>Runs on</strong>: Desktop GPU, Datacenter GPU, or NVIDIA Jetson edge modules; NVIDIA only<br />
<strong>Formats</strong>: safetensors, 2 shards<br />
<strong>On disk</strong>: 6.91GB</p>
]]></description><link>https://forum.objects.foundation/post/180</link><guid isPermaLink="true">https://forum.objects.foundation/post/180</guid><dc:creator><![CDATA[montez]]></dc:creator><pubDate>Sun, 30 Aug 2026 04:00:29 GMT</pubDate></item><item><title><![CDATA[Reply to Embodied Foundation Models on Sun, 30 Aug 2026 01:14:00 GMT]]></title><description><![CDATA[<h1>Tencent Hy-Embodied-RxBrain-1.0</h1>
<p dir="auto"><strong>Hugging Face</strong>: <a href="https://huggingface.co/tencent/Hy-Embodied-RxBrain-1.0" target="_blank" rel="noopener noreferrer nofollow ugc">https://huggingface.co/tencent/Hy-Embodied-RxBrain-1.0</a><br />
<strong>GitHub</strong>: <a href="https://github.com/Tencent-Hunyuan/Hy-Embodied-RxBrain-1.0" target="_blank" rel="noopener noreferrer nofollow ugc">https://github.com/Tencent-Hunyuan/Hy-Embodied-RxBrain-1.0</a><br />
<strong>Website</strong>: <a href="https://tairos.tencent.com/openSourceModels/hy-embodied-rxbrain-1.0" target="_blank" rel="noopener noreferrer nofollow ugc">https://tairos.tencent.com/openSourceModels/hy-embodied-rxbrain-1.0</a><br />
<strong>Technical report</strong>: <a href="https://arxiv.org/abs/2607.14187" target="_blank" rel="noopener noreferrer nofollow ugc">https://arxiv.org/abs/2607.14187</a></p>
<p dir="auto"><strong>Developer</strong>: Tencent Robotics X, Futian Laboratory, Tencent Hy Team<br />
<strong>Released</strong>: July 2026<br />
<strong>Parameters</strong>: ~6.2B<br />
<strong>Architecture</strong>: Unified Mixture-of-Transformers, modality-specific text, vision, and generation pathways<br />
<strong>License</strong>: Apache 2.0<br />
<strong>Modalities</strong>: Text + Image + Video<br />
<strong>Runs on</strong>: NVIDIA GPU, CUDA 12.x, Linux recommended<br />
<strong>Formats</strong>: safetensors<br />
<strong>On disk</strong>: 12.42GB safetensors</p>
]]></description><link>https://forum.objects.foundation/post/177</link><guid isPermaLink="true">https://forum.objects.foundation/post/177</guid><dc:creator><![CDATA[montez]]></dc:creator><pubDate>Sun, 30 Aug 2026 01:14:00 GMT</pubDate></item><item><title><![CDATA[Reply to Embodied Foundation Models on Sun, 30 Aug 2026 01:13:47 GMT]]></title><description><![CDATA[<h1>NVIDIA Cosmos 3 Edge</h1>
<p dir="auto"><strong>Hugging Face</strong>: <a href="https://huggingface.co/nvidia/Cosmos3-Edge" target="_blank" rel="noopener noreferrer nofollow ugc">https://huggingface.co/nvidia/Cosmos3-Edge</a>, <a href="https://huggingface.co/collections/nvidia/cosmos3" target="_blank" rel="noopener noreferrer nofollow ugc">https://huggingface.co/collections/nvidia/cosmos3</a><br />
<strong>GitHub</strong>: <a href="https://github.com/nvidia/cosmos" target="_blank" rel="noopener noreferrer nofollow ugc">https://github.com/nvidia/cosmos</a><br />
<strong>Website</strong>: <a href="https://research.nvidia.com/labs/cosmos-lab/cosmos3/" target="_blank" rel="noopener noreferrer nofollow ugc">https://research.nvidia.com/labs/cosmos-lab/cosmos3/</a><br />
<strong>White paper</strong>: <a href="https://research.nvidia.com/labs/cosmos-lab/cosmos3/technical-report.pdf" target="_blank" rel="noopener noreferrer nofollow ugc">https://research.nvidia.com/labs/cosmos-lab/cosmos3/technical-report.pdf</a></p>
<p dir="auto"><strong>Developer</strong>: NVIDIA<br />
<strong>Released</strong>: July 2026<br />
<strong>Variants</strong>: Cosmos3-Edge, Cosmos3-Edge-Policy-DROID<br />
<strong>Parameters</strong>: 4B<br />
<strong>Context</strong>: Reasoner: 256,000 tokens / Generator text input: 4,096 tokens<br />
<strong>Architecture</strong>: Mixture-of-Transformers, two towers: an autoregressive transformer for text, a diffusion transformer for image/video/action generation<br />
<strong>License</strong>: OpenMDW 1.1<br />
<strong>Modalities</strong>: Text + Image + Video + Action trajectory in / Text + Image + Video + Action out<br />
<strong>Formats</strong>: safetensors via Hugging Face, PyTorch (NVIDIA-proprietary inference only, BF16)<br />
<strong>Runs on</strong>: NVIDIA consumer GPU (Linux), Jetson (Linux)</p>
]]></description><link>https://forum.objects.foundation/post/176</link><guid isPermaLink="true">https://forum.objects.foundation/post/176</guid><dc:creator><![CDATA[montez]]></dc:creator><pubDate>Sun, 30 Aug 2026 01:13:47 GMT</pubDate></item></channel></rss>