Skip to content
  • Categories
  • Recent
  • Tags
  • Popular
  • Users
Menu
montezM

montez

@montez
Unfollow Follow
About
Posts
321
Topics
19
Shares
0
Groups
1
Followers
0
Following
0

Posts

Recent Best Controversial

  • Local LMs
    montezM montez

    Qwen-Image-2.1

    Hugging Face: https://huggingface.co/Qwen/Qwen-Image-2.1, https://huggingface.co/Qwen/Qwen-Image-2.1-Turbo
    GitHub: https://github.com/QwenLM/Qwen-Image-2.1
    Website: https://qwen.ai/blog?id=qwen-image-2.1

    Developer: Alibaba Cloud, Qwen team
    Released: September 2026
    Variants: Qwen-Image-2.1 (40 default denoising steps) / Qwen-Image-2.1-Turbo (8 denoising steps)
    Parameters: 7B visual generation component, plus Qwen3-VL 8B text encoder
    Resolution: native 2K, 2048x2048 default; recommended sizes include 2752x1536 and 1536x2752
    Architecture: single-stream DiT, 32 layers, block-causal attention with prefix KV cache reuse; Qwen3-VL 8B text encoder; 64-channel RGBA VAE with 16x spatial compression; flow matching with Euler scheduler
    License: Qwen Research License Agreement
    Modalities: Text-to-Image + Image editing (up to 10 reference images), native RGBA transparency
    Runs on: Desktop GPU, 24GB+ VRAM with model CPU offload
    Formats: BF16 safetensors, Diffusers pipeline
    On disk: Qwen-Image-2.1: 33.13GB (14.23GB transformer, 17.53GB text encoder, 1.35GB VAE) / Turbo: 32.49GB (14.23GB transformer, 17.53GB text encoder, 0.68GB VAE)

    AI & Software llm foundation models

  • Local LMs
    montezM montez

    NII LLM-jp-4

    Hugging Face: https://huggingface.co/llm-jp
    GitHub: https://github.com/llm-jp/llm-jp-4-cookbook
    Website: https://llm-jp.nii.ac.jp/blog/llm-jp-4-1/

    Developer: National Institute of Informatics, Research and Development Center for Large Language Models (LLMC), LLM-jp
    Released: April-September 2026, rolling family
    Variants: LLM-jp-4 8B and 32B-A3B (April 2026) / LLM-jp-4 33B (August 2026) / LLM-jp-4-VL 9B (September 2026) / LLM-jp-4.1 8B, 32B-A3B, 33B Thinking (September 2026)
    Parameters: 8B: 8.59B / 32B-A3B: 32.14B total, 3.83B active / 33B: 33.22B / VL 9B: 8.6B language model + 0.4B vision encoder
    Context: 65,536 tokens, text models
    Architecture: 8B and 33B: dense Llama architecture, 32 and 64 layers / 32B-A3B: Qwen3-MoE architecture, 32 layers, 128 routed experts with 8 active / VL 9B: llm-jp-4-8b-thinking, SigLIP 2 So400m vision encoder, 2-layer MLP projector
    License: Apache 2.0
    Modalities: 8B, 32B-A3B, 33B: Text / VL 9B: Text + Image
    Runs on: 8B: Laptop / 32B-A3B and 33B: Laptop, 32GB+ unified memory / VL 9B: Laptop, 24GB+ unified memory
    Formats: safetensors BF16, GGUF (BF16, Q4_K_M; LLM-jp-4.1 text models)
    On disk: LLM-jp-4.1 8B: 5.50GB Q4_K_M / 32B-A3B: 21.52GB Q4_K_M / 33B: 20.41GB Q4_K_M / VL 9B: 18.11GB BF16

    AI & Software llm foundation models

  • Embodied Foundation Models
    montezM montez

    Metacognition NavGPT3

    Hugging Face: https://huggingface.co/Metacognition-AI/NavGPT3-8B, https://huggingface.co/Metacognition-AI/NavGPT3-4B
    GitHub: https://github.com/metacognitionai/NavGPT-3
    Website: https://metacognitionai.github.io/NavGPT3/
    Technical report: https://arxiv.org/abs/2610.10787

    Developer: Metacognition
    Released: October 2026
    Variants: 4B, 8B
    Parameters: 4.44B (4B checkpoint), 8.77B (8B checkpoint)
    Architecture: Qwen3-VL fine-tune with a two-layer MLP action head on the last prompt token’s hidden state; one 3,072-token visual budget shared across a four-view, up to 16-step image history
    License: GNU Affero General Public License v3.0
    Modalities: Text + Four-view RGB image history in / 8 waypoints (x, y, theta) out
    Runs on: NVIDIA GPU (Linux)
    Formats: safetensors
    On disk: 9.66GB (4B checkpoint), 17.54GB (8B checkpoint)

    AI & Software embodied foundation models robotics

  • Local LMs
    montezM montez

    OpenBMB MiniCPM5

    Hugging Face: https://huggingface.co/openbmb/MiniCPM5-2B, https://huggingface.co/openbmb/MiniCPM5-1B
    GitHub: https://github.com/OpenBMB/MiniCPM

    Developer: OpenBMB
    Released: May-September 2026, rolling family
    Variants: MiniCPM5-2B (September 2026) / MiniCPM5-1B (May 2026)
    Parameters: 2B: 2.52B, 1.98B non-embedding / 1B: 1.08B, 0.68B non-embedding
    Context: 131,072 tokens
    Architecture: dense LlamaForCausalLM, GQA with 16 Q and 2 KV heads; 2B: 42 layers / 1B: 24 layers
    License: Apache 2.0
    Modalities: Text
    Runs on: Smartphone, Laptop
    Formats: safetensors BF16, GGUF (F16, Q8_0, Q4_K_M), MLX 4-bit, GPTQ 4-bit (2B)
    On disk: 2B: 1.56GB Q4_K_M, 5.03GB BF16 / 1B: 0.69GB Q4_K_M, 2.16GB BF16

    AI & Software llm foundation models

  • Embodied Foundation Models
    montezM montez

    DeepCybo PhysBrain 1.5

    Hugging Face: https://huggingface.co/DeepCybo/PhysBrain1.5-8B, https://huggingface.co/DeepCybo/PhysBrain1.5-2B
    GitHub: https://github.com/DeepCybo-PhysAI/PhysBrain-1.5
    Website: https://deepcybo-physai.github.io/PhysBrain-1.5/
    Technical report: https://arxiv.org/abs/2609.14973

    Developer: DeepCybo, Zhongguancun Academy, Zhongguancun Institute of Artificial Intelligence
    Released: September 2026
    Variants: 2B, 8B
    Parameters: 2.16B (2B checkpoint), 8.90B (8B checkpoint)
    Context: 262,144 tokens
    Architecture: Qwen3-VL backbone extended with action and visual-state tokens; one shared autoregressive transformer, single next-token objective, no task-specific heads; ActionPiece action tokens with one action codebook across robot setups
    Modalities: Text + Image + Video + Action history in / Text + Spatial grounding + Action chunk + Future-state image, depth map, robot mask out
    Runs on: 2B: Laptop, Desktop GPU / 8B: Desktop GPU
    Formats: safetensors
    On disk: 4.32GB (2B checkpoint), 17.81GB (8B checkpoint)

    AI & Software embodied foundation models robotics

  • Local LMs
    montezM montez

    SparkLLM Spark-X2.5

    Hugging Face: https://huggingface.co/XHToken/Spark-X2.5-4B, https://huggingface.co/XHToken/Spark-X2.5-1.7B
    GitHub: https://github.com/XHToken/Spark-X2.5
    Website: https://dev.to/sparkllm/spark-x25-4b-17b-the-only-on-device-models-with-native-1m-token-context-now-open-source-d9o

    Developer: SparkLLM
    Released: September 2026
    Variants: 4B / 1.7B
    Parameters: 4B / 1.7B
    Context: 1,048,576 tokens
    Architecture: dense; hybrid attention with 1 full-attention layer to 3 sliding-window layers, window 512; 4B: 36 layers, GQA 16 Q and 4 KV heads / 1.7B: 28 layers, GQA 8 Q and 2 KV heads
    License: Apache 2.0
    Modalities: Text
    Runs on: Smartphone, Laptop, Edge device
    Formats: safetensors BF16, FP8, INT8, GGUF (Q4_K_M, Q8_0, F16)
    On disk: 4B: 2.60GB Q4_K_M / 1.7B: 1.11GB Q4_K_M

    AI & Software llm foundation models

  • Embodied Foundation Models
    montezM montez

    Meta FAIR RoboJEPA

    GitHub: https://github.com/facebookresearch/robo_jepa
    Website: https://robojepa.github.io
    Technical report: https://arxiv.org/abs/2610.10515

    Developer: Meta FAIR
    Released: October 2026
    Variants: 22M / 50M / 100M / 300M / 1B / 2B / 4B / 8B / 8B DROID 720p (3 views)
    Parameters: 22M to 8B predictors on a frozen V-JEPA 2.1 ViT-G/384 encoder
    Architecture: Action-conditioned JEPA latent world model: a transformer predictor over frozen V-JEPA 2.1 features, with latent CEM/MPC planning toward a goal image and an optional diffusion video decoder
    License: CC BY-NC-SA 4.0
    Modalities: Camera video + action sequence in / predicted future latent states out (optional decoded video); no action head, actions come from latent MPC planning
    Runs on: NVIDIA GPU (CUDA 12.6), one GPU per model size for evaluation; no VRAM requirement published
    Formats: PyTorch (.pth.tar) from dl.fbaipublicfiles.com
    On disk: 8B 29.99GB, 1B 4.07GB, 300M 1.24GB (300k-step checkpoints), plus V-JEPA 2.1 ViT-G/384 encoder 30.24GB

    AI & Software embodied foundation models robotics

  • Local LMs
    montezM montez

    China Telecom Xing4.0

    Hugging Face: https://huggingface.co/XingChen-AGI/Xing4.0-29B-A4B, https://huggingface.co/XingChen-AGI/Xing4.0-29B-A4B-GGUF, https://huggingface.co/XingChen-AGI/Xing4.0-29B-A4B-FP8
    GitHub: https://github.com/XingChen-AGI/Xing4.0-29B-A4B

    Developer: China Telecom Artificial Intelligence Technology Co., Ltd.
    Released: September 2026
    Parameters: 29B total, 4B active
    Context: 262,144 tokens, extensible to 512K
    Architecture: MoE with mHC, MLA attention and MTP; 40 layers, 64 routed experts with 4 active plus 1 shared; trained on Ascend NPUs with MindSpore
    License: Apache 2.0
    Modalities: Text
    Runs on: Desktop GPU
    Formats: safetensors BF16, FP8, GGUF IQ4_NL
    On disk: 20.1GB IQ4_NL / 33.17GB FP8 / 62.43GB BF16

    AI & Software llm foundation models

  • Embodied Foundation Models
    montezM montez

    Microsoft Rho

    Hugging Face: https://huggingface.co/microsoft/rho-base, https://huggingface.co/collections/microsoft/rho
    GitHub: https://github.com/microsoft/rhobotics
    Website: https://microsoft.github.io/rhobotics/
    Technical report: https://arxiv.org/abs/2609.38164

    Developer: Microsoft Research
    Released: September 2026
    Variants: rho-base / rho-yam-box / rho-ur-ai-trainer / rho-fr3-duo / rho-libero / rho-roboeval
    Parameters: 5B (Phi-Phy VLM 4.68B / flow-matching action expert 542M)
    Architecture: Vision-language-action model: a physically grounded Phi-family VLM (Phi-Phy) with a 12-block flow-matching action expert that cross-attends to a VLM decoder layer, midtrained per embodiment
    License: MIT
    Modalities: Multi-camera images + language instruction + robot state in / action chunk out (up to 50 steps in pretraining)
    Runs on: NVIDIA GPU (Linux, CUDA; FlashAttention 2 by default); no VRAM requirement published
    Formats: safetensors (BF16, 3 shards)
    On disk: rho-base 10.50GB

    AI & Software embodied foundation models robotics

  • Local LMs
    montezM montez

    AI Singapore Nemotron-SEA-LION v4.8 30B-A3B

    Hugging Face: https://huggingface.co/aisingapore/Nemotron-SEA-LION-v4.8-30B-A3B, https://huggingface.co/aisingapore/Nemotron-SEA-LION-v4.8-30B-A3B-GGUF, https://huggingface.co/aisingapore/Nemotron-SEA-LION-v4.8-30B-A3B-FP8, https://huggingface.co/aisingapore/Nemotron-SEA-LION-v4.8-30B-A3B-NVFP4
    Website: https://sea-lion.ai/blog/uplifting-ai-in-southeast-asia-sea-announcing-nemotron-sea-lion-v4-8-in-collaboration-with-nvidia/
    Technical report: https://arxiv.org/abs/2609.18310

    Developer: AI Singapore, with NVIDIA
    Released: September 2026
    Parameters: 30B total, 3B active
    Context: 262,144 tokens
    Architecture: Mamba2-Transformer hybrid MoE; NVIDIA Nemotron 3 Nano 30B-A3B base, continued pretraining on 150B tokens, then SFT and on-policy distillation; English plus 7 Southeast Asian languages
    License: MIT
    Modalities: Text
    Runs on: Laptop, 32GB+ unified memory
    Formats: safetensors BF16, FP8, NVFP4, GGUF (Q4_K_M, Q6_K, Q8_0, F16)
    On disk: 25.43GB Q4_K_M / 22.94GB NVFP4 / 34.96GB FP8 / 65.83GB BF16

    AI & Software llm foundation models

  • Embodied Foundation Models
    montezM montez

    TeleAI SMART-VLA

    Hugging Face: https://huggingface.co/TeleEmbodied/SMART-VLA
    GitHub: https://github.com/TeleHuman/PRTS
    Website: https://teamillusion-smart.github.io/
    Technical report: https://arxiv.org/abs/2610.07652

    Developer: TeleAI (China Telecom)
    Released: October 2026
    Parameters: 4.44B
    Architecture: PRTS-architecture VLA on a Qwen3-VL-4B-Instruct backbone with a flow-matching DiT action head, pretrained on synthetic articulated-object manipulation data
    License: MIT
    Modalities: Camera images + language instruction in / 50-step action chunk (up to 32 action dimensions) out
    Formats: safetensors (BF16, 2 shards)
    On disk: 9.67GB (two safetensors shards)

    AI & Software embodied foundation models robotics

  • Local LMs
    montezM montez

    IFM K2 Horizon

    Hugging Face: https://huggingface.co/IFM/K2-Horizon-0.9B, https://huggingface.co/IFM/K2-Horizon-3.7B, https://huggingface.co/IFM/K2-Horizon-7B, https://huggingface.co/IFM/K2-Horizon-32B, https://huggingface.co/IFM/K2-Horizon-MoVA-36B-A4B
    GitHub: https://github.com/ifm-ai/xllm
    Website: https://ifm.ai/k2/

    Developer: Institute of Foundation Models (IFM), MBZUAI
    Released: September 2026
    Variants: 0.9B / 3.7B / 7B / 32B / MoVA-36B-A4B
    Parameters: 0.9B / 3.7B / 7B / 32B / 36B total, 4B active
    Context: 0.9B: 131,072 tokens / 3.7B, 7B, 32B and MoVA-36B-A4B: 524,288 tokens
    Architecture: 0.9B, 3.7B, 7B, 32B: dense decoder-only, GQA with 8 KV heads, 28 / 36 / 36 / 64 layers; MoVA-36B-A4B: MoE with Mixture-of-Values attention, 48 layers, 100 FFN experts with 8 per token plus 1 shared, 64 value experts with 4 per token
    License: Apache 2.0
    Modalities: Text
    Runs on: 0.9B: Smartwatch, Smart glasses, Smartphone / 3.7B and 7B: Smartphone, Laptop / 32B and MoVA-36B-A4B: Laptop, 32GB+ unified memory
    Formats: safetensors BF16, GGUF (Q4_K_M, Q5_0, Q5_K_M, Q6_K, Q8_0, BF16), FP8 (7B, 32B, MoVA-36B-A4B), NVFP4 (32B)
    On disk: 0.9B: 0.67GB Q4_K_M / 3.7B: 3.16GB Q4_K_M / 7B: 5.59GB Q4_K_M / 32B: 21.08GB Q4_K_M / MoVA-36B-A4B: 22.37GB Q4_K_M

    AI & Software llm foundation models

  • Embodied Foundation Models
    montezM montez

    Xiaomi Robotics Xiaomi-Robotics-U0

    Hugging Face: https://huggingface.co/XiaomiRobotics/Xiaomi-Robotics-U0, https://huggingface.co/XiaomiRobotics/Xiaomi-Robotics-U0-FlashAR, https://huggingface.co/XiaomiRobotics/Xiaomi-Robotics-U0-4B, https://huggingface.co/XiaomiRobotics/Xiaomi-Robotics-U0-Sequence, https://huggingface.co/XiaomiRobotics/Xiaomi-Robotics-U0-4B-Sequence
    GitHub: https://github.com/XiaomiRobotics/Xiaomi-Robotics-U0
    Website: http://robotics.xiaomi.com/xiaomi-robotics-u0.html
    Technical report: https://arxiv.org/abs/2607.11643

    Developer: Xiaomi Robotics
    Released: July 2026
    Variants: U0, U0-FlashAR, U0-4B, U0-Sequence, U0-4B-Sequence
    Parameters: 34B (U0, U0-Sequence), 4B (U0-4B, U0-4B-Sequence)
    Architecture: Autoregressive transformer initialized from EMU3.5; shared discrete visual tokenizer, single next-token objective across text and image sequences; FlashAR decodes visual tokens in anti-diagonal groups; Sequence checkpoints interleave subtask text and observations
    License: Apache-2.0
    Modalities: Text + Image + Video
    Runs on: U0-4B variants: Desktop GPU / U0, U0-Sequence, U0-FlashAR: Datacenter GPU, vendor benchmark on one NVIDIA H20
    Formats: safetensors
    On disk: 68.21GB (U0, U0-Sequence), 75.02GB (U0-FlashAR), 10.16GB (U0-4B, U0-4B-Sequence)

    AI & Software embodied foundation models robotics

  • Local LMs
    montezM montez

    Edge0

    Hugging Face: https://huggingface.co/Edge0/Edge0-35B-A3B-preview, https://huggingface.co/Edge0/Edge0-8B-A1B-preview
    GitHub: https://github.com/Edge0-AI/Edge0
    Technical report: https://arxiv.org/abs/2609.18063

    Developer: Edge0 AI
    Released: September 2026
    Variants: Edge0-35B-A3B-preview / Edge0-8B-A1B-preview
    Parameters: 35B total, 3B active / 8B total, 1B active
    Context: 262,144 tokens / 131,072 tokens
    Architecture: early preview release; sparse MoE; 35B-A3B: Qwen3.6-35B-A3B base, 40 layers, 256 experts, 8 per token / 8B-A1B: Ling 3.0 tiny base, 24 layers, 128 experts, 8 per token; int4 checkpoint plus Recover-LoRA adapters and prerouter heads, with experts streamed from SSD by the edge0 framework
    License: Apache 2.0
    Modalities: Text
    Runs on: 35B-A3B: Smartphone, 12GB+ RAM on Android / 8B-A1B: Smartphone, 8GB+ RAM on Android; edge0 engines only (iOS, macOS, Android, Windows)
    Formats: MLX 4-bit safetensors, LoRA and prerouter adapter safetensors
    On disk: 35B-A3B: 19.51GB int4 checkpoint plus 0.18GB adapters / 8B-A1B: 4.51GB int4 checkpoint plus 0.06GB adapters

    AI & Software llm foundation models

  • Embodied Foundation Models
    montezM montez

    Alibaba DAMO RynnWorld-Latent

    Hugging Face: https://huggingface.co/Alibaba-DAMO-Academy/RynnWorld-Latent
    GitHub: https://github.com/alibaba-damo-academy/RynnWorld-Latent
    Website: https://alibaba-damo-academy.github.io/RynnWorld-Latent.github.io/

    Developer: Alibaba DAMO Academy
    Released: October 2026
    Parameters: 3.46B
    Architecture: Full fine-tune of NVIDIA Cosmos3-Edge (Mixture-of-Transformers, Nemotron-2B backbone, SigLIP2 vision tower) into a latent-action-conditioned video world model; rectified-flow video generation on a frozen Wan2.2 VAE; 608-dimension RynnLAM latent actions shared by human and robot video
    Modalities: First-frame image + Latent action sequence in / Video out
    Formats: safetensors
    On disk: 20.75GB safetensors (4 shards, BF16 weights plus F32 EMA copy)

    AI & Software embodied foundation models robotics

  • Local LMs
    montezM montez

    Swiss AI Apertus 1.5

    Hugging Face: https://huggingface.co/swiss-ai/Apertus-v1.5-8B, https://huggingface.co/swiss-ai/Apertus-v1.5-70B
    Website: https://www.apertus-ai.org/articles/2026-07-apertus-1-5
    Docs: https://www.apertus-ai.org/docs

    Developer: Swiss AI Initiative (EPFL, ETH Zurich, CSCS)
    Released: July 2026
    Variants: 8B / 70B
    Parameters: 8B / 70B
    Context: 262,144 tokens
    Architecture: decoder-only transformer with xIELU activation, trained with the AdEMAMix optimizer; continued pretraining of Apertus 1.0 on 4T added tokens (8B) and 2T added tokens (70B); optional thinking mode and tool calling
    License: Apache 2.0 with Acceptable Use Policy
    Modalities: Text + Image + Audio (audio experimental), text output
    Runs on: 8B: Laptop, Desktop GPU / 70B: Server-class hardware
    Formats: safetensors
    On disk: 8B: 18.40GB safetensors / 70B: 144.60GB safetensors

    AI & Software llm foundation models

  • Embodied Foundation Models
    montezM montez

    Alibaba DAMO RynnValue

    Hugging Face: https://huggingface.co/Alibaba-DAMO-Academy/RynnValue-8B-Quantile, https://huggingface.co/Alibaba-DAMO-Academy/RynnValue-4B-Quantile, https://huggingface.co/Alibaba-DAMO-Academy/RynnValue-8B, https://huggingface.co/Alibaba-DAMO-Academy/RynnValue-4B
    GitHub: https://github.com/alibaba-damo-academy/RynnValue
    Website: https://alibaba-damo-academy.github.io/RynnValue.github.io/
    Technical report: https://arxiv.org/abs/2608.09853

    Developer: Alibaba DAMO Academy
    Released: August 2026
    Variants: RynnValue-4B, RynnValue-8B, RynnValue-4B-Quantile, RynnValue-8B-Quantile
    Parameters: 5.14B (4B checkpoints), 9.57B (8B checkpoints)
    Architecture: RynnBrain backbone on the Qwen3-VL architecture; absolute and relative distributional value heads (fixed-bin or 256-bin quantile) plus a language head; predicts remaining time to task completion per frame
    License: Apache License 2.0
    Modalities: Image/video + language in / per-frame remaining-time value in seconds, video analysis text out
    Runs on: Desktop GPU, single GPU with 24GB+ memory for the 8B model in bf16
    Formats: safetensors
    On disk: 10.29GB (4B checkpoints), 19.15GB (8B checkpoints)

    AI & Software embodied foundation models robotics

  • Local LMs
    montezM montez

    PrismML Bonsai 2 27B

    Hugging Face: https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-mlx-2bit, https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf
    GitHub: https://github.com/PrismML-Eng/Bonsai-demo
    Website: https://prismml.com/news/bonsai-2-27b
    White paper: https://github.com/PrismML-Eng/Bonsai-demo/blob/main/bonsai-2-27b-whitepaper.pdf

    Developer: PrismML
    Released: September 2026
    Parameters: 27.36B total, 24.35B language + 2.54B embedding and LM head + 0.47B vision tower
    Context: 262,144 tokens
    Architecture: Qwen3.8-27B base; 64 blocks, 24Q/4KV heads, hybrid attention with about 75% linear and 25% full attention; ternary {-1, 0, +1} weights with FP16 group-wise scales at group size 128, 1.76 effective bits per weight
    License: Apache 2.0
    Modalities: Text + Image
    Runs on: Laptop, Desktop GPU; Apple Silicon or NVIDIA GPU only
    Formats: MLX 2-bit ternary, GGUF PTQ1_0, GGUF PQ2_0, optional mmproj vision pack
    On disk: 5.95GB PTQ1_0 GGUF, 7.21GB PQ2_0 GGUF, 8.60GB MLX / 0.63GB mmproj vision pack

    AI & Software llm foundation models

  • Embodied Foundation Models
    montezM montez

    InternRobotics InternW0-Delta

    Hugging Face: https://huggingface.co/InternRobotics/InternW0-Delta-Base, https://huggingface.co/InternRobotics/InternW0-Delta-Libero, https://huggingface.co/InternRobotics/InternW0-Delta-RoboTwin, https://huggingface.co/InternRobotics/InternW0-Delta-RoboDojo
    GitHub: https://github.com/InternRobotics/InternW0-Delta
    Website: https://internrobotics.github.io/InternW0-Delta/
    Technical report: https://arxiv.org/abs/2609.31394

    Developer: Shanghai AI Laboratory (InternRobotics)
    Released: September 2026
    Variants: Base / Libero / RoboTwin / RoboDojo
    Architecture: World-action model: a pretrained Wan2.2-TI2V-5B video expert and an ActionDiT action expert coupled through 30 directed Mixture-of-Transformers blocks, with a frozen RynnBrain1.1-2B VLM conditioning the action expert and training-only 4D distillation
    License: MIT
    Modalities: Multi-view RGB + language instruction + proprioceptive state in / action chunk out (canonical 80-D action space)
    Runs on: NVIDIA GPU (Linux, CUDA); no inference requirement published
    Formats: PyTorch (.pt)
    On disk: 12.42GB per checkpoint (Base pretrain.pt)

    AI & Software embodied foundation models robotics

  • Local LMs
    montezM montez

    JetBrains Mellum2.1

    Hugging Face: https://huggingface.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking, https://huggingface.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking-GGUF, https://huggingface.co/JetBrains/Mellum2-12B-A2.5B-Base, https://huggingface.co/JetBrains/Mellum2-12B-A2.5B-Instruct, https://huggingface.co/JetBrains/Mellum2-12B-A2.5B-Thinking
    Website: https://blog.jetbrains.com/ai/2026/10/mellum2-1-gets-to-work-a-fast-open-model-for-coding-agents/
    Technical report: https://arxiv.org/abs/2605.31268

    Developer: JetBrains
    Released: October 2026
    Variants: Mellum2.1 Thinking / Mellum2 Base, Instruct, Thinking (June 2026)
    Parameters: 12B total, 2.5B active
    Context: 131,072 tokens
    Architecture: MoE, 28 layers, 64 experts with 8 active, GQA with 32 Q and 4 KV heads, sliding window 1,024 on 3 of every 4 layers; Mellum2.1 keeps the Mellum2 architecture and adds reinforcement learning post-training
    License: Apache 2.0
    Modalities: Text
    Runs on: Laptop, 16GB+ unified memory
    Formats: safetensors BF16, GGUF (BF16, Q8_0, Q6_K, Q4_K_M, MXFP4_MOE)
    On disk: Mellum2.1 Thinking: 8.07GB Q4_K_M, 24.31GB BF16

    AI & Software llm foundation models
  • Login

  • Don't have an account? Register

  • Login or register to search.
  • First post
    Last post
  • 0
    • Categories
    • Recent
    • Tags
    • Popular
    • Users