Skip to content
  • Categories
  • Recent
  • Tags
  • Popular
  • Users
Menu

administrators

Private

Posts


  • Local LMs
    montezM montez

    Qwen-Image-2.1

    Hugging Face: https://huggingface.co/Qwen/Qwen-Image-2.1, https://huggingface.co/Qwen/Qwen-Image-2.1-Turbo
    GitHub: https://github.com/QwenLM/Qwen-Image-2.1
    Website: https://qwen.ai/blog?id=qwen-image-2.1

    Developer: Alibaba Cloud, Qwen team
    Released: September 2026
    Variants: Qwen-Image-2.1 (40 default denoising steps) / Qwen-Image-2.1-Turbo (8 denoising steps)
    Parameters: 7B visual generation component, plus Qwen3-VL 8B text encoder
    Resolution: native 2K, 2048x2048 default; recommended sizes include 2752x1536 and 1536x2752
    Architecture: single-stream DiT, 32 layers, block-causal attention with prefix KV cache reuse; Qwen3-VL 8B text encoder; 64-channel RGBA VAE with 16x spatial compression; flow matching with Euler scheduler
    License: Qwen Research License Agreement
    Modalities: Text-to-Image + Image editing (up to 10 reference images), native RGBA transparency
    Runs on: Desktop GPU, 24GB+ VRAM with model CPU offload
    Formats: BF16 safetensors, Diffusers pipeline
    On disk: Qwen-Image-2.1: 33.13GB (14.23GB transformer, 17.53GB text encoder, 1.35GB VAE) / Turbo: 32.49GB (14.23GB transformer, 17.53GB text encoder, 0.68GB VAE)

    AI & Software llm foundation models

  • Local LMs
    montezM montez

    NII LLM-jp-4

    Hugging Face: https://huggingface.co/llm-jp
    GitHub: https://github.com/llm-jp/llm-jp-4-cookbook
    Website: https://llm-jp.nii.ac.jp/blog/llm-jp-4-1/

    Developer: National Institute of Informatics, Research and Development Center for Large Language Models (LLMC), LLM-jp
    Released: April-September 2026, rolling family
    Variants: LLM-jp-4 8B and 32B-A3B (April 2026) / LLM-jp-4 33B (August 2026) / LLM-jp-4-VL 9B (September 2026) / LLM-jp-4.1 8B, 32B-A3B, 33B Thinking (September 2026)
    Parameters: 8B: 8.59B / 32B-A3B: 32.14B total, 3.83B active / 33B: 33.22B / VL 9B: 8.6B language model + 0.4B vision encoder
    Context: 65,536 tokens, text models
    Architecture: 8B and 33B: dense Llama architecture, 32 and 64 layers / 32B-A3B: Qwen3-MoE architecture, 32 layers, 128 routed experts with 8 active / VL 9B: llm-jp-4-8b-thinking, SigLIP 2 So400m vision encoder, 2-layer MLP projector
    License: Apache 2.0
    Modalities: 8B, 32B-A3B, 33B: Text / VL 9B: Text + Image
    Runs on: 8B: Laptop / 32B-A3B and 33B: Laptop, 32GB+ unified memory / VL 9B: Laptop, 24GB+ unified memory
    Formats: safetensors BF16, GGUF (BF16, Q4_K_M; LLM-jp-4.1 text models)
    On disk: LLM-jp-4.1 8B: 5.50GB Q4_K_M / 32B-A3B: 21.52GB Q4_K_M / 33B: 20.41GB Q4_K_M / VL 9B: 18.11GB BF16

    AI & Software llm foundation models

  • Embodied Foundation Models
    montezM montez

    Metacognition NavGPT3

    Hugging Face: https://huggingface.co/Metacognition-AI/NavGPT3-8B, https://huggingface.co/Metacognition-AI/NavGPT3-4B
    GitHub: https://github.com/metacognitionai/NavGPT-3
    Website: https://metacognitionai.github.io/NavGPT3/
    Technical report: https://arxiv.org/abs/2610.10787

    Developer: Metacognition
    Released: October 2026
    Variants: 4B, 8B
    Parameters: 4.44B (4B checkpoint), 8.77B (8B checkpoint)
    Architecture: Qwen3-VL fine-tune with a two-layer MLP action head on the last prompt token’s hidden state; one 3,072-token visual budget shared across a four-view, up to 16-step image history
    License: GNU Affero General Public License v3.0
    Modalities: Text + Four-view RGB image history in / 8 waypoints (x, y, theta) out
    Runs on: NVIDIA GPU (Linux)
    Formats: safetensors
    On disk: 9.66GB (4B checkpoint), 17.54GB (8B checkpoint)

    AI & Software embodied foundation models robotics

  • Local LMs
    montezM montez

    OpenBMB MiniCPM5

    Hugging Face: https://huggingface.co/openbmb/MiniCPM5-2B, https://huggingface.co/openbmb/MiniCPM5-1B
    GitHub: https://github.com/OpenBMB/MiniCPM

    Developer: OpenBMB
    Released: May-September 2026, rolling family
    Variants: MiniCPM5-2B (September 2026) / MiniCPM5-1B (May 2026)
    Parameters: 2B: 2.52B, 1.98B non-embedding / 1B: 1.08B, 0.68B non-embedding
    Context: 131,072 tokens
    Architecture: dense LlamaForCausalLM, GQA with 16 Q and 2 KV heads; 2B: 42 layers / 1B: 24 layers
    License: Apache 2.0
    Modalities: Text
    Runs on: Smartphone, Laptop
    Formats: safetensors BF16, GGUF (F16, Q8_0, Q4_K_M), MLX 4-bit, GPTQ 4-bit (2B)
    On disk: 2B: 1.56GB Q4_K_M, 5.03GB BF16 / 1B: 0.69GB Q4_K_M, 2.16GB BF16

    AI & Software llm foundation models

  • Embodied Foundation Models
    montezM montez

    DeepCybo PhysBrain 1.5

    Hugging Face: https://huggingface.co/DeepCybo/PhysBrain1.5-8B, https://huggingface.co/DeepCybo/PhysBrain1.5-2B
    GitHub: https://github.com/DeepCybo-PhysAI/PhysBrain-1.5
    Website: https://deepcybo-physai.github.io/PhysBrain-1.5/
    Technical report: https://arxiv.org/abs/2609.14973

    Developer: DeepCybo, Zhongguancun Academy, Zhongguancun Institute of Artificial Intelligence
    Released: September 2026
    Variants: 2B, 8B
    Parameters: 2.16B (2B checkpoint), 8.90B (8B checkpoint)
    Context: 262,144 tokens
    Architecture: Qwen3-VL backbone extended with action and visual-state tokens; one shared autoregressive transformer, single next-token objective, no task-specific heads; ActionPiece action tokens with one action codebook across robot setups
    Modalities: Text + Image + Video + Action history in / Text + Spatial grounding + Action chunk + Future-state image, depth map, robot mask out
    Runs on: 2B: Laptop, Desktop GPU / 8B: Desktop GPU
    Formats: safetensors
    On disk: 4.32GB (2B checkpoint), 17.81GB (8B checkpoint)

    AI & Software embodied foundation models robotics

  • Local LMs
    montezM montez

    SparkLLM Spark-X2.5

    Hugging Face: https://huggingface.co/XHToken/Spark-X2.5-4B, https://huggingface.co/XHToken/Spark-X2.5-1.7B
    GitHub: https://github.com/XHToken/Spark-X2.5
    Website: https://dev.to/sparkllm/spark-x25-4b-17b-the-only-on-device-models-with-native-1m-token-context-now-open-source-d9o

    Developer: SparkLLM
    Released: September 2026
    Variants: 4B / 1.7B
    Parameters: 4B / 1.7B
    Context: 1,048,576 tokens
    Architecture: dense; hybrid attention with 1 full-attention layer to 3 sliding-window layers, window 512; 4B: 36 layers, GQA 16 Q and 4 KV heads / 1.7B: 28 layers, GQA 8 Q and 2 KV heads
    License: Apache 2.0
    Modalities: Text
    Runs on: Smartphone, Laptop, Edge device
    Formats: safetensors BF16, FP8, INT8, GGUF (Q4_K_M, Q8_0, F16)
    On disk: 4B: 2.60GB Q4_K_M / 1.7B: 1.11GB Q4_K_M

    AI & Software llm foundation models

  • Embodied Foundation Models
    montezM montez

    Meta FAIR RoboJEPA

    GitHub: https://github.com/facebookresearch/robo_jepa
    Website: https://robojepa.github.io
    Technical report: https://arxiv.org/abs/2610.10515

    Developer: Meta FAIR
    Released: October 2026
    Variants: 22M / 50M / 100M / 300M / 1B / 2B / 4B / 8B / 8B DROID 720p (3 views)
    Parameters: 22M to 8B predictors on a frozen V-JEPA 2.1 ViT-G/384 encoder
    Architecture: Action-conditioned JEPA latent world model: a transformer predictor over frozen V-JEPA 2.1 features, with latent CEM/MPC planning toward a goal image and an optional diffusion video decoder
    License: CC BY-NC-SA 4.0
    Modalities: Camera video + action sequence in / predicted future latent states out (optional decoded video); no action head, actions come from latent MPC planning
    Runs on: NVIDIA GPU (CUDA 12.6), one GPU per model size for evaluation; no VRAM requirement published
    Formats: PyTorch (.pth.tar) from dl.fbaipublicfiles.com
    On disk: 8B 29.99GB, 1B 4.07GB, 300M 1.24GB (300k-step checkpoints), plus V-JEPA 2.1 ViT-G/384 encoder 30.24GB

    AI & Software embodied foundation models robotics

  • Local LMs
    montezM montez

    China Telecom Xing4.0

    Hugging Face: https://huggingface.co/XingChen-AGI/Xing4.0-29B-A4B, https://huggingface.co/XingChen-AGI/Xing4.0-29B-A4B-GGUF, https://huggingface.co/XingChen-AGI/Xing4.0-29B-A4B-FP8
    GitHub: https://github.com/XingChen-AGI/Xing4.0-29B-A4B

    Developer: China Telecom Artificial Intelligence Technology Co., Ltd.
    Released: September 2026
    Parameters: 29B total, 4B active
    Context: 262,144 tokens, extensible to 512K
    Architecture: MoE with mHC, MLA attention and MTP; 40 layers, 64 routed experts with 4 active plus 1 shared; trained on Ascend NPUs with MindSpore
    License: Apache 2.0
    Modalities: Text
    Runs on: Desktop GPU
    Formats: safetensors BF16, FP8, GGUF IQ4_NL
    On disk: 20.1GB IQ4_NL / 33.17GB FP8 / 62.43GB BF16

    AI & Software llm foundation models

  • Embodied Foundation Models
    montezM montez

    Microsoft Rho

    Hugging Face: https://huggingface.co/microsoft/rho-base, https://huggingface.co/collections/microsoft/rho
    GitHub: https://github.com/microsoft/rhobotics
    Website: https://microsoft.github.io/rhobotics/
    Technical report: https://arxiv.org/abs/2609.38164

    Developer: Microsoft Research
    Released: September 2026
    Variants: rho-base / rho-yam-box / rho-ur-ai-trainer / rho-fr3-duo / rho-libero / rho-roboeval
    Parameters: 5B (Phi-Phy VLM 4.68B / flow-matching action expert 542M)
    Architecture: Vision-language-action model: a physically grounded Phi-family VLM (Phi-Phy) with a 12-block flow-matching action expert that cross-attends to a VLM decoder layer, midtrained per embodiment
    License: MIT
    Modalities: Multi-camera images + language instruction + robot state in / action chunk out (up to 50 steps in pretraining)
    Runs on: NVIDIA GPU (Linux, CUDA; FlashAttention 2 by default); no VRAM requirement published
    Formats: safetensors (BF16, 3 shards)
    On disk: rho-base 10.50GB

    AI & Software embodied foundation models robotics

  • Local LMs
    montezM montez

    AI Singapore Nemotron-SEA-LION v4.8 30B-A3B

    Hugging Face: https://huggingface.co/aisingapore/Nemotron-SEA-LION-v4.8-30B-A3B, https://huggingface.co/aisingapore/Nemotron-SEA-LION-v4.8-30B-A3B-GGUF, https://huggingface.co/aisingapore/Nemotron-SEA-LION-v4.8-30B-A3B-FP8, https://huggingface.co/aisingapore/Nemotron-SEA-LION-v4.8-30B-A3B-NVFP4
    Website: https://sea-lion.ai/blog/uplifting-ai-in-southeast-asia-sea-announcing-nemotron-sea-lion-v4-8-in-collaboration-with-nvidia/
    Technical report: https://arxiv.org/abs/2609.18310

    Developer: AI Singapore, with NVIDIA
    Released: September 2026
    Parameters: 30B total, 3B active
    Context: 262,144 tokens
    Architecture: Mamba2-Transformer hybrid MoE; NVIDIA Nemotron 3 Nano 30B-A3B base, continued pretraining on 150B tokens, then SFT and on-policy distillation; English plus 7 Southeast Asian languages
    License: MIT
    Modalities: Text
    Runs on: Laptop, 32GB+ unified memory
    Formats: safetensors BF16, FP8, NVFP4, GGUF (Q4_K_M, Q6_K, Q8_0, F16)
    On disk: 25.43GB Q4_K_M / 22.94GB NVFP4 / 34.96GB FP8 / 65.83GB BF16

    AI & Software llm foundation models

Member List

montezM montez
  • Login

  • Don't have an account? Register

  • Login or register to search.
  • First post
    Last post
  • 0
    • Categories
    • Recent
    • Tags
    • Popular
    • Users