Local LMs
-
FireRedTeam FireRedAudio
Hugging Face: https://huggingface.co/FireRedTeam/FireRedAudio
GitHub: https://github.com/FireRedTeam/FireRedAudio
Website: https://fireredteam.github.io/demos/fireredaudio/
Technical report: https://arxiv.org/abs/2608.24168Developer: FireRedTeam
Released: August 2026
Parameters: 9B
Architecture: shared 9B LLM; decoupled Audio Encoder for understanding and RedAE-Patch plus flow-matching DiT pathway for generation
License: Apache 2.0
Modalities: Text + Audio
Runs on: Desktop GPU, NVIDIA only
Formats: safetensors, PyTorch .pt
On disk: FireRedAudio: 21.23GB safetensors / RedAE decoder: 8.40GB PyTorch .pt / 29.63GB total, estimated -
Cohere Labs Tiny Aya L2-Thinker
Hugging Face: https://huggingface.co/CohereLabs/tiny-aya-l2-thinker
Technical report: https://arxiv.org/abs/2609.10445Developer: Cohere Labs
Released: September 2026
Parameters: 3.35B
Architecture: Cohere2 decoder-only transformer; supervised fine-tuning for in-language reasoning
License: CC-BY-NC-4.0
Modalities: Text
Runs on: Laptop
Formats: safetensors
On disk: 6.70GB BF16 safetensors, estimated -
InternLM Intern Lumina U2
Hugging Face: https://huggingface.co/internlm/InternLumina-U2
GitHub: https://github.com/InternLM/InternLumina-U2
Website: https://internlm.github.io/InternLumina-U2/Developer: Shanghai AI Laboratory / InternLM
Released: September 2026
Parameters: 16B total, 1B active
Architecture: LLaDA-2.0 MoE diffusion LLM backbone; 8-codebook fully-discrete AToken visual representation; spatial-parallel denoising plus codebook-depth autoregressive head
License: Apache 2.0
Modalities: Text + Image + Video + 3D
Runs on: Huawei Ascend NPU only
Formats: safetensors
On disk: 33.81GB Ascend safetensors, estimated -
Microsoft FrogNano
Hugging Face: https://huggingface.co/microsoft/FrogNano-4B-2609
GitHub: https://github.com/microsoft/FrogNano
Technical report: https://arxiv.org/abs/2609.07925Developer: Microsoft
Released: September 2026
Parameters: 4.66B
Context: 131K tokens in the evaluated configuration
Architecture: Qwen3.5-4B base; dense, 32 layers, hybrid Gated DeltaNet + gated attention; reinforcement-learning post-training on about 1,500 synthetic software-engineering tasks; repository-level coding agent run through the five-tool Leaf harness
License: MIT
Modalities: Text
Runs on: Laptop, 16GB+ unified memory
Formats: BF16 safetensors
On disk: 9.32GB BF16 -
H Company Holo4
Hugging Face: https://huggingface.co/Hcompany/Holo4-27B, https://huggingface.co/Hcompany/Holo4-35B-A3B, https://huggingface.co/Hcompany/Holotron4-30B-A3B
Website: https://hcompany.ai/newsroom/holo4Developer: H Company
Released: September 2026
Variants: Holo4-27B / Holo4-35B-A3B / Holotron4-30B-A3B
Parameters: 27B dense / 35B total, 3B active / 30B total, 3B active
Context: 262,144 tokens
Architecture: computer-use agent VLMs; 27B: Qwen3.8 base, dense, 64 layers / 35B-A3B: Qwen3.6 base, MoE, 40 layers, 256 experts, 8 routed active / Holotron4: NVIDIA Nemotron 3 Nano Omni base, hybrid Mamba + MoE + attention, 52 layers, 128 routed experts, 6 active; Holo4: supervised fine-tuning plus two merged reinforcement-learning LoRA experts
License: 27B: CC-BY-NC-4.0 / 35B-A3B: Apache 2.0 / Holotron4: NVIDIA Open Model Agreement
Modalities: Text + Image
Runs on: 27B: Laptop, 32GB+ unified memory / 35B-A3B: Laptop, 32GB+ unified memory / Holotron4: Workstation GPU, 48GB+ memory
Formats: 27B and 35B-A3B: BF16 safetensors, FP8, NVFP4, Q4_K_M GGUF with mmproj / Holotron4: BF16 safetensors, FP8
On disk: 27B: 16.88GB Q4_K_M GGUF, 54.71GB BF16 / 35B-A3B: 21.30GB Q4_K_M GGUF, 70.21GB BF16 / Holotron4: 35.20GB FP8, 66.03GB BF16 -
JetBrains Mellum2.1
Hugging Face: https://huggingface.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking, https://huggingface.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking-GGUF, https://huggingface.co/JetBrains/Mellum2-12B-A2.5B-Base, https://huggingface.co/JetBrains/Mellum2-12B-A2.5B-Instruct, https://huggingface.co/JetBrains/Mellum2-12B-A2.5B-Thinking
Website: https://blog.jetbrains.com/ai/2026/10/mellum2-1-gets-to-work-a-fast-open-model-for-coding-agents/
Technical report: https://arxiv.org/abs/2605.31268Developer: JetBrains
Released: October 2026
Variants: Mellum2.1 Thinking / Mellum2 Base, Instruct, Thinking (June 2026)
Parameters: 12B total, 2.5B active
Context: 131,072 tokens
Architecture: MoE, 28 layers, 64 experts with 8 active, GQA with 32 Q and 4 KV heads, sliding window 1,024 on 3 of every 4 layers; Mellum2.1 keeps the Mellum2 architecture and adds reinforcement learning post-training
License: Apache 2.0
Modalities: Text
Runs on: Laptop, 16GB+ unified memory
Formats: safetensors BF16, GGUF (BF16, Q8_0, Q6_K, Q4_K_M, MXFP4_MOE)
On disk: Mellum2.1 Thinking: 8.07GB Q4_K_M, 24.31GB BF16 -
PrismML Bonsai 2 27B
Hugging Face: https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-mlx-2bit, https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf
GitHub: https://github.com/PrismML-Eng/Bonsai-demo
Website: https://prismml.com/news/bonsai-2-27b
White paper: https://github.com/PrismML-Eng/Bonsai-demo/blob/main/bonsai-2-27b-whitepaper.pdfDeveloper: PrismML
Released: September 2026
Parameters: 27.36B total, 24.35B language + 2.54B embedding and LM head + 0.47B vision tower
Context: 262,144 tokens
Architecture: Qwen3.8-27B base; 64 blocks, 24Q/4KV heads, hybrid attention with about 75% linear and 25% full attention; ternary {-1, 0, +1} weights with FP16 group-wise scales at group size 128, 1.76 effective bits per weight
License: Apache 2.0
Modalities: Text + Image
Runs on: Laptop, Desktop GPU; Apple Silicon or NVIDIA GPU only
Formats: MLX 2-bit ternary, GGUF PTQ1_0, GGUF PQ2_0, optional mmproj vision pack
On disk: 5.95GB PTQ1_0 GGUF, 7.21GB PQ2_0 GGUF, 8.60GB MLX / 0.63GB mmproj vision pack -
Swiss AI Apertus 1.5
Hugging Face: https://huggingface.co/swiss-ai/Apertus-v1.5-8B, https://huggingface.co/swiss-ai/Apertus-v1.5-70B
Website: https://www.apertus-ai.org/articles/2026-07-apertus-1-5
Docs: https://www.apertus-ai.org/docsDeveloper: Swiss AI Initiative (EPFL, ETH Zurich, CSCS)
Released: July 2026
Variants: 8B / 70B
Parameters: 8B / 70B
Context: 262,144 tokens
Architecture: decoder-only transformer with xIELU activation, trained with the AdEMAMix optimizer; continued pretraining of Apertus 1.0 on 4T added tokens (8B) and 2T added tokens (70B); optional thinking mode and tool calling
License: Apache 2.0 with Acceptable Use Policy
Modalities: Text + Image + Audio (audio experimental), text output
Runs on: 8B: Laptop, Desktop GPU / 70B: Server-class hardware
Formats: safetensors
On disk: 8B: 18.40GB safetensors / 70B: 144.60GB safetensors -
Edge0
Hugging Face: https://huggingface.co/Edge0/Edge0-35B-A3B-preview, https://huggingface.co/Edge0/Edge0-8B-A1B-preview
GitHub: https://github.com/Edge0-AI/Edge0
Technical report: https://arxiv.org/abs/2609.18063Developer: Edge0 AI
Released: September 2026
Variants: Edge0-35B-A3B-preview / Edge0-8B-A1B-preview
Parameters: 35B total, 3B active / 8B total, 1B active
Context: 262,144 tokens / 131,072 tokens
Architecture: early preview release; sparse MoE; 35B-A3B: Qwen3.6-35B-A3B base, 40 layers, 256 experts, 8 per token / 8B-A1B: Ling 3.0 tiny base, 24 layers, 128 experts, 8 per token; int4 checkpoint plus Recover-LoRA adapters and prerouter heads, with experts streamed from SSD by the edge0 framework
License: Apache 2.0
Modalities: Text
Runs on: 35B-A3B: Smartphone, 12GB+ RAM on Android / 8B-A1B: Smartphone, 8GB+ RAM on Android; edge0 engines only (iOS, macOS, Android, Windows)
Formats: MLX 4-bit safetensors, LoRA and prerouter adapter safetensors
On disk: 35B-A3B: 19.51GB int4 checkpoint plus 0.18GB adapters / 8B-A1B: 4.51GB int4 checkpoint plus 0.06GB adapters -
IFM K2 Horizon
Hugging Face: https://huggingface.co/IFM/K2-Horizon-0.9B, https://huggingface.co/IFM/K2-Horizon-3.7B, https://huggingface.co/IFM/K2-Horizon-7B, https://huggingface.co/IFM/K2-Horizon-32B, https://huggingface.co/IFM/K2-Horizon-MoVA-36B-A4B
GitHub: https://github.com/ifm-ai/xllm
Website: https://ifm.ai/k2/Developer: Institute of Foundation Models (IFM), MBZUAI
Released: September 2026
Variants: 0.9B / 3.7B / 7B / 32B / MoVA-36B-A4B
Parameters: 0.9B / 3.7B / 7B / 32B / 36B total, 4B active
Context: 0.9B: 131,072 tokens / 3.7B, 7B, 32B and MoVA-36B-A4B: 524,288 tokens
Architecture: 0.9B, 3.7B, 7B, 32B: dense decoder-only, GQA with 8 KV heads, 28 / 36 / 36 / 64 layers; MoVA-36B-A4B: MoE with Mixture-of-Values attention, 48 layers, 100 FFN experts with 8 per token plus 1 shared, 64 value experts with 4 per token
License: Apache 2.0
Modalities: Text
Runs on: 0.9B: Smartwatch, Smart glasses, Smartphone / 3.7B and 7B: Smartphone, Laptop / 32B and MoVA-36B-A4B: Laptop, 32GB+ unified memory
Formats: safetensors BF16, GGUF (Q4_K_M, Q5_0, Q5_K_M, Q6_K, Q8_0, BF16), FP8 (7B, 32B, MoVA-36B-A4B), NVFP4 (32B)
On disk: 0.9B: 0.67GB Q4_K_M / 3.7B: 3.16GB Q4_K_M / 7B: 5.59GB Q4_K_M / 32B: 21.08GB Q4_K_M / MoVA-36B-A4B: 22.37GB Q4_K_M -
AI Singapore Nemotron-SEA-LION v4.8 30B-A3B
Hugging Face: https://huggingface.co/aisingapore/Nemotron-SEA-LION-v4.8-30B-A3B, https://huggingface.co/aisingapore/Nemotron-SEA-LION-v4.8-30B-A3B-GGUF, https://huggingface.co/aisingapore/Nemotron-SEA-LION-v4.8-30B-A3B-FP8, https://huggingface.co/aisingapore/Nemotron-SEA-LION-v4.8-30B-A3B-NVFP4
Website: https://sea-lion.ai/blog/uplifting-ai-in-southeast-asia-sea-announcing-nemotron-sea-lion-v4-8-in-collaboration-with-nvidia/
Technical report: https://arxiv.org/abs/2609.18310Developer: AI Singapore, with NVIDIA
Released: September 2026
Parameters: 30B total, 3B active
Context: 262,144 tokens
Architecture: Mamba2-Transformer hybrid MoE; NVIDIA Nemotron 3 Nano 30B-A3B base, continued pretraining on 150B tokens, then SFT and on-policy distillation; English plus 7 Southeast Asian languages
License: MIT
Modalities: Text
Runs on: Laptop, 32GB+ unified memory
Formats: safetensors BF16, FP8, NVFP4, GGUF (Q4_K_M, Q6_K, Q8_0, F16)
On disk: 25.43GB Q4_K_M / 22.94GB NVFP4 / 34.96GB FP8 / 65.83GB BF16 -
China Telecom Xing4.0
Hugging Face: https://huggingface.co/XingChen-AGI/Xing4.0-29B-A4B, https://huggingface.co/XingChen-AGI/Xing4.0-29B-A4B-GGUF, https://huggingface.co/XingChen-AGI/Xing4.0-29B-A4B-FP8
GitHub: https://github.com/XingChen-AGI/Xing4.0-29B-A4BDeveloper: China Telecom Artificial Intelligence Technology Co., Ltd.
Released: September 2026
Parameters: 29B total, 4B active
Context: 262,144 tokens, extensible to 512K
Architecture: MoE with mHC, MLA attention and MTP; 40 layers, 64 routed experts with 4 active plus 1 shared; trained on Ascend NPUs with MindSpore
License: Apache 2.0
Modalities: Text
Runs on: Desktop GPU
Formats: safetensors BF16, FP8, GGUF IQ4_NL
On disk: 20.1GB IQ4_NL / 33.17GB FP8 / 62.43GB BF16 -
SparkLLM Spark-X2.5
Hugging Face: https://huggingface.co/XHToken/Spark-X2.5-4B, https://huggingface.co/XHToken/Spark-X2.5-1.7B
GitHub: https://github.com/XHToken/Spark-X2.5
Website: https://dev.to/sparkllm/spark-x25-4b-17b-the-only-on-device-models-with-native-1m-token-context-now-open-source-d9oDeveloper: SparkLLM
Released: September 2026
Variants: 4B / 1.7B
Parameters: 4B / 1.7B
Context: 1,048,576 tokens
Architecture: dense; hybrid attention with 1 full-attention layer to 3 sliding-window layers, window 512; 4B: 36 layers, GQA 16 Q and 4 KV heads / 1.7B: 28 layers, GQA 8 Q and 2 KV heads
License: Apache 2.0
Modalities: Text
Runs on: Smartphone, Laptop, Edge device
Formats: safetensors BF16, FP8, INT8, GGUF (Q4_K_M, Q8_0, F16)
On disk: 4B: 2.60GB Q4_K_M / 1.7B: 1.11GB Q4_K_M -
OpenBMB MiniCPM5
Hugging Face: https://huggingface.co/openbmb/MiniCPM5-2B, https://huggingface.co/openbmb/MiniCPM5-1B
GitHub: https://github.com/OpenBMB/MiniCPMDeveloper: OpenBMB
Released: May-September 2026, rolling family
Variants: MiniCPM5-2B (September 2026) / MiniCPM5-1B (May 2026)
Parameters: 2B: 2.52B, 1.98B non-embedding / 1B: 1.08B, 0.68B non-embedding
Context: 131,072 tokens
Architecture: dense LlamaForCausalLM, GQA with 16 Q and 2 KV heads; 2B: 42 layers / 1B: 24 layers
License: Apache 2.0
Modalities: Text
Runs on: Smartphone, Laptop
Formats: safetensors BF16, GGUF (F16, Q8_0, Q4_K_M), MLX 4-bit, GPTQ 4-bit (2B)
On disk: 2B: 1.56GB Q4_K_M, 5.03GB BF16 / 1B: 0.69GB Q4_K_M, 2.16GB BF16 -
NII LLM-jp-4
Hugging Face: https://huggingface.co/llm-jp
GitHub: https://github.com/llm-jp/llm-jp-4-cookbook
Website: https://llm-jp.nii.ac.jp/blog/llm-jp-4-1/Developer: National Institute of Informatics, Research and Development Center for Large Language Models (LLMC), LLM-jp
Released: April-September 2026, rolling family
Variants: LLM-jp-4 8B and 32B-A3B (April 2026) / LLM-jp-4 33B (August 2026) / LLM-jp-4-VL 9B (September 2026) / LLM-jp-4.1 8B, 32B-A3B, 33B Thinking (September 2026)
Parameters: 8B: 8.59B / 32B-A3B: 32.14B total, 3.83B active / 33B: 33.22B / VL 9B: 8.6B language model + 0.4B vision encoder
Context: 65,536 tokens, text models
Architecture: 8B and 33B: dense Llama architecture, 32 and 64 layers / 32B-A3B: Qwen3-MoE architecture, 32 layers, 128 routed experts with 8 active / VL 9B: llm-jp-4-8b-thinking, SigLIP 2 So400m vision encoder, 2-layer MLP projector
License: Apache 2.0
Modalities: 8B, 32B-A3B, 33B: Text / VL 9B: Text + Image
Runs on: 8B: Laptop / 32B-A3B and 33B: Laptop, 32GB+ unified memory / VL 9B: Laptop, 24GB+ unified memory
Formats: safetensors BF16, GGUF (BF16, Q4_K_M; LLM-jp-4.1 text models)
On disk: LLM-jp-4.1 8B: 5.50GB Q4_K_M / 32B-A3B: 21.52GB Q4_K_M / 33B: 20.41GB Q4_K_M / VL 9B: 18.11GB BF16 -
Qwen-Image-2.1
Hugging Face: https://huggingface.co/Qwen/Qwen-Image-2.1, https://huggingface.co/Qwen/Qwen-Image-2.1-Turbo
GitHub: https://github.com/QwenLM/Qwen-Image-2.1
Website: https://qwen.ai/blog?id=qwen-image-2.1Developer: Alibaba Cloud, Qwen team
Released: September 2026
Variants: Qwen-Image-2.1 (40 default denoising steps) / Qwen-Image-2.1-Turbo (8 denoising steps)
Parameters: 7B visual generation component, plus Qwen3-VL 8B text encoder
Resolution: native 2K, 2048x2048 default; recommended sizes include 2752x1536 and 1536x2752
Architecture: single-stream DiT, 32 layers, block-causal attention with prefix KV cache reuse; Qwen3-VL 8B text encoder; 64-channel RGBA VAE with 16x spatial compression; flow matching with Euler scheduler
License: Qwen Research License Agreement
Modalities: Text-to-Image + Image editing (up to 10 reference images), native RGBA transparency
Runs on: Desktop GPU, 24GB+ VRAM with model CPU offload
Formats: BF16 safetensors, Diffusers pipeline
On disk: Qwen-Image-2.1: 33.13GB (14.23GB transformer, 17.53GB text encoder, 1.35GB VAE) / Turbo: 32.49GB (14.23GB transformer, 17.53GB text encoder, 0.68GB VAE)