Local LMs
-
IBM Granite Vision 4.1
Hugging Face: https://huggingface.co/ibm-granite/granite-vision-4.1-4b, https://huggingface.co/ibm-granite/granite-vision-4.1-4b-GGUF
GitHub: https://github.com/ibm-granite/granite-vision-models
Website: https://www.ibm.com/granite/docs/models/vision
Announcement: https://research.ibm.com/blog/granite-4-1-ai-foundation-modelsDeveloper: IBM
Released: April 2026
Parameters: 4B total, Granite 4.1 3B language model plus vision encoder and projectors
Context: 131,072 tokens
Architecture: SigLIP2 so400m patch16-384 vision encoder over 384x384 image tiles, windowed Q-Former projectors compressing each 4x4 patch window to 2x2 tokens, and a Granite 4.1 3B language model with rank-256 LoRA across all self-attention projections
License: Apache 2.0
Modalities: Text + Image
Runs on: Smartphone, Laptop, Edge device
Formats: safetensors, GGUF Q4_K_M, Q5_K_M, Q6_K, Q8_0, bf16, with f16 mmproj
On disk: 2.10GB Q4_K_M plus 1.16GB f16 mmproj / 6.81GB bf16 -
Microsoft Mage-VL
Hugging Face: https://huggingface.co/microsoft/Mage-VL
GitHub: https://github.com/microsoft/Mage
Website: https://microsoft.github.io/Mage/vl/
Technical report: https://arxiv.org/abs/2607.24904Developer: Microsoft Mage Team
Released: July 2026
Parameters: 4B
Context: 262,144 tokens
Architecture: Mage-ViT codec-native visual encoder trained from scratch, 24 layers, feeding a two-layer MLP projector into a Qwen3-4B-Instruct-2507 causal decoder; separate cognition gate for proactive streaming
License: Apache 2.0
Modalities: Text + Image + Video
Runs on: Laptop, 12GB+ memory at BF16, estimated / Desktop GPU
Formats: safetensors, bundled streaming gate and neural codec
On disk: 9.48GB safetensors -
Cohere Labs North Micro Vision
Hugging Face: https://huggingface.co/CohereLabs/North-Micro-Vision-Instruct
Technical report: https://huggingface.co/blog/CohereLabs/meet-north-micro-vision-instructDeveloper: Cohere Labs
Released: August 2026
Parameters: 2.4B total, 2B language model + 400M vision encoder
Context: 128,000 tokens, multimodal validated to 8,192
Architecture: custom native-resolution vision encoder with DeepStack patch embeddings injected into early decoder layers, projector, and the Command A+ style North Micro LLM: three sliding-window attention layers with RoPE plus one global layer without positional embeddings
License: Apache 2.0
Modalities: Text + Image
Runs on: Smartphone, Laptop, Edge device, with quantization
Formats: safetensors BF16, MLX 4-bit and 8-bit community conversions
On disk: 4.97GB BF16 safetensors -
Cactus Compute Needle 2
Hugging Face: https://huggingface.co/Cactus-Compute/needle2
GitHub: https://github.com/cactus-compute/needle
X: https://x.com/cactuscompute/status/2086865960669983035
Website: https://cactuscompute.com/needleDeveloper: Cactus Compute
Released: August 2026
Parameters: 45M
Context: 2,048 tokens
Architecture: Simple Attention Network, 27 layers, hidden 512, 8Q/4KV GQA, Hadamard MLP, engram sites, CQ2 quantization at 2.2 effective bits; byte-level grammar-constrained decoding and a tool-retrieval head
License: Apache 2.0
Modalities: Text
Runs on: Smartphone, Headset, Edge device, Microcontroller
Formats: cact single binary; ARM64, x86-64, ARMv7, RISC-V and WebAssembly builds
On disk: 13.7MB cact -
Syzygy Mach-1 Additive 35B
Hugging Face: https://huggingface.co/SyzygyResearch/Mach-1-Additive-35B
X: https://x.com/syzygyeng/status/2084350792841195992
Website: https://withsyzygy.com/mach-1
Docs: https://withsyzygy.com/docs/machDeveloper: Syzygy Research
Released: August 2026
Parameters: 35B total, 8 of 256 experts active
Context: 262,144 tokens
Architecture: Qwen3.5 MoE topology, 40 layers, 256 experts, 8 active, hybrid linear attention with full attention every 4th layer; additive 1.7-bit weights with no weight multiplication
License: Apache 2.0
Modalities: Text
Runs on: Laptop, 16GB+ unified memory; Apple Silicon only
Formats: packed 1.7-bit safetensors, MLX
On disk: 7.0GB -
OpenMOSS MOSS-VL
Hugging Face: https://huggingface.co/OpenMOSS-Team/MOSS-VL-Instruct-0708-FP8, https://huggingface.co/OpenMOSS-Team/MOSS-VL-Realtime-FP8
GitHub: https://github.com/OpenMOSS/MOSS-VL
Website: https://openmoss.ai/MOSS-VL/
Technical report: https://arxiv.org/abs/2606.07639Developer: OpenMOSS Team
Released: August 2026
Variants: Instruct-0708, Realtime
Parameters: 11B
Context: 262,144 tokens
Architecture: unified cross-attention multimodal model, 48 language layers with 12 cross-attention layers, XRoPE 3D spatiotemporal positions, absolute frame timestamps for streaming video
License: Apache 2.0
Modalities: Text + Image + Video
Runs on: Desktop GPU; Instruct: 24GB VRAM / Realtime: 26GB+ VRAM; NVIDIA only
Formats: FP8 compressed-tensors with BF16 cross-attention and vision, HQQ INT8 KV cache
On disk: 15.73GB FP8 safetensors -
IBM Granite 4.2
Hugging Face: https://huggingface.co/ibm-granite/granite-4.2-3b, https://huggingface.co/ibm-granite/granite-4.2-8b, https://huggingface.co/ibm-granite/granite-4.2-30b
GitHub: https://github.com/ibm-granite/granite-4.2-language-models
Website: https://www.ibm.com/granite/docs/models/granite4-2Developer: IBM
Released: August 2026
Variants: 3B / 8B / 30B
Parameters: 3B / 8B / 30B
Context: 3B and 8B: 131,072 tokens / 30B: 131,072 tokens native, 512,000 tokens extended
Architecture: dense decoder-only transformer; GQA, RoPE, SwiGLU and RMSNorm; native reasoning modes and tool calling
License: Apache 2.0
Modalities: Text
Formats: safetensors
On disk: 3B: 7.32GB BF16 / 8B: 17.58GB BF16 / 30B: 58.55GB BF16 -
FireRedTeam FireRedAudio
Hugging Face: https://huggingface.co/FireRedTeam/FireRedAudio
GitHub: https://github.com/FireRedTeam/FireRedAudio
Website: https://fireredteam.github.io/demos/fireredaudio/
Technical report: https://arxiv.org/abs/2608.24168Developer: FireRedTeam
Released: August 2026
Parameters: 9B
Architecture: shared 9B LLM; decoupled Audio Encoder for understanding and RedAE-Patch plus flow-matching DiT pathway for generation
License: Apache 2.0
Modalities: Text + Audio
Runs on: Desktop GPU, NVIDIA only
Formats: safetensors, PyTorch .pt
On disk: FireRedAudio: 21.23GB safetensors / RedAE decoder: 8.40GB PyTorch .pt / 29.63GB total, estimated -
Cohere Labs Tiny Aya L2-Thinker
Hugging Face: https://huggingface.co/CohereLabs/tiny-aya-l2-thinker
Technical report: https://arxiv.org/abs/2609.10445Developer: Cohere Labs
Released: September 2026
Parameters: 3.35B
Architecture: Cohere2 decoder-only transformer; supervised fine-tuning for in-language reasoning
License: CC-BY-NC-4.0
Modalities: Text
Runs on: Laptop
Formats: safetensors
On disk: 6.70GB BF16 safetensors, estimated -
InternLM Intern Lumina U2
Hugging Face: https://huggingface.co/internlm/InternLumina-U2
GitHub: https://github.com/InternLM/InternLumina-U2
Website: https://internlm.github.io/InternLumina-U2/Developer: Shanghai AI Laboratory / InternLM
Released: September 2026
Parameters: 16B total, 1B active
Architecture: LLaDA-2.0 MoE diffusion LLM backbone; 8-codebook fully-discrete AToken visual representation; spatial-parallel denoising plus codebook-depth autoregressive head
License: Apache 2.0
Modalities: Text + Image + Video + 3D
Runs on: Huawei Ascend NPU only
Formats: safetensors
On disk: 33.81GB Ascend safetensors, estimated -
Microsoft FrogNano
Hugging Face: https://huggingface.co/microsoft/FrogNano-4B-2609
GitHub: https://github.com/microsoft/FrogNano
Technical report: https://arxiv.org/abs/2609.07925Developer: Microsoft
Released: September 2026
Parameters: 4.66B
Context: 131K tokens in the evaluated configuration
Architecture: Qwen3.5-4B base; dense, 32 layers, hybrid Gated DeltaNet + gated attention; reinforcement-learning post-training on about 1,500 synthetic software-engineering tasks; repository-level coding agent run through the five-tool Leaf harness
License: MIT
Modalities: Text
Runs on: Laptop, 16GB+ unified memory
Formats: BF16 safetensors
On disk: 9.32GB BF16 -
H Company Holo4
Hugging Face: https://huggingface.co/Hcompany/Holo4-27B, https://huggingface.co/Hcompany/Holo4-35B-A3B, https://huggingface.co/Hcompany/Holotron4-30B-A3B
Website: https://hcompany.ai/newsroom/holo4Developer: H Company
Released: September 2026
Variants: Holo4-27B / Holo4-35B-A3B / Holotron4-30B-A3B
Parameters: 27B dense / 35B total, 3B active / 30B total, 3B active
Context: 262,144 tokens
Architecture: computer-use agent VLMs; 27B: Qwen3.8 base, dense, 64 layers / 35B-A3B: Qwen3.6 base, MoE, 40 layers, 256 experts, 8 routed active / Holotron4: NVIDIA Nemotron 3 Nano Omni base, hybrid Mamba + MoE + attention, 52 layers, 128 routed experts, 6 active; Holo4: supervised fine-tuning plus two merged reinforcement-learning LoRA experts
License: 27B: CC-BY-NC-4.0 / 35B-A3B: Apache 2.0 / Holotron4: NVIDIA Open Model Agreement
Modalities: Text + Image
Runs on: 27B: Laptop, 32GB+ unified memory / 35B-A3B: Laptop, 32GB+ unified memory / Holotron4: Workstation GPU, 48GB+ memory
Formats: 27B and 35B-A3B: BF16 safetensors, FP8, NVFP4, Q4_K_M GGUF with mmproj / Holotron4: BF16 safetensors, FP8
On disk: 27B: 16.88GB Q4_K_M GGUF, 54.71GB BF16 / 35B-A3B: 21.30GB Q4_K_M GGUF, 70.21GB BF16 / Holotron4: 35.20GB FP8, 66.03GB BF16 -
JetBrains Mellum2.1
Hugging Face: https://huggingface.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking, https://huggingface.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking-GGUF, https://huggingface.co/JetBrains/Mellum2-12B-A2.5B-Base, https://huggingface.co/JetBrains/Mellum2-12B-A2.5B-Instruct, https://huggingface.co/JetBrains/Mellum2-12B-A2.5B-Thinking
Website: https://blog.jetbrains.com/ai/2026/10/mellum2-1-gets-to-work-a-fast-open-model-for-coding-agents/
Technical report: https://arxiv.org/abs/2605.31268Developer: JetBrains
Released: October 2026
Variants: Mellum2.1 Thinking / Mellum2 Base, Instruct, Thinking (June 2026)
Parameters: 12B total, 2.5B active
Context: 131,072 tokens
Architecture: MoE, 28 layers, 64 experts with 8 active, GQA with 32 Q and 4 KV heads, sliding window 1,024 on 3 of every 4 layers; Mellum2.1 keeps the Mellum2 architecture and adds reinforcement learning post-training
License: Apache 2.0
Modalities: Text
Runs on: Laptop, 16GB+ unified memory
Formats: safetensors BF16, GGUF (BF16, Q8_0, Q6_K, Q4_K_M, MXFP4_MOE)
On disk: Mellum2.1 Thinking: 8.07GB Q4_K_M, 24.31GB BF16 -
PrismML Bonsai 2 27B
Hugging Face: https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-mlx-2bit, https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf
GitHub: https://github.com/PrismML-Eng/Bonsai-demo
Website: https://prismml.com/news/bonsai-2-27b
White paper: https://github.com/PrismML-Eng/Bonsai-demo/blob/main/bonsai-2-27b-whitepaper.pdfDeveloper: PrismML
Released: September 2026
Parameters: 27.36B total, 24.35B language + 2.54B embedding and LM head + 0.47B vision tower
Context: 262,144 tokens
Architecture: Qwen3.8-27B base; 64 blocks, 24Q/4KV heads, hybrid attention with about 75% linear and 25% full attention; ternary {-1, 0, +1} weights with FP16 group-wise scales at group size 128, 1.76 effective bits per weight
License: Apache 2.0
Modalities: Text + Image
Runs on: Laptop, Desktop GPU; Apple Silicon or NVIDIA GPU only
Formats: MLX 2-bit ternary, GGUF PTQ1_0, GGUF PQ2_0, optional mmproj vision pack
On disk: 5.95GB PTQ1_0 GGUF, 7.21GB PQ2_0 GGUF, 8.60GB MLX / 0.63GB mmproj vision pack -
Swiss AI Apertus 1.5
Hugging Face: https://huggingface.co/swiss-ai/Apertus-v1.5-8B, https://huggingface.co/swiss-ai/Apertus-v1.5-70B
Website: https://www.apertus-ai.org/articles/2026-07-apertus-1-5
Docs: https://www.apertus-ai.org/docsDeveloper: Swiss AI Initiative (EPFL, ETH Zurich, CSCS)
Released: July 2026
Variants: 8B / 70B
Parameters: 8B / 70B
Context: 262,144 tokens
Architecture: decoder-only transformer with xIELU activation, trained with the AdEMAMix optimizer; continued pretraining of Apertus 1.0 on 4T added tokens (8B) and 2T added tokens (70B); optional thinking mode and tool calling
License: Apache 2.0 with Acceptable Use Policy
Modalities: Text + Image + Audio (audio experimental), text output
Runs on: 8B: Laptop, Desktop GPU / 70B: Server-class hardware
Formats: safetensors
On disk: 8B: 18.40GB safetensors / 70B: 144.60GB safetensors -
Edge0
Hugging Face: https://huggingface.co/Edge0/Edge0-35B-A3B-preview, https://huggingface.co/Edge0/Edge0-8B-A1B-preview
GitHub: https://github.com/Edge0-AI/Edge0
Technical report: https://arxiv.org/abs/2609.18063Developer: Edge0 AI
Released: September 2026
Variants: Edge0-35B-A3B-preview / Edge0-8B-A1B-preview
Parameters: 35B total, 3B active / 8B total, 1B active
Context: 262,144 tokens / 131,072 tokens
Architecture: early preview release; sparse MoE; 35B-A3B: Qwen3.6-35B-A3B base, 40 layers, 256 experts, 8 per token / 8B-A1B: Ling 3.0 tiny base, 24 layers, 128 experts, 8 per token; int4 checkpoint plus Recover-LoRA adapters and prerouter heads, with experts streamed from SSD by the edge0 framework
License: Apache 2.0
Modalities: Text
Runs on: 35B-A3B: Smartphone, 12GB+ RAM on Android / 8B-A1B: Smartphone, 8GB+ RAM on Android; edge0 engines only (iOS, macOS, Android, Windows)
Formats: MLX 4-bit safetensors, LoRA and prerouter adapter safetensors
On disk: 35B-A3B: 19.51GB int4 checkpoint plus 0.18GB adapters / 8B-A1B: 4.51GB int4 checkpoint plus 0.06GB adapters -
IFM K2 Horizon
Hugging Face: https://huggingface.co/IFM/K2-Horizon-0.9B, https://huggingface.co/IFM/K2-Horizon-3.7B, https://huggingface.co/IFM/K2-Horizon-7B, https://huggingface.co/IFM/K2-Horizon-32B, https://huggingface.co/IFM/K2-Horizon-MoVA-36B-A4B
GitHub: https://github.com/ifm-ai/xllm
Website: https://ifm.ai/k2/Developer: Institute of Foundation Models (IFM), MBZUAI
Released: September 2026
Variants: 0.9B / 3.7B / 7B / 32B / MoVA-36B-A4B
Parameters: 0.9B / 3.7B / 7B / 32B / 36B total, 4B active
Context: 0.9B: 131,072 tokens / 3.7B, 7B, 32B and MoVA-36B-A4B: 524,288 tokens
Architecture: 0.9B, 3.7B, 7B, 32B: dense decoder-only, GQA with 8 KV heads, 28 / 36 / 36 / 64 layers; MoVA-36B-A4B: MoE with Mixture-of-Values attention, 48 layers, 100 FFN experts with 8 per token plus 1 shared, 64 value experts with 4 per token
License: Apache 2.0
Modalities: Text
Runs on: 0.9B: Smartwatch, Smart glasses, Smartphone / 3.7B and 7B: Smartphone, Laptop / 32B and MoVA-36B-A4B: Laptop, 32GB+ unified memory
Formats: safetensors BF16, GGUF (Q4_K_M, Q5_0, Q5_K_M, Q6_K, Q8_0, BF16), FP8 (7B, 32B, MoVA-36B-A4B), NVFP4 (32B)
On disk: 0.9B: 0.67GB Q4_K_M / 3.7B: 3.16GB Q4_K_M / 7B: 5.59GB Q4_K_M / 32B: 21.08GB Q4_K_M / MoVA-36B-A4B: 22.37GB Q4_K_M -
AI Singapore Nemotron-SEA-LION v4.8 30B-A3B
Hugging Face: https://huggingface.co/aisingapore/Nemotron-SEA-LION-v4.8-30B-A3B, https://huggingface.co/aisingapore/Nemotron-SEA-LION-v4.8-30B-A3B-GGUF, https://huggingface.co/aisingapore/Nemotron-SEA-LION-v4.8-30B-A3B-FP8, https://huggingface.co/aisingapore/Nemotron-SEA-LION-v4.8-30B-A3B-NVFP4
Website: https://sea-lion.ai/blog/uplifting-ai-in-southeast-asia-sea-announcing-nemotron-sea-lion-v4-8-in-collaboration-with-nvidia/
Technical report: https://arxiv.org/abs/2609.18310Developer: AI Singapore, with NVIDIA
Released: September 2026
Parameters: 30B total, 3B active
Context: 262,144 tokens
Architecture: Mamba2-Transformer hybrid MoE; NVIDIA Nemotron 3 Nano 30B-A3B base, continued pretraining on 150B tokens, then SFT and on-policy distillation; English plus 7 Southeast Asian languages
License: MIT
Modalities: Text
Runs on: Laptop, 32GB+ unified memory
Formats: safetensors BF16, FP8, NVFP4, GGUF (Q4_K_M, Q6_K, Q8_0, F16)
On disk: 25.43GB Q4_K_M / 22.94GB NVFP4 / 34.96GB FP8 / 65.83GB BF16 -
China Telecom Xing4.0
Hugging Face: https://huggingface.co/XingChen-AGI/Xing4.0-29B-A4B, https://huggingface.co/XingChen-AGI/Xing4.0-29B-A4B-GGUF, https://huggingface.co/XingChen-AGI/Xing4.0-29B-A4B-FP8
GitHub: https://github.com/XingChen-AGI/Xing4.0-29B-A4BDeveloper: China Telecom Artificial Intelligence Technology Co., Ltd.
Released: September 2026
Parameters: 29B total, 4B active
Context: 262,144 tokens, extensible to 512K
Architecture: MoE with mHC, MLA attention and MTP; 40 layers, 64 routed experts with 4 active plus 1 shared; trained on Ascend NPUs with MindSpore
License: Apache 2.0
Modalities: Text
Runs on: Desktop GPU
Formats: safetensors BF16, FP8, GGUF IQ4_NL
On disk: 20.1GB IQ4_NL / 33.17GB FP8 / 62.43GB BF16 -
SparkLLM Spark-X2.5
Hugging Face: https://huggingface.co/XHToken/Spark-X2.5-4B, https://huggingface.co/XHToken/Spark-X2.5-1.7B
GitHub: https://github.com/XHToken/Spark-X2.5
Website: https://dev.to/sparkllm/spark-x25-4b-17b-the-only-on-device-models-with-native-1m-token-context-now-open-source-d9oDeveloper: SparkLLM
Released: September 2026
Variants: 4B / 1.7B
Parameters: 4B / 1.7B
Context: 1,048,576 tokens
Architecture: dense; hybrid attention with 1 full-attention layer to 3 sliding-window layers, window 512; 4B: 36 layers, GQA 16 Q and 4 KV heads / 1.7B: 28 layers, GQA 8 Q and 2 KV heads
License: Apache 2.0
Modalities: Text
Runs on: Smartphone, Laptop, Edge device
Formats: safetensors BF16, FP8, INT8, GGUF (Q4_K_M, Q8_0, F16)
On disk: 4B: 2.60GB Q4_K_M / 1.7B: 1.11GB Q4_K_M