Local LMs
-
IBM Granite 4.1 3B
Hugging Face: https://huggingface.co/ibm-granite/granite-4.1-3b
GitHub: https://github.com/ibm-granite/granite-4.1-language-models
Announcement: https://huggingface.co/blog/ibm-granite/granite-4-1Developer: IBM
Released: April 2026
Parameters: 3B
Context: 512,000 tokens
Architecture: dense decoder-only transformer
License: Apache 2.0
Modalities: Text
Runs on: Smartphone, Laptop, Edge device
Formats: GGUF Q4_K_M, Q5_K_M, Q6_K, Q8_0, MLX/Core ML via conversion
On disk: 2.2GB Q4_K_M, estimated -
Liquid AI LFM2.5-VL-450M
Hugging Face: https://huggingface.co/LiquidAI/LFM2.5-VL-450M
Website: https://www.liquid.ai
Announcement: https://www.liquid.ai/blog/lfm2-5-vl-450mDeveloper: Liquid AI
Released: April 2026
Parameters: 450M
Context: 32,768 tokens
Architecture: hybrid gated-convolution + grouped-query attention LM backbone, LFM2.5-350M, with SigLIP2 NaFlex 86M vision encoder
License: LFM Open License v1.0
Modalities: Text + Vision, bounding box / object detection support
Runs on: Smartphone, Laptop, Edge device
Formats: GGUF, ONNX, MLX 4-bit/5-bit/6-bit/8-bit/bf16, safetensors
On disk: 270MB Q4, estimated -
Liquid AI LFM2-24B-A2B
Hugging Face: https://huggingface.co/LiquidAI/LFM2-24B-A2B
GitHub: https://github.com/Liquid4All/cookbook
Website: https://www.liquid.ai
Announcement: https://www.liquid.ai/blog/lfm2-24b-a2bDeveloper: Liquid AI
Released: February 2026
Parameters: 24B total, 2.3B active
Context: 32,768 tokens
Architecture: MoE, 40 layers: 30 conv + 10 attention, 64 experts, top-4 routing
License: LFM Open License v1.0
Modalities: Text
Runs on: Laptop, 14GB+ unified memory minimum, 32GB+ recommended
Formats: GGUF Q4_K_M, Q5_K_M, Q6_K, safetensors, ONNX
On disk: 14.44GB Q4_K_M / 16.93GB Q5_K_M / 19.58GB Q6_K -
Qwen3.6
Hugging Face: https://huggingface.co/Qwen/Qwen3.6-27B, https://huggingface.co/Qwen/Qwen3.6-35B-A3B
GitHub: https://github.com/QwenLM/Qwen3.6
Website: https://qwen.ai/Developer: Alibaba Cloud, Qwen team
Released: April 2026
Parameters: 27B / 35B total, 3B active
Context: 262,144 tokens, extensible to 1,010,000
Architecture: Dense / MoE, 256 experts, 8 routed + 1 shared active
License: Apache 2.0
Modalities: Text + Vision
Runs on: Laptop, 24GB+ unified memory minimum
Formats: GGUF, MLX 4-bit, community quantizations
On disk: 27B: 16GB Q4_K_M, estimated / 35B-A3B: 21GB Q4_K_M, estimated -
PrismML Bonsai Image 4B
Hugging Face: https://huggingface.co/prism-ml/bonsai-image-binary-4B-mlx-1bit, https://huggingface.co/prism-ml/bonsai-image-ternary-4B-mlx-2bit
GitHub: https://github.com/PrismML-Eng/Bonsai-Image-Demo
Website: https://prismml.com
Announcement: https://prismml.com/news/bonsai-image-4bDeveloper: PrismML
Released: May 2026
Parameters: 4B, transformer trunk
Architecture: MMDiT diffusion transformer, base architecture FLUX.2 Klein 4B, 25 blocks: 5 double-stream + 20 single-stream
Variants: Binary, 1-bit / Ternary, 2-bit
License: Apache 2.0
Modalities: Text-to-Image
Runs on: Smartphone, Laptop
Formats: MLX 1-bit, MLX 2-bit, Gemlite 1-bit/2-bit for CUDA, safetensors
On disk: Binary: 0.93GB transformer, 3.42GB total deployment payload / Ternary: 1.21GB transformer -
Google Gemma 4 Effective
Hugging Face: https://huggingface.co/collections/google/gemma-4
GitHub: https://github.com/google-gemma
Website: https://ai.google.dev/gemma/docs/core
Announcement: https://blog.google/innovation-and-ai/technology/developers-tools/gemma-4/Developer: Google DeepMind
Released: April 2026
Variants: E2B, E4B
Parameters: E2B: 2.3B effective, 5.1B with embeddings / E4B: 4.5B effective, 8B with embeddings
Context: 128,000 tokens
Architecture: Dense, hybrid local sliding window + global attention, Per-Layer Embeddings for on-device efficiency
License: Apache 2.0
Modalities: Text + Image + Audio
Runs on: Smartphone, Laptop, Edge device
Formats: GGUF via Ollama/LM Studio, safetensors via Hugging Face
On disk: E2B: 1.4GB Q4, estimated / E4B: 2.7GB Q4, estimated -
Google Gemma 4
Hugging Face: https://huggingface.co/collections/google/gemma-4
GitHub: https://github.com/google-gemma
Website: https://ai.google.dev/gemma/docs/core
Announcement: https://blog.google/innovation-and-ai/technology/developers-tools/gemma-4/, https://blog.google/innovation-and-ai/technology/developers-tools/introducing-gemma-4-12B/Developer: Google DeepMind
Released: 26B A4B and 31B: April 2026 / 12B Unified: June 2026
Variants: 12B Unified, 26B A4B, 31B
Parameters: 12B Unified: 11.95B / 26B A4B: 25.2B total, 3.8B active / 31B: 30.7B
Context: 256,000 tokens
Architecture: 31B: Dense / 26B A4B: MoE, 8 active of 128 experts plus 1 shared / 12B: Unified encoder-free dense, multimodal input projected directly into the decoder
License: Apache 2.0
Modalities: 12B: Text + Image + Audio / 26B A4B and 31B: Text + Image
Runs on: Laptop, consumer GPU/workstation class
Formats: GGUF via Ollama/LM Studio, safetensors via Hugging Face, Docker
On disk: 12B: ~7GB Q4 / 26B A4B: ~15GB Q4 / 31B: ~18GB Q4 -
NVIDIA Cosmos 3 Edge
Hugging Face: https://huggingface.co/nvidia/Cosmos3-Edge, https://huggingface.co/collections/nvidia/cosmos3
GitHub: https://github.com/nvidia/cosmos
Website: https://research.nvidia.com/labs/cosmos-lab/cosmos3/
White paper: https://research.nvidia.com/labs/cosmos-lab/cosmos3/technical-report.pdfDeveloper: NVIDIA
Released: July 2026
Variants: Cosmos3-Edge, Cosmos3-Edge-Policy-DROID
Parameters: 4B
Context: Reasoner: 256,000 tokens / Generator text input: 4,096 tokens
Architecture: Mixture-of-Transformers, two towers: an autoregressive transformer for text, a diffusion transformer for image/video/action generation
License: OpenMDW 1.1
Modalities: Text + Image + Video + Action trajectory in / Text + Image + Video + Action out
Formats: safetensors via Hugging Face, PyTorch (NVIDIA-proprietary inference only, BF16)
Runs on: NVIDIA consumer GPU (Linux), Jetson (Linux) -
Agents-A1
Hugging Face: https://huggingface.co/InternScience/Agents-A1, https://huggingface.co/InternScience/Agents-A1-4B
GitHub: https://github.com/InternScience/Agents-A1
X: https://x.com/intern_lm
Website: https://internscience.github.io/Agents-A1/
Technical report: https://arxiv.org/abs/2606.30616Developer: InternScience, Shanghai Artificial Intelligence Laboratory
Released: 35B-A3B: June 2026 / 4B: July 2026
Parameters: 35B total, 3B active / 4B dense
Context: 262,144 tokens
Architecture: Qwen3.5 base, hybrid linear + full attention (full attention every 4th layer); 35B-A3B: MoE, 40 layers, 256 experts, 8 routed + 1 shared active / 4B: dense, 32 layers
License: Apache 2.0
Modalities: Text + Vision
Runs on: 35B-A3B: Laptop, 24GB+ unified memory / 4B: Smartphone, Laptop
Formats: GGUF, MLX (community), safetensors
On disk: 35B-A3B: 19.7GB Q4_K_M / 4B: 2.5GB Q4_K_M, 4.2GB Q8_0 -
Ornith-1.0
Hugging Face: https://huggingface.co/deepreinforce-ai/Ornith-1.0-9B, https://huggingface.co/deepreinforce-ai/Ornith-1.0-35B
GitHub: https://github.com/deepreinforce-ai/Ornith-1
Website: https://deep-reinforce.com/ornith.htmlDeveloper: DeepReinforce
Released: June 2026
Parameters: 9B dense / 35B total, 8 of 256 experts active
Context: 262,144 tokens
Architecture: Qwen3.5 base; 9B: dense, 32 layers / 35B: MoE, 40 layers, 256 experts, 8 active
License: MIT
Modalities: Text + Vision
Runs on: 9B: Smartphone, Laptop / 35B: Laptop, 24GB+ unified memory
Formats: GGUF, FP8, safetensors
On disk: 9B: 5.9GB Q4_K_M / 35B: 21.4GB Q4_K_M