Local Foundation Models
-
Kakao Kanana-2
Hugging Face: https://huggingface.co/kakaocorp/kanana-2-3b-instruct, https://huggingface.co/kakaocorp/kanana-2-1.3b-instruct
Website: https://tech.kakao.com/posts/826Developer: Kakao, Kanana LLM
Released: July 2026
Variants: 3B, 1.3B
Parameters: 3B / 1.3B
Context: 32,768 tokens
Architecture: 3B: dense, pretrained from scratch, SFT + RL / 1.3B: cascade-pruned and distilled from 3B, sliding-window attention
License: Kanana Open License
Modalities: Text
Runs on: Smartphone, Laptop
Formats: safetensors, MLX 4-bit, 6-bit, 8-bit
On disk: 3B: 1.99GB MLX 4-bit -
Microsoft Fara1.5
Hugging Face: https://huggingface.co/microsoft/Fara1.5-4B, https://huggingface.co/microsoft/Fara1.5-9B, https://huggingface.co/microsoft/Fara1.5-27B
GitHub: https://github.com/microsoft/fara
Website: https://labs.ai.azure.com/innovations/fara1-5/Developer: Microsoft Research AI Frontiers
Released: May 2026
Variants: 4B, 9B, 27B
Parameters: 4B / 9B / 27B
Context: 262,144 tokens
Architecture: multimodal decoder-only LM, image + text to text
License: MIT
Modalities: Text + Image
Formats: safetensors
On disk: 4B: 9.08GB / 9B: 18.82GB / 27B: 54.71GB safetensors -
AMD Instella-MoE-16B-A3B-Think
Hugging Face: https://huggingface.co/amd/Instella-MoE-16B-A3B-Think
GitHub: https://github.com/AMD-AGI/Instella-MoE
Website: https://rocm.blogs.amd.com/artificial-intelligence/instella-moe/README.htmlDeveloper: AMD
Released: July 2026
Parameters: 16B total, 2.8B active
Context: 32,768 tokens
Architecture: decoder-only MoE, Gated Multi-head Latent Attention, FarSkip-Collective connectivity
License: Research RAIL
Modalities: Text
Formats: safetensors
On disk: 31.73GB safetensors -
Cisco Antares
Hugging Face: https://huggingface.co/fdtn-ai/antares-1b, https://huggingface.co/fdtn-ai/antares-350m
Website: https://cisco-foundation-ai.github.io/antares/
Announcement: https://cisco-foundation-ai.github.io/blogs/antares-beyond-vlocbench/Developer: Cisco Foundation AI
Released: July 2026
Variants: 1b, 350m
Parameters: 1B / 350M
Architecture: fine-tuned from IBM Granite 4.0, GraniteMoEHybrid
License: Apache 2.0
Modalities: Text
Runs on: Smartphone, Laptop, Edge device
Formats: safetensors
On disk: 1b: 3.67GB safetensors / 350m: 0.70GB safetensors -
AI9Stars G9v3-3B
Hugging Face: https://huggingface.co/ai9stars/G9v3-3B
GitHub: https://github.com/AI9StarsDeveloper: AI9Stars
Parameters: ~3B
Context: 131,072 tokens
Architecture: dense causal LM, LlamaForCausalLM
License: Apache 2.0
Modalities: Text
Formats: safetensors
On disk: 5.99GB safetensors -
Tencent Hy-Embodied-RxBrain-1.0
Hugging Face: https://huggingface.co/tencent/Hy-Embodied-RxBrain-1.0
GitHub: https://github.com/Tencent-Hunyuan/Hy-Embodied-RxBrain-1.0
Website: https://tairos.tencent.com/openSourceModels/hy-embodied-rxbrain-1.0
Technical report: https://arxiv.org/abs/2607.14187Developer: Tencent Robotics X, Futian Laboratory, Tencent Hy Team
Released: July 2026
Parameters: ~6.2B
Architecture: Unified Mixture-of-Transformers, modality-specific text, vision, and generation pathways
License: Apache 2.0
Modalities: Text + Image + Video
Runs on: NVIDIA GPU, CUDA 12.x, Linux recommended
Formats: safetensors
On disk: 12.42GB safetensors -
Poolside Laguna XS 2.1
Hugging Face: https://huggingface.co/poolside/Laguna-XS-2.1, https://huggingface.co/poolside/Laguna-XS-2.1-GGUF
Website: https://poolside.ai/blog/introducing-laguna-xs-2-1Developer: Poolside
Released: July 2026
Variants: BF16, FP8, NVFP4, INT4; GGUF BF16, Q4_K_M
Parameters: 33B total, 3B active
Context: 262,144 tokens
Architecture: MoE, 40 layers: 10 global-attention + 30 sliding-window-attention; 256 experts + 1 shared expert
License: OpenMDW-1.1
Modalities: Text
Runs on: Mac with 36GB RAM
Formats: safetensors, GGUF
On disk: 20.27GB Q4_K_M GGUF / 66.89GB BF16 safetensors -
Meta Muse Glimmer-30B
Hugging Face: https://huggingface.co/meta-models/Muse-Glimmer-30B
Announcement: https://research.meta.ai/blog/introducing-muse-glimmer-open-agentic-model
Technical report: https://research.meta.ai/static/muse-glimmer-methodologyDeveloper: Meta Superintelligence Lab
Released: August 2026
Parameters: 29.6B total, including perception encoder
Context: 131,072+ tokens
Architecture: Dense causal transformer with ViT-G/14 perception encoder; 52 layers, GQA, SwiGLU, RoPE
License: Apache 2.0
Modalities: Text + Image
Runs on: MacBook M4 Max/M5 Max, RTX 5090; 24GB+ memory with 4-bit weights
Formats: BF16 safetensors, 4-bit quantized weights
On disk: 17GB K-Quant -
webAI TwIL-LM
Hugging Face: https://huggingface.co/webAI-Official/TwIL-LM, https://huggingface.co/webAI-Official/TwIL-LM3
Website: https://www.webai.com/blog/webai-releases-twil-lm-a-family-of-formal-logic-models-that-outreason-a-120b-model-and-run-on-an-iphoneDeveloper: webAI Intelligence Lab
Released: August 2026
Variants: TwIL-LM 1.7B, TwIL-LM3 3B
Parameters: 1.7B: 1.78B total, 1.71B backbone + 72M LoRA / 3B: 3B
Context: 1.7B: 8,192 tokens / 3B: 65,536 tokens
Architecture: 1.7B: SmolLM2-1.7B-Instruct base, dense Llama, 24 layers, 32 heads, LoRA rank 64 SFT / 3B: SmolLM3-3B base, dense, 36 layers, GQA 16Q/4KV, NoPE every 4th layer; LoRA SFT, checkpoint fusion, WiSE-FT interpolation, GRPO reinforcement learning
License: webAI Non-Commercial License ver. 1.0
Modalities: Text
Runs on: 1.7B: Smartphone, Laptop / 3B: Laptop, 4GB VRAM or CPU
Formats: 1.7B: merged GGUF Q4_K_M, Q5_K_M, Q8_0, f16 / 3B: safetensors, GGUF Q4_K_M, Q5_K_M, Q6_K, Q8_0, F16
On disk: 1.7B: 1.06GB Q4_K_M / 3B: 1.92GB Q4_K_M -
Ling 3.0
Hugging Face: https://huggingface.co/inclusionAI/Ling-3.0-flash, https://huggingface.co/inclusionAI/Ling-3.0-tiny, https://huggingface.co/inclusionAI/Ling-3.0-tiny-int4, https://huggingface.co/inclusionAI/Ling-3.0-tiny-fp8
GitHub: https://github.com/inclusionAI/Ling
X: https://x.com/AntLingAGI/status/2080351022028095681
Website: https://www.ant-ling.com/en
Docs: https://github.com/inclusionAI/ling-cookbookDeveloper: Ant Group, InclusionAI
Released: July 2026
Variants: Ling-3.0-flash, Ling-3.0-tiny
Parameters: flash: 124B total, 5.1B active / tiny: 7.9B total, 1.3B active
Context: flash: 262,144 tokens / tiny: 131,072 tokens, 262,144 with YaRN
Architecture: BailingMoE hybrid; flash: linear KDA + MLA attention with sparse MoE / tiny: 3:1 KDA to MLA blocks, 128 routed experts, 8 routed + 1 shared active
License: MIT
Modalities: Text
Runs on: flash: NVIDIA DGX Spark / tiny: Laptop, 48GB unified memory
Formats: safetensors BF16, safetensors FP8, safetensors INT4, GGUF Q4_K_M
On disk: flash: 60.5GB Q4_K_M GGUF, 254.98GB BF16 safetensors / tiny: 5.81GB INT4, 8.41GB FP8, 15.79GB BF16 safetensors -
Qwen3.8
Hugging Face: https://huggingface.co/Qwen/Qwen3.8-27B, https://huggingface.co/Qwen/Qwen3.8-27B-FP8
GitHub: https://github.com/QwenLM/Qwen3
X: https://x.com/Alibaba_Qwen/status/2088280182356611304
Website: https://qwen.ai/blog?id=qwen3.8Developer: Alibaba Cloud, Qwen team
Released: August 2026
Parameters: 27B
Context: 262,144 tokens
Architecture: Qwen3.5 foundation, dense native vision-language model; hybrid linear + full attention
License: Apache 2.0
Modalities: Text + Image + Video
Runs on: Laptop, 24GB+ unified memory, estimated / Desktop GPU
Formats: safetensors BF16, safetensors FP8, GGUF and MLX community quantizations
On disk: 16.05GB MLX 4-bit / 30.87GB FP8 safetensors / 55.56GB BF16 safetensors -
NVIDIA Nemotron 3.5 Lightning
Hugging Face: https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16, https://huggingface.co/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4
GitHub: https://github.com/NVIDIA-NeMo/Nemotron
Website: https://build.nvidia.com/nvidia/nemotron-3.5-lightning-30b-a3bDeveloper: NVIDIA
Released: August 2026
Variants: Instruct, Base, DSpark and DFlash speculative drafters
Parameters: 30B total, 3B active
Context: 1,048,576 tokens
Architecture: hybrid LatentMoE interleaving Mamba-2, MoE and attention; 52 layers, 128 routed experts, 6 active, 32Q/2KV heads, Multi-Token Prediction
License: OpenMDW License Agreement, version 1.1
Modalities: Text
Runs on: Desktop GPU, Edge device, DGX Spark; NVIDIA only
Formats: safetensors BF16, safetensors NVFP4, GGUF community conversion
On disk: 17.82GB NVFP4 safetensors / 31.58GB BF16 safetensors -
Google DiffusionGemma
Hugging Face: https://huggingface.co/google/diffusiongemma-26B-A4B-it
GitHub: https://github.com/google-gemma
Website: https://ai.google.dev/gemma/docs/diffusiongemma
Announcement: https://blog.google/innovation-and-ai/technology/developers-tools/diffusion-gemma-faster-text-generation/Developer: Google DeepMind
Released: June 2026
Parameters: 25.2B total, 3.8B active
Context: 256,000 tokens
Architecture: discrete text diffusion on the Gemma 4 26B A4B MoE foundation; autoregressive encoder prefills the prompt into a KV cache, decoder applies bidirectional attention over a 256-token canvas, block-autoregressive multi-canvas sampling; 30 layers, 8 active of 128 experts plus 1 shared, 1,024 sliding window, 550M vision encoder
License: Apache 2.0
Modalities: Text + Image + Video in, Text out
Runs on: Desktop GPU, 18GB+ VRAM
Formats: safetensors
On disk: 51.65GB safetensors -
IBM Granite Swash
Hugging Face: https://huggingface.co/ibm-granite/granite-swash-2b, https://huggingface.co/ibm-granite/granite-swash-3b-a600m
GitHub: https://github.com/ibm-granite/granite-4.1-language-modelsDeveloper: IBM Granite Team
Released: July 2026
Variants: SWASH-2B, SWASH-3B-A600M
Parameters: 2B / 3B total, 600M active
Context: 8,192 tokens
Architecture: sliding window attention with learnable per-head attention sinks, LSE-scaled; 2B: dense decoder-only, 24 layers, 7 full-attention + 17 sliding-window layers, window 128, GQA, SwiGLU, RoPE, RMSNorm / 3B-A600M: MoE, 28 layers, 48 experts, 4 routed active
License: Apache 2.0
Modalities: Text
Runs on: Smartphone, Laptop, Edge device
Formats: safetensors
On disk: 2B: 4.29GB safetensors / 3B-A600M: 6.04GB safetensors -
IBM Granite Vision 4.1
Hugging Face: https://huggingface.co/ibm-granite/granite-vision-4.1-4b, https://huggingface.co/ibm-granite/granite-vision-4.1-4b-GGUF
GitHub: https://github.com/ibm-granite/granite-vision-models
Website: https://www.ibm.com/granite/docs/models/vision
Announcement: https://research.ibm.com/blog/granite-4-1-ai-foundation-modelsDeveloper: IBM
Released: April 2026
Parameters: 4B total, Granite 4.1 3B language model plus vision encoder and projectors
Context: 131,072 tokens
Architecture: SigLIP2 so400m patch16-384 vision encoder over 384x384 image tiles, windowed Q-Former projectors compressing each 4x4 patch window to 2x2 tokens, and a Granite 4.1 3B language model with rank-256 LoRA across all self-attention projections
License: Apache 2.0
Modalities: Text + Image
Runs on: Smartphone, Laptop, Edge device
Formats: safetensors, GGUF Q4_K_M, Q5_K_M, Q6_K, Q8_0, bf16, with f16 mmproj
On disk: 2.10GB Q4_K_M plus 1.16GB f16 mmproj / 6.81GB bf16 -
Microsoft Mage-VL
Hugging Face: https://huggingface.co/microsoft/Mage-VL
GitHub: https://github.com/microsoft/Mage
Website: https://microsoft.github.io/Mage/vl/
Technical report: https://arxiv.org/abs/2607.24904Developer: Microsoft Mage Team
Released: July 2026
Parameters: 4B
Context: 262,144 tokens
Architecture: Mage-ViT codec-native visual encoder trained from scratch, 24 layers, feeding a two-layer MLP projector into a Qwen3-4B-Instruct-2507 causal decoder; separate cognition gate for proactive streaming
License: Apache 2.0
Modalities: Text + Image + Video
Runs on: Laptop, 12GB+ memory at BF16, estimated / Desktop GPU
Formats: safetensors, bundled streaming gate and neural codec
On disk: 9.48GB safetensors -
Cohere Labs North Micro Vision
Hugging Face: https://huggingface.co/CohereLabs/North-Micro-Vision-Instruct
Technical report: https://huggingface.co/blog/CohereLabs/meet-north-micro-vision-instructDeveloper: Cohere Labs
Released: August 2026
Parameters: 2.4B total, 2B language model + 400M vision encoder
Context: 128,000 tokens, multimodal validated to 8,192
Architecture: custom native-resolution vision encoder with DeepStack patch embeddings injected into early decoder layers, projector, and the Command A+ style North Micro LLM: three sliding-window attention layers with RoPE plus one global layer without positional embeddings
License: Apache 2.0
Modalities: Text + Image
Runs on: Smartphone, Laptop, Edge device, with quantization
Formats: safetensors BF16, MLX 4-bit and 8-bit community conversions
On disk: 4.97GB BF16 safetensors -
Cactus Compute Needle 2
Hugging Face: https://huggingface.co/Cactus-Compute/needle2
GitHub: https://github.com/cactus-compute/needle
X: https://x.com/cactuscompute/status/2086865960669983035
Website: https://cactuscompute.com/needleDeveloper: Cactus Compute
Released: August 2026
Parameters: 45M
Context: 2,048 tokens
Architecture: Simple Attention Network, 27 layers, hidden 512, 8Q/4KV GQA, Hadamard MLP, engram sites, CQ2 quantization at 2.2 effective bits; byte-level grammar-constrained decoding and a tool-retrieval head
License: Apache 2.0
Modalities: Text
Runs on: Smartphone, Headset, Edge device, Microcontroller
Formats: cact single binary; ARM64, x86-64, ARMv7, RISC-V and WebAssembly builds
On disk: 13.7MB cact -
Syzygy Mach-1 Additive 35B
Hugging Face: https://huggingface.co/SyzygyResearch/Mach-1-Additive-35B
X: https://x.com/syzygyeng/status/2084350792841195992
Website: https://withsyzygy.com/mach-1
Docs: https://withsyzygy.com/docs/machDeveloper: Syzygy Research
Released: August 2026
Parameters: 35B total, 8 of 256 experts active
Context: 262,144 tokens
Architecture: Qwen3.5 MoE topology, 40 layers, 256 experts, 8 active, hybrid linear attention with full attention every 4th layer; additive 1.7-bit weights with no weight multiplication
License: Apache 2.0
Modalities: Text
Runs on: Laptop, 16GB+ unified memory; Apple Silicon only
Formats: packed 1.7-bit safetensors, MLX
On disk: 7.0GB -
OpenMOSS MOSS-VL
Hugging Face: https://huggingface.co/OpenMOSS-Team/MOSS-VL-Instruct-0708-FP8, https://huggingface.co/OpenMOSS-Team/MOSS-VL-Realtime-FP8
GitHub: https://github.com/OpenMOSS/MOSS-VL
Website: https://openmoss.ai/MOSS-VL/
Technical report: https://arxiv.org/abs/2606.07639Developer: OpenMOSS Team
Released: August 2026
Variants: Instruct-0708, Realtime
Parameters: 11B
Context: 262,144 tokens
Architecture: unified cross-attention multimodal model, 48 language layers with 12 cross-attention layers, XRoPE 3D spatiotemporal positions, absolute frame timestamps for streaming video
License: Apache 2.0
Modalities: Text + Image + Video
Runs on: Desktop GPU; Instruct: 24GB VRAM / Realtime: 26GB+ VRAM; NVIDIA only
Formats: FP8 compressed-tensors with BF16 cross-attention and vision, HQQ INT8 KV cache
On disk: 15.73GB FP8 safetensors