Key Takeaways
- Best reasoning at small scale: Phi-4 Mini 3.8B -- 68% MMLU, 70% HumanEval, runs on 4 GB RAM.
- Best small/edge model: Gemma 4 E2B -- MMLU Pro 60%, 128K context, ~2 GB RAM, Apache 2.0.
- Best small coding model: Qwen2.5 3B -- 65% HumanEval at ~2 GB RAM.
- Best general-purpose 3B: Llama 3.2 3B -- most community support, 128K context, 2.5 GB RAM.
- No sub-2B-effective model produces output quality suitable for professional tasks. Use 3B+ (or Gemma 4 E2B) for real work.
📍 In One Sentence
Small local LLMs (1B-4B parameters) run on 4-8 GB RAM machines at 30-70 tokens/sec, with Phi-4 Mini 3.8B leading reasoning, Gemma 4 E2B best for small/edge devices with 128K context, Qwen2.5 3B best for coding, and Llama 3.2 3B best for general use.
💬 In Plain Terms
If your laptop only has 4-8 GB of RAM, you can still run AI models -- just smaller ones. These sub-4B models are fast enough for real-time chat and handle simple tasks like summarizing, basic Q&A, and translating short text well, but they struggle with multi-step reasoning or writing long, coherent documents. For most tasks, though, they're a solid fit for entry-level hardware.
What Is a "Small" Local LLM and When Should You Use One?
A small local LLM is typically defined as a model with fewer than 4 billion parameters. At Q4_K_M quantization, these models require 1.5-3 GB of RAM -- well within the constraints of entry-level laptops with 4-8 GB total memory.
Small models are appropriate for: quick summarization, simple Q&A, code snippet explanation, translation of short texts, and classification tasks. They are not suitable for multi-step reasoning, complex code generation, or writing long-form coherent documents.
The quality gap between a 3B and 7B model is significant -- roughly equivalent to the gap between GPT-4o mini and GPT-5.5. For users with 8 GB RAM, a 7B model at Q4_K_M is almost always the better choice if the machine has headroom. See Best Beginner Local LLM Models for 7B recommendations.
Which Model Should You Use? Quick Decision Guide?
Phi-4 Mini 3.8B -- Best Reasoning Performance in the Sub-4B Class
Microsoft Phi-4 Mini achieves 68% on MMLU and 70% on HumanEval -- scores that exceed many 7B models released before 2025. This is possible because Phi-4 Mini was trained on a curated synthetic dataset focused on reasoning and problem-solving, rather than broad web text.
Phi-4 Mini is the recommended choice for users who primarily need reasoning (math, logic, step-by-step explanations) or coding assistance on hardware with 4-6 GB RAM.
Spec | Value |
|---|---|
| MMLU | 68% |
| HumanEval | 70% |
| RAM (Q4_K_M) | ~2.5 GB |
| Context | 128K tokens |
| CPU speed | 30-50 tok/sec |
| Ollama command | ollama run phi4-mini |
Gemma 4 E2B -- Best Small/Edge Local LLM
Google Gemma 4 E2B is the current generation small model, replacing the older Gemma 2 2B. It scores 60% on MMLU Pro (a harder benchmark than the original MMLU), 44% on LiveCodeBench v6, and 43.4% on GPQA Diamond, per Google's own model card. Effective parameters are 2.3B (5.1B including embeddings, via a Per-Layer Embeddings architecture), fitting in ~2 GB of RAM or VRAM.
The 128K context window is a major upgrade over Gemma 2 2B's 8K -- it now matches Phi-4 Mini and Llama 3.2 3B, removing the old context-length trade-off for long documents. CPU throughput varies by hardware and quantization (Google reports ~7.6 tok/sec decode on a Raspberry Pi 5; independent x86 reports range 8-48 tok/sec) -- check current figures for your specific device rather than assuming a fixed number. Apache 2.0 license.
Spec | Value |
|---|---|
| MMLU Pro | 60% |
| LiveCodeBench v6 | 44% |
| GPQA Diamond | 43.4% |
| RAM (Q4_K_M) | ~2 GB |
| Context | 128K tokens |
| Ollama command | ollama pull gemma4:e2b |
Qwen2.5 3B -- Best Small Model for Coding Tasks
Qwen2.5 3B scores 65% on HumanEval -- 5 percentage points above Llama 3.2 3B -- making it the best choice for coding tasks at the 3B scale. It includes JSON mode and function calling support, and natively handles 29 languages.
For non-coding tasks in English, Llama 3.2 3B and Phi-4 Mini produce more natural prose. Choose Qwen2.5 3B specifically when coding or multilingual output is the primary use case.
Spec | Value |
|---|---|
| MMLU | 62% |
| HumanEval | 65% |
| RAM (Q4_K_M) | ~2 GB |
| Context | 128K tokens |
| CPU speed | 25-40 tok/sec |
| Ollama command | ollama run qwen2.5:3b |
Llama 3.2 3B -- Best General-Purpose Small Model
Meta Llama 3.2 3B is the most widely documented and community-supported 3B model. It scores 58% on MMLU and 60% on HumanEval -- slightly below Phi-4 Mini on both -- but has the widest tool support, the most fine-tunes available, and the largest collection of community guides.
The 128K context window is the same as larger Llama 3.x models, making it suitable for summarizing medium-length documents. For a first small model, Llama 3.2 3B remains the safest choice due to predictable behavior and extensive documentation.
Spec | Value |
|---|---|
| MMLU | 58% |
| RAM (Q4_K_M) | ~2.5 GB |
| Context | 128K tokens |
| CPU speed | 25-45 tok/sec |
| Ollama command | ollama run llama3.2:3b |
Llama 3.2 1B -- Absolute Minimum for Any Useful Output
Llama 3.2 1B is the smallest Llama model Meta publishes, requiring only 1.3 GB of RAM and generating 60-90 tok/sec on CPU -- the fastest locally-runnable model on this page. Output quality is marginal: it handles very simple classification and keyword extraction but struggles with coherent multi-sentence responses. Use Llama 3.2 1B only when RAM is genuinely the binding constraint (under 3 GB available) or for testing tool integrations.
Full Comparison: Best Small Local LLMs Under 4B Parameters
Model | MMLU | HumanEval | RAM | Context | Best For |
|---|---|---|---|---|---|
| Phi-4 Mini 3.8B | 68% | 70% | 2.5 GB | 128K | Reasoning, coding |
| Qwen2.5 3B | 62% | 65% | 2 GB | 128K | Coding, multilingual |
| Llama 3.2 3B | 58% | 60% | 2.5 GB | 128K | General use, first model |
| Gemma 4 E2B | — | — | 2 GB | 128K | Edge/SBC (MMLU Pro 60%) |
| Llama 3.2 1B | 32% | 28% | 1.3 GB | 128K | Absolute minimum RAM |

Small Local LLMs by Region
EU / GDPR: For EU professionals running AI on constrained hardware -- field work, air-gapped environments, older enterprise laptops -- small local models provide GDPR-compliant inference with zero data egress. A Phi-4 Mini 3.8B running on a standard-issue corporate laptop (8 GB RAM) keeps all processed text on-device under GDPR Article 5 (data minimization). For German BSI compliance documentation: Phi-4 Mini (Microsoft, MIT licence), Gemma 4 E2B (Google, Apache 2.0), and Llama 3.2 3B (Meta, Llama Community licence) all provide versioned model identifiers via their Ollama tags, satisfying AI tool documentation requirements. Mistral does not currently offer a sub-4B model. For EU organizations preferring an EU-origin model at this size class, options are limited until Mistral releases a sub-4B variant.
Japan (METI): For Japanese-language tasks at the small model tier, Qwen2.5 3B is the only model in this comparison with native Japanese tokenization. Llama 3.2 3B handles Japanese but with lower token efficiency. For Japanese summarization or translation on constrained hardware: `ollama run qwen2.5:3b`. The speed advantage of small models is particularly relevant for Japanese enterprise use: 25-40 tok/sec on CPU provides adequate real-time response for chat interfaces on standard-issue office hardware.
China: Qwen2.5 3B (Alibaba, Apache 2.0) is the natural choice for Chinese-language small model deployment. Native Chinese tokenization processes Mandarin text 30-40% more efficiently than Llama at equivalent parameter count. For IoT and edge deployments under China's Data Security Law (数据安全法): `ollama run qwen2.5:3b` runs on any Linux device with 4 GB RAM and processes all text on-device with no external API calls.
What Are the Common Mistakes When Running Small Local LLMs?
- Using Q8_0 quantization instead of Q4_K_M: Q8_0 requires nearly double the RAM of Q4_K_M for minimal quality improvement at small scale. A Llama 3.2 3B model at Q8_0 needs ~3.8 GB RAM vs ~2.5 GB for Q4_K_M. On a 4 GB machine, Q8_0 may trigger swap usage and make inference 3-5x slower. Always use Q4_K_M as the default for sub-4B models.
- Running a base model instead of the instruct variant: Base models (e.g., `llama3.2:3b-text`) are pre-fine-tuning checkpoints trained to predict the next token in text. They do not follow instructions. When you ask a base model "What is 2+2?", it may complete the sentence as a quiz rather than answer "4". Always use the instruct variant: `llama3.2:3b` (Ollama defaults to instruct for named models).
- Expecting 7B model quality from a 3B model: A 3B model at 68% MMLU (Phi-4 Mini) performs similarly to a 2023-era GPT-3.5 Mini on general tasks. Complex reasoning chains, long-form writing, and nuanced code generation will produce noticeably lower quality than a 7B model. If output quality is insufficient, upgrade to a 7B model -- the RAM difference is ~2 GB (2.5 GB → 4.5 GB).
Understanding Quantization: RAM vs Quality Trade-off

Common Questions About Small Local LLM Models
What is the smallest local LLM that produces useful output?
The practical minimum for useful output is a 3B model at Q4_K_M quantization. Llama 3.2 1B produces coherent single sentences but struggles with multi-step instructions, longer responses, and complex reasoning. Gemma 4 E2B (2.3B effective) is more capable than that -- usable for summarization and simple Q&A, and its 128K context handles longer documents. For anything more complex, start with a 3B model.
Can a 3B model run on a phone?
Yes -- Llama 3.2 1B and 3B are specifically designed for on-device mobile deployment. Meta provides optimized builds for iOS (via MLC LLM) and Android. Inference on a modern phone (Snapdragon 8 Gen 3 or Apple A17 Pro) produces 15-30 tok/sec for 1B models. LM Studio and Ollama do not currently run on iOS or Android -- mobile requires separate frameworks.
Are small models good for summarization?
Yes -- summarization is one of the strongest use cases for small models. Gemma 4 E2B and Llama 3.2 3B reliably produce accurate summaries, and Gemma 4 E2B's upgrade to a 128K context window (from the older Gemma 2's 8K) means document length is rarely the limiting factor at this scale anymore.
How much faster is a 2B model than a 7B model on the same hardware?
Meaningfully faster on CPU, though the multiple depends heavily on hardware and quantization -- Google's own figures for Gemma 4 E2B show ~7.6 tok/sec decode on a Raspberry Pi 5, while independent x86 CPU reports range 8-48 tok/sec depending on quantization level. On a GPU, the speed advantage narrows because GPU throughput is less constrained by model size. Check current benchmarks for your specific hardware rather than assuming a fixed multiplier.
Do small models support function calling?
Some do. Qwen2.5 3B supports function calling and JSON mode. Llama 3.2 3B has basic tool use support. Check each model's own documentation before building a pipeline that depends on structured output, since support varies by version.
Which small model is best for languages other than English?
Qwen2.5 3B supports 29 languages natively including Chinese, Japanese, Korean, and Arabic. Phi-4 Mini is primarily English-optimized. For non-English tasks at the small model scale, Qwen2.5 3B is the clear choice. See Qwen vs Llama vs Mistral multilingual comparison for a full language comparison.
What is the difference between Phi-4 Mini and Llama 3.2 3B for everyday tasks?
Phi-4 Mini outperforms Llama 3.2 3B on reasoning, math, and coding (68% vs 58% MMLU, 70% vs 60% HumanEval) at nearly identical RAM (2.5 GB each). For everyday tasks -- Q&A, summarization, simple explanations -- the quality gap is noticeable but not dramatic. Llama 3.2 3B has broader community support and more fine-tunes available. Choose Phi-4 Mini for structured reasoning; Llama 3.2 3B for general chat and broader compatibility.
Can I run two small models simultaneously?
Yes, if total RAM permits. Two 3B models at Q4_K_M use ~5 GB combined -- feasible on an 8 GB machine with a lean OS. Ollama loads one model at a time per process by default. Run two Ollama instances on different ports (OLLAMA_HOST=:11434 and OLLAMA_HOST=:11435) to serve two models in parallel. This is useful for A/B testing outputs.
Do small models work for RAG (retrieval-augmented generation)?
Yes for simple RAG. Llama 3.2 3B and Phi-4 Mini can answer questions over retrieved document chunks reliably. For RAG over large knowledge bases requiring multi-hop reasoning, 7B+ models perform more consistently. GPT4All's LocalDocs feature uses a 3B model for document Q&A and works well for personal document collections.
Is Phi-4 Mini better than Llama 3.2 3B for coding?
Yes. Phi-4 Mini scores 70% on HumanEval vs 60% for Llama 3.2 3B -- a meaningful 10-point gap at this scale. For coding assistance on 4-6 GB RAM machines, Phi-4 Mini is the recommended choice. For multilingual coding (non-Python), Qwen2.5 3B at 65% HumanEval is competitive with Phi-4 Mini while also supporting function calling.
Using Q8_0 quantization instead of Q4_K_M
Q8_0 requires nearly double the RAM of Q4_K_M for minimal quality improvement at small scale. A Llama 3.2 3B model at Q8_0 needs ~3.8 GB RAM vs ~2.5 GB for Q4_K_M. On a 4 GB machine, Q8_0 may trigger swap usage and make inference 3-5× slower. Always use Q4_K_M as the default for sub-4B models.
Running a base model instead of the instruct variant
Base models (e.g., llama3.2:3b-text) are pre-fine-tuning checkpoints trained to predict the next token in text. They do not follow instructions. When you ask a base model "What is 2+2?", it may complete the sentence as a quiz rather than answer "4". Always use the instruct variant: llama3.2:3b (Ollama defaults to instruct for named models).
Expecting 7B model quality from a 3B model
A 3B model at 68% MMLU (Phi-4 Mini) performs similarly to a 2023-era GPT-3.5 Mini on general tasks. Complex reasoning chains, long-form writing, and nuanced code generation will produce noticeably lower quality than a 7B model. If output quality is insufficient, upgrade to a 7B model -- the RAM difference is ~2 GB (2.5 GB → 4.5 GB).
What is the smallest local LLM that is actually usable?
Around 1B parameters is the floor for coherent output, and Llama 3.2 1B is the smallest model here worth running for real work. Below that, tiny models will load and generate on almost anything, but the text stops holding together across a few sentences, so they are better treated as autocomplete than as an assistant. If your constraint is disk and RAM rather than quality, the smallest usable local LLM with a small footprint is a sub-2B model at Q4; if you have 4GB or more to spend, a 4B model is a large step up in coherence for very little extra memory.
What is the smallest LLM model you can run locally?
Llama 3.2 1B is the smallest Llama model and the smallest practical option in this comparison -- it needs only 1.3 GB of RAM and runs at 60-90 tok/sec on CPU, faster than any other model on this page. It handles short classification and keyword-extraction tasks but breaks down on longer, multi-step output. If you want a small model (sometimes called a 'mini LLM') that's actually usable day-to-day, treat Llama 3.2 1B as the floor for hardware, not for quality -- step up to Gemma 4 E2B or a 3B model like Qwen2.5 3B or Llama 3.2 3B for real work.
Sources
- Hugging Face Open LLM Leaderboard -- open-llm-leaderboard.hf.space (MMLU and HumanEval scores)
- Microsoft Phi-4 Technical Report -- microsoft.com/en-us/research/publication/phi-4-technical-report/
- Meta Llama 3.2 Model Card -- huggingface.co/meta-llama/Llama-3.2-3B-Instruct
- Google Gemma 4 Model Card -- huggingface.co/google/gemma-4-e2b-it
