Best Ollama Models Right Now?

Quick Answer
For general use, qwen3.5:9b is the strongest current pick at a 6.6 GB download. For coding, qwen2.5-coder:7b at 4.7 GB does the job on far less hardware than most guides claim. For compact setups, llama3.2:3b runs in 2.0 GB. All sizes read from the Ollama library on 28 August 2026.
- βΈBest general: qwen3.5:9b β 6.6 GB
- βΈBest coding: qwen2.5-coder:7b β 4.7 GB
- βΈBest compact: llama3.2:3b β 2.0 GB
Key Takeaways
- βBest general use: qwen3.5:9b β 6.6 GB download, the newest Qwen generation actually available in Ollama
- βBest coding: qwen2.5-coder:7b β 4.7 GB, and it is the value pick because qwen3-coder starts at 19 GB
- βBest compact: llama3.2:3b β 2.0 GB, or deepseek-r1:1.5b at 1.1 GB if you want reasoning traces
- βWatch the figure you are comparing: these are download sizes from the Ollama library, not VRAM requirements. Leave headroom above the download size for your context window
- βA model from six months ago with mature quantization often beats a brand-new release with limited community support
The Three Tier Leaders
For general use the current pick is qwen3.5:9b at a 6.6 GB download. Sizes below were read directly from the Ollama library on 28 August 2026.
"Best" in practice means the highest balance of output quality, inference speed, and memory efficiency β not raw benchmark score alone. A 9B model you can actually fit is more useful than a 30B model that swaps to disk.
One number worth correcting: coding models are cheaper than most guides suggest. qwen2.5-coder:7b is a 4.7 GB download and handles Python, TypeScript and Go without special prompting. The newer qwen3-coder starts at 19 GB, so it is a different class of hardware entirely, not a drop-in upgrade.
| Tier | Model | Download | Why It Leads |
|---|---|---|---|
| Compact | llama3.2:3b | 2.0 GB | Best quality-per-GB at the small end; 81.6M pulls |
| General | qwen3.5:9b | 6.6 GB | Newest Qwen generation available in Ollama |
| Coding | qwen2.5-coder:7b | 4.7 GB | Strong coding output at a fraction of qwen3-coder size |
| Reasoning | deepseek-r1:8b | 5.2 GB | Second most-pulled model on Ollama at 92M pulls |
When Newer Isn't Better
A new model release does not automatically become the best Ollama pick. Quantization quality, community fine-tunes, and Ollama integration maturity take 4β8 weeks to catch up with a fresh release.
The pull counts make the point better than any benchmark. llama3.1 is still the most-downloaded model in the library at 118.9M pulls, and llama3.2 sits at 81.6M β both ahead of every newer arrival. That is not inertia; it is people picking the thing whose quantizations are well-optimized and whose behaviour is predictable across hardware.
If you want the newer generation, gemma4 (7.2 GB) and gpt-oss:20b (14 GB, 128K context) are both worth a look. Give a model 6+ weeks at the top before you rely on it in production. For a deeper look at evaluating models for your specific workload, see the top open-source models for Ollama.
Related Guides
- βΈBest VPN for Downloading AI Models -- VPN for AI downloads
- βΈOllama 128K Context Models -- long context models
- βΈOllama Latest Version: What's New? -- Ollama updates
- βΈMistral Small 24B vs Qwen3 14B vs Llama 3.1 8B -- model comparison
Quick Answers About Ollama Models
Should I always use the newest Ollama model?βΎ
Which Ollama model is best for coding right now?βΎ
How much VRAM do I actually need?βΎ
Are Qwen models better than Llama models in 2026?βΎ
Want the full breakdown?
Read the complete guide βRelated Prompt Bites