Skip to main content
PromptQuorumBuilt for humans. Structured for AI.

Best Ollama Models Right Now?

Best Ollama Models Right Now?

Quick Answer

For general use, qwen3.5:9b is the strongest current pick at a 6.6 GB download. For coding, qwen2.5-coder:7b at 4.7 GB does the job on far less hardware than most guides claim. For compact setups, llama3.2:3b runs in 2.0 GB. All sizes read from the Ollama library on 28 August 2026.

  • β–ΈBest general: qwen3.5:9b β€” 6.6 GB
  • β–ΈBest coding: qwen2.5-coder:7b β€” 4.7 GB
  • β–ΈBest compact: llama3.2:3b β€” 2.0 GB
OllamaBeginner

Key Takeaways

  • βœ“Best general use: qwen3.5:9b β€” 6.6 GB download, the newest Qwen generation actually available in Ollama
  • βœ“Best coding: qwen2.5-coder:7b β€” 4.7 GB, and it is the value pick because qwen3-coder starts at 19 GB
  • βœ“Best compact: llama3.2:3b β€” 2.0 GB, or deepseek-r1:1.5b at 1.1 GB if you want reasoning traces
  • βœ“Watch the figure you are comparing: these are download sizes from the Ollama library, not VRAM requirements. Leave headroom above the download size for your context window
  • βœ“A model from six months ago with mature quantization often beats a brand-new release with limited community support

The Three Tier Leaders

For general use the current pick is qwen3.5:9b at a 6.6 GB download. Sizes below were read directly from the Ollama library on 28 August 2026.

"Best" in practice means the highest balance of output quality, inference speed, and memory efficiency β€” not raw benchmark score alone. A 9B model you can actually fit is more useful than a 30B model that swaps to disk.

One number worth correcting: coding models are cheaper than most guides suggest. qwen2.5-coder:7b is a 4.7 GB download and handles Python, TypeScript and Go without special prompting. The newer qwen3-coder starts at 19 GB, so it is a different class of hardware entirely, not a drop-in upgrade.

TierModelDownloadWhy It Leads
Compactllama3.2:3b2.0 GBBest quality-per-GB at the small end; 81.6M pulls
Generalqwen3.5:9b6.6 GBNewest Qwen generation available in Ollama
Codingqwen2.5-coder:7b4.7 GBStrong coding output at a fraction of qwen3-coder size
Reasoningdeepseek-r1:8b5.2 GBSecond most-pulled model on Ollama at 92M pulls

When Newer Isn't Better

A new model release does not automatically become the best Ollama pick. Quantization quality, community fine-tunes, and Ollama integration maturity take 4–8 weeks to catch up with a fresh release.

The pull counts make the point better than any benchmark. llama3.1 is still the most-downloaded model in the library at 118.9M pulls, and llama3.2 sits at 81.6M β€” both ahead of every newer arrival. That is not inertia; it is people picking the thing whose quantizations are well-optimized and whose behaviour is predictable across hardware.

If you want the newer generation, gemma4 (7.2 GB) and gpt-oss:20b (14 GB, 128K context) are both worth a look. Give a model 6+ weeks at the top before you rely on it in production. For a deeper look at evaluating models for your specific workload, see the top open-source models for Ollama.

Sizes verified against ollama.com/library on 28 August 2026. Model tags change often β€” run ollama pull and check the reported size before planning around any figure here.

Related Guides

Quick Answers About Ollama Models

Should I always use the newest Ollama model?β–Ύ
Not automatically. New releases need 4–8 weeks for community quantizations, fine-tunes, and Ollama integration to mature β€” which is why llama3.1 is still the most-pulled model in the library. Check the table above for current vetted picks. For CPU-only setups, see best Ollama models for CPU-only use.
Which Ollama model is best for coding right now?β–Ύ
qwen2.5-coder:7b at a 4.7 GB download. It handles Python, TypeScript and Go without special prompting and fits comfortably on an 8 GB card. Note that qwen3-coder, despite the higher version number, starts at 19 GB β€” so it is not a drop-in upgrade unless you have the hardware for it.
How much VRAM do I actually need?β–Ύ
The sizes quoted here are download sizes, not VRAM requirements. As a rule of thumb, allow a couple of GB above the download size for the context window and overhead. A 6.6 GB model like qwen3.5:9b is comfortable on a 12 GB card and workable on 8 GB with a modest context.
Are Qwen models better than Llama models in 2026?β–Ύ
For coding, the Qwen line leads β€” qwen2.5-coder and qwen3-coder have no direct Llama equivalent in the library. For general use it is closer than the version numbers suggest: llama3.1 and llama3.2 remain the two most-downloaded models on Ollama, ahead of every newer release, because their quantizations are mature and their behaviour is predictable.