Skip to main content
PromptQuorum

Best Mini PC for an Always-On Ollama Server (2026)

Best Mini PC for an Always-On Ollama Server (2026)

This page contains links to third-party products for reference. PromptQuorum is not enrolled in any affiliate program — these are plain links that earn no commission. Clicking links and your next steps are entirely your own responsibility. These links do not represent any endorsement or verification by PromptQuorum.

Hardware & PerformanceIntermediate

Key Takeaways

  • Mini PCs draw 15–45 W vs 200–350 W for desktop GPUs — 24/7 savings matter
  • UM890 Pro runs 7B models CPU-only at 12–18 tok/s; fine for API server use
  • AOOSTAR GEM12 Pro + OCuLink eGPU unlocks GPU acceleration without a desktop PC
  • Mac Mini M4 Pro: 48 GB unified memory runs 32B models — best macOS option
  • Beelink SER8 is the <$400 starting point — 32 GB RAM for 7B and 13B

Best Mini PCs for Always-On Ollama Server — Ranked

1

Minisforum UM890 Pro — Best Overall

The Minisforum UM890 Pro runs the AMD Ryzen 9 8945HS (8-core, up to 5.2 GHz) and supports up to 96 GB DDR5-5600 dual-channel RAM — enough to load Llama 3.3 70B at Q4 entirely in RAM. The Radeon 780M iGPU (12 RDNA3 CUs) accelerates 7B and 13B models at 8–14 tok/s via ROCm. CPU-only 7B Q4 speed: ~12–18 tok/s. Idle power: ~15 W. Load (GPU active): ~35–45 W. Price: ~$500 (32 GB/1 TB) to $600 (64 GB).

Minisforum UM890 Pro on Amazonproduct link · disclosedMinisforum UM890 Pro on Minisforum.comproduct link · disclosed
2

AOOSTAR GEM12 Pro OCuLink — Best for eGPU

The AOOSTAR GEM12 Pro OCuLink adds an OCuLink port that connects to an external GPU at PCIe 4.0 x4 bandwidth (64 Gbps) — enough to run an RTX 3090 at full speed for Ollama. Without eGPU: same as other mini PCs, 13–18 tok/s CPU-only. With RTX 3090 eGPU: 60–80 tok/s on 7B Q4. The mini PC itself runs AMD Ryzen 7 Pro 8845HS with 32–96 GB DDR5. OCuLink to PCIe adapter required (~$30). Barebone price: ~$380; configured (32 GB/1 TB): ~$480.

AOOSTAR GEM12 Pro OCuLink on Amazonproduct link · disclosed
3

Beelink SER8 — Best Budget Pick

The Beelink SER8 runs the AMD Ryzen 7 8745HS (8-core, up to 4.9 GHz) with 32 GB DDR5 and 1 TB NVMe SSD for ~$450. CPU-only Ollama speed: ~10–15 tok/s on 7B Q4. Idle power: 10–15 W. RAM is user-upgradeable (not soldered on current 8745HS revision). If you want a solid entry into always-on Ollama without spending $500+, the SER8 covers 7B and 13B models well.

Beelink SER8 on Amazonproduct link · disclosed
4

Apple Mac Mini M4 Pro — Best for macOS

The Mac Mini M4 Pro (24-core GPU, 48 GB unified memory, ~$1399) is the only mini PC that runs 32B models at GPU speed out of the box. Ollama on Apple Silicon uses Metal, not CUDA — the 48 GB unified memory loads Qwen3 32B Q4 (~18 GB) and runs at 20–30 tok/s. Power: 18–30 W under Ollama load. Ideal for macOS users who want a silent, always-on home server that doubles as a desk machine. The price premium is real but justified if 32B performance matters.

Apple Mac Mini M4 Pro on Amazonproduct link · disclosedApple Mac Mini M4 Pro on Apple.comproduct link · disclosed

Always-On Electricity Cost Comparison

At $0.15/kWh (US average), running 24/7 for 30 days:

DeviceAvg Load PowerMonthly Cost (24/7)
Minisforum UM890 Pro35 W~$3.78/mo
Beelink SER825 W~$2.70/mo
Mac Mini M4 Pro25 W~$2.70/mo
Desktop RTX 4060 Ti (comparison)200 W~$21.60/mo
Cloud API (GPT-5.5-mini, 1M tok/day)N/A~$45–90/mo

Related Guides

Quick Answers

Can a mini PC run 13B or larger models at useful speed?
Yes — with enough RAM. The Minisforum UM890 Pro with 64 GB runs Llama 3.3 13B Q8 entirely in RAM at ~8–12 tok/s CPU-only. With the Radeon 780M iGPU accelerating, Q4 models run at 10–18 tok/s — usable for background summarization or API calls. Interactive chat benefits from at least 12–15 tok/s. For 30B+ models, the Mac Mini M4 Pro (48 GB unified memory) is the only mini PC option under $1500.
Does Ollama work well as a network server on a mini PC?
Yes. Set OLLAMA_HOST=0.0.0.0 in your environment and Ollama serves requests from any device on your LAN. Pair with Open WebUI (Docker container) for a browser-based interface accessible from phones, tablets, and PCs. The mini PC draws low power, runs silently, and handles one concurrent request at a time without issue.
What about eGPU setups — are they worth it?
For Ollama specifically, an OCuLink eGPU (AOOSTAR GEM12 Pro + RTX 3090 enclosure) is the best of both worlds: desktop GPU speed with mini PC power draw when idle. OCuLink (PCIe 4.0 x4) delivers ~80% of the bandwidth of a direct PCIe x16 slot — enough for LLM inference with minimal bottleneck. Thunderbolt eGPUs are slower (~40% bandwidth) and not recommended for GPU-intensive inference.

Want the full breakdown?

Read the complete guide →