Skip to main content
PromptQuorum
Home/The Complete Local LLM Software Directory: 224 Tools to Run AI on Your Own Hardware (2026)
Overview & Reference

The Complete Local LLM Software Directory: 224 Tools to Run AI on Your Own Hardware (2026)

·28 min read·By Hans Kuepper · Founder of PromptQuorum, multi-model AI dispatch tool · PromptQuorum

224 local LLM tools across 7 categories, each with its licence, price, and primary URL. Use the filters below to find the right one for your stack.

224 local LLM tools across 7 categories — Run & Serve, Chat & Assistants, Code & Development, Knowledge & Retrieval, Voice & Audio, Images & Video, and Train & Operate. Filter, search, and compare below.

224 tools

Fully local: 121Hybrid: 103Cloud: 0

Ollama

Inference engines
100% local

Easiest overall — one-command install, OpenAI-compatible API, huge model library

Runs its own engineFree
Depends on the model you run
Command lineDesktop appmacOSWindowsLinux
180,722283 articles
Feature Article

llama.cpp

Inference engines
100% local

Foundational C++ engine behind most other tools, runs anywhere including Apple Silicon

Runs its own engineFree
Depends on the model you run
Command lineLibrary / SDKmacOSWindowsLinux
126,800150 articles
Feature Article

vLLM

Inference engines
100% local

High-throughput serving for multi-user GPU deployments

Runs its own engineFree
Depends on the model you run
Command lineLibrary / SDKLinux
90,80091 articles
Feature Article

text-generation-webui

Inference engines
100% local

Power-user UI with extensive plugin ecosystem

Runs its own engineFree
Depends on the model you run
Web appCommand lineWindowsLinuxmacOS
47,60015 articles
Feature Article

exo

Inference engines
100% local

Distributed inference — run large models by pooling compute across multiple everyday devices

Runs its own engineFree
Depends on the model you run
Command linemacOSLinux
47,2602 articles
Feature Article

SGLang

Inference engines
100% local

Structured inference serving for agent pipelines

Runs its own engineFree
Depends on the model you run
Command lineLibrary / SDKLinux
33,5009 articles
Feature Article

Llamafile

Inference engines
100% local

Single-file portable LLM execution by Mozilla

Runs its own engineFree
Depends on the model you run
Command linemacOSWindowsLinux
25,9006 articles
Feature Article

MLC LLM

Inference engines
100% local

Mobile and edge device deployment runtime

Runs its own engineFree
Depends on the model you run
Library / SDKMobile appmacOSWindowsLinuxiOSAndroid
23,1346 articles
Feature Article

oMLX

Inference engines
100% local

MLX-based local inference server for Apple Silicon with paged SSD caching, built to power local coding agents

Runs its own engineFree
Claude CodeCursorOpenClaw
Depends on the model you run
Desktop appCommand linemacOS
21,8596 articles
Feature Article

TensorRT-LLM

Inference engines
100% local

NVIDIA-optimized inference for enterprise GPU rigs

Runs its own engineFree
Depends on the model you run
Command lineLibrary / SDKWindowsLinux
14,50010 articles
Feature Article

OpenLLM

Inference engines
Hybrid

Self-hostable inference server that runs open-source LLMs like DeepSeek and Llama behind an OpenAI-compatible API

Runs its own engineFree
LlamaDeepSeekQwen
Depends on the model you run
Command lineLibrary / SDK
12,5351 article
Feature Article

KoboldCpp

Inference engines
100% local

Lightweight llama.cpp wrapper with built-in UI

Runs its own engineFree
Depends on the model you run
Command lineWeb appWindowsLinuxmacOS
11,60013 articles
Feature Article

MLX-LM

Inference engines
100% local

Apple Silicon-native runtime by Apple research

Runs its own engineFree
Depends on the model you run
Command lineLibrary / SDKmacOS
8,90010 articles
Feature Article

NVIDIA Dynamo

Inference engines
100% local

Self-hosted, datacenter-scale distributed inference serving framework for large LLM deployments across multiple GPUs and nodes

Needs Ollama/LM StudioFree
vLLMTensorRT-LLMSGLangKubernetes
Depends on the model you run
Command lineLibrary / SDKLinux
8,1133 articles
Feature Article

LMDeploy

Inference engines
100% local

Self-hosted toolkit for compressing, quantizing, and serving LLMs with a high-throughput OpenAI-compatible inference engine

Runs its own engineFree
InternLMLlamaQwenBaichuanDeepSeek
Depends on the model you run
Command lineLibrary / SDKLinux
8,0805 articles
Feature Article

TurboFieldfare

Inference engines
100% local

Native Swift and Metal runtime that runs Gemma 4 26B-A4B locally in about 2 GB of RAM on any M-series MacBook

Runs its own engineFree
Gemma
Min. 2 GB RAM
Desktop appCommand linemacOS
6,7682 articles
Feature Article

Shimmy

Inference engines
100% local

Single-binary, pure-Rust OpenAI-compatible inference server that serves local GGUF and SafeTensors models with no Python or llama.cpp dependency

Runs its own engineFree
GGUFSafeTensors
Depends on the model you run
Command linemacOSWindowsLinux
5,8902 articles
Feature Article

ExLlamaV2

Inference engines
100% local

Fast quantized inference optimized for RTX GPUs

Runs its own engineFree
Min. 8 GB VRAM
Command lineLibrary / SDKWindowsLinux
4,6004 articles
Feature Article

LoRAX

Inference engines
100% local

Self-hosted inference server that serves thousands of fine-tuned LoRA adapters on a single GPU without a per-adapter cost

Own engine + externalFree
HuggingFacePEFTLudwigDocker
Depends on the model you run
Command lineLibrary / SDKLinux
3,8322 articles
Feature Article

Rapid-MLX

Inference engines
100% local

OpenAI-compatible local inference engine for Apple Silicon, built as a drop-in backend for Claude Code, Cursor, and Aider

Runs its own engineFree
Claude CodeCursorAiderCline
Min. 8 GB RAM
Command linemacOS
3,7734 articles
Feature Article

claude-code-local

Inference engines
100% local

MLX-native local server that lets Claude Code run 100% on-device on Apple Silicon, no cloud or API fees

Runs its own engineFree
Claude Code
Depends on the model you run
Command linemacOS
3,3141 article
Feature Article

Lucebox

Inference engines
100% local

Local LLM inference server with hand-tuned kernels and speculative decoding for specific consumer GPUs

Runs its own engineFree
Claude CodeCodexOpenCodeOpen WebUIGGUF
Depends on the model you run
Command lineLibrary / SDKLinux
2,8671 article
Feature Article

vllm-mlx

Inference engines
100% local

OpenAI- and Anthropic-compatible inference server bringing vLLM-style continuous batching to Apple Silicon via native MLX.

Runs its own engineFreeSupports MCP
Claude CodeOpenAI APIAnthropic API
Depends on the model you run
Command linemacOS
1,5803 articles
Feature Article

mlx-serve

Inference engines
100% local

Native Zig inference server for Apple Silicon serving MLX and GGUF models through OpenAI- and Anthropic-compatible APIs, with a bundled macOS menu-bar app.

Runs its own engineFreeSupports MCP
Claude CodeContinueCursorOpen WebUI
Depends on the model you run
Command lineDesktop appmacOS
1,3731 article
Feature Article

This page contains links to third-party products for reference. PromptQuorum is not enrolled in any affiliate program — these are plain links that earn no commission. Clicking links and your next steps are entirely your own responsibility. These links do not represent any endorsement or verification by PromptQuorum.

Common Real-World Stacks

For readers who do not want to read all seven categories, pick the closest stack and copy it. Each row pairs a real goal with a tested combination and the hardware floor it actually runs on.

Goal
Stack
Hardware floor
Just chat casuallyLM Studio standalone16 GB RAM, no GPU
Best balance for power usersOllama + Open WebUI16 GB RAM, optional GPU
Document chatOllama + AnythingLLM16 GB RAM, optional GPU
CodingOllama + Cline16 GB RAM + GPU recommended
Roleplay / creativeKoboldCpp + SillyTavern16 GB RAM, GPU recommended
Privacy-first businessOllama + Open WebUI + PrivateGPT32 GB RAM + 12 GB VRAM
Mobile / on-the-goMLC Chat or PocketPal AIiPhone 13+ / Pixel 7+
Apple SiliconOllama (MLX backend) or LM StudioM2/M3/M4/M5 with 16+ GB unified
Multi-user teamvLLM + Open WebUI32+ GB RAM + multi-GPU
Image generationStable Diffusion + ComfyUI or Invoke AI6+ GB VRAM GPU
Voice assistantOllama + Whisper.cpp + Piper TTS8 GB RAM, CPU-only possible
10+ common real-world local AI stacks by goal: from LM Studio standalone (16 GB RAM, no GPU) to vLLM + Open WebUI for multi-user teams (32 GB RAM + multi-GPU), Stable Diffusion for images (6 GB VRAM), and Ollama + Whisper + Piper for fully-offline voice assistants. Ollama + Open WebUI is the best-balance default at 16 GB RAM.
10+ common real-world local AI stacks by goal: from LM Studio standalone (16 GB RAM, no GPU) to vLLM + Open WebUI for multi-user teams (32 GB RAM + multi-GPU), Stable Diffusion for images (6 GB VRAM), and Ollama + Whisper + Piper for fully-offline voice assistants. Ollama + Open WebUI is the best-balance default at 16 GB RAM.

Key Takeaways

  • 7 categories, 224 projects, one map. Run & Serve, Chat & Assistants, Code & Development, Knowledge & Retrieval, Voice & Audio, Images & Video, and Train & Operate — most popular projects in 2026 fit in one primary category, and many span more than one.
  • Pick a runtime first. Ollama is the right default for ~95% of readers; llama.cpp is the foundational engine underneath most other tools; vLLM is the production-serving pick for multi-user setups.
  • Most categories above Run & Serve are optional. A desktop app OR a web UI is enough for chat. Add a code assistant or CLI tool only when you want code help; add a RAG system only when you want to chat with your own documents; add an agent framework only when one-shot calls stop being enough; add Images & Video only when you need visual output.
  • Licence matters for commercial use. MIT and Apache 2.0 dominate the ecosystem. AGPL appears on a handful of UIs (text-generation-webui, KoboldCpp, Jan, SillyTavern) — fine for personal use, more deliberate for commercial deployments. The "License" column below names every one explicitly; see AI & open-source software licenses explained for what MIT, Apache-2.0, AGPL-3.0, and the other license types below actually require.
  • Multi-tool stacks are normal. Ollama + Open WebUI + AnythingLLM + Cline + Stable Diffusion is a single-machine setup that covers chat, RAG, coding, and image generation without compromise. The "Common Real-World Stacks" table below names the recipes that actually work in 2026.
The 7 categories of a local LLM stack: 224 actively-maintained projects spanning Run & Serve (Ollama, llama.cpp, vLLM), Chat & Assistants (LM Studio, Jan, GPT4All), Code & Development (Cline, LangChain, CrewAI), Knowledge & Retrieval (AnythingLLM, PrivateGPT), Voice & Audio (Whisper.cpp, Piper), Images & Video (Stable Diffusion, ComfyUI), and Train & Operate.
The 7 categories of a local LLM stack: 224 actively-maintained projects spanning Run & Serve (Ollama, llama.cpp, vLLM), Chat & Assistants (LM Studio, Jan, GPT4All), Code & Development (Cline, LangChain, CrewAI), Knowledge & Retrieval (AnythingLLM, PrivateGPT), Voice & Audio (Whisper.cpp, Piper), Images & Video (Stable Diffusion, ComfyUI), and Train & Operate.

How This Directory Stays Current

This directory is reviewed every three months, with focused monthly updates in between reviews. Recent expansions added dozens of new tools across every category, split voice/multimodal into three focused categories (speech-to-text, text-to-speech, vision), split coding assistants into IDE integrations and terminal tools, and added an entirely new Images & Video category. All links and licenses were reverified; new entries (PearAI, Windsurf, Sourcegraph Cody, SuperAGI, Leon AI, Draw Things, Fooocus, StableSwarmUI, and others) were validated for active maintenance. Inclusion criteria: project is actively maintained (commits in the last 90 days), has a verifiable open-source licence or a clear commercial-use statement, and either holds meaningful user share in 2026 or fills a category that would otherwise be empty. Projects that go inactive for more than two release cycles are removed; new entrants that pass the criteria are added at the next review. To suggest a project for inclusion, open an issue or PR against the PromptQuorum repository — include the project URL, licence, and a one-sentence description in the format above.

Sources

Frequently Asked Questions

What is the difference between a local LLM runtime and a desktop app?

A runtime (Ollama, llama.cpp, vLLM) is the engine that loads model weights and serves an API — typically OpenAI-compatible. A desktop app (LM Studio, Jan, GPT4All) is a chat UI that calls a runtime under the hood. Some apps bundle their own runtime (LM Studio embeds llama.cpp), others require you to install a runtime separately (Open WebUI calls Ollama). The runtime decides what is possible; the app decides what is convenient.

Can I use multiple tools from this list at the same time?

Yes — most stacks combine 2-4 tools. A common setup: Ollama as the runtime, Open WebUI for chat, AnythingLLM for document chat, and Cline for coding — all four run against the same Ollama instance on a single machine. The "Common Real-World Stacks" table above lists the recipes that work without conflict.

Which tools work fully offline with no telemetry?

Ollama, llama.cpp, vLLM, Jan, GPT4All, Open WebUI, AnythingLLM, PrivateGPT, Cline, Aider, KoboldCpp, Llamafile, MLX-LM, and most of the AGPL/MIT-licensed apps in this directory work fully offline once the model is downloaded. LM Studio and several closed-source tools have optional analytics that can be disabled in settings — verify by running a packet capture once after install. Browser-based UIs (Open WebUI, LibreChat) are local-only when configured to use a local backend.

Are any of these commercial-licensed (not free for commercial use)?

A handful: LM Studio, Msty, Backyard AI, Layla, and Cursor are closed-source — generally free to use but not redistributable, and commercial terms vary. Private LLM is paid. AGPL-licensed tools (Jan, KoboldCpp, text-generation-webui, SillyTavern, Khoj, Copilot for Obsidian) are free for any use including commercial, but the AGPL terms require source disclosure if you modify and host them publicly. Apache 2.0 and MIT projects (the majority) are usable in any context including commercial without attribution constraints beyond the licence text.

Which tools support Apple Silicon (M-series chips) natively?

Ollama, llama.cpp, MLX-LM, LM Studio, Jan, Enchanted, GPT4All, MLC Chat, AnythingLLM, and most Electron/Tauri apps run natively on Apple Silicon and use the Metal backend. MLX-LM is Apple-specific and the fastest for large models on M-series. vLLM, TensorRT-LLM, and ExLlamaV2 are NVIDIA-focused and either do not run or run poorly on Apple Silicon — for Apple users, Ollama with the Metal backend is the default.

Do all these tools support GGUF model format?

GGUF is the native format for llama.cpp and any tool that wraps it (Ollama, LM Studio, Jan, GPT4All, KoboldCpp, Llamafile). vLLM and TensorRT-LLM use their own optimised formats (typically AWQ or FP16) for higher throughput. ExLlamaV2 uses EXL2 quantisation. MLX-LM uses MLX-converted weights. Most listed tools accept GGUF; a few (vLLM, TensorRT-LLM, ExLlamaV2, MLX-LM) require a one-time conversion step from the original Hugging Face weights.

Which tools are best for users with no coding experience?

GPT4All has the simplest install (one click, runs on 8 GB RAM), though its own updates have slowed recently. LM Studio is the most feature-rich without requiring a terminal. Jan is the most privacy-conscious of the no-code options. For document chat without command-line work, AnythingLLM is the easiest. All four are listed in the Desktop GUI Apps category above.

Can I run these tools on a server and access them remotely?

Most server-capable tools (Ollama, vLLM, LocalAI, Open WebUI, LibreChat, PrivateGPT, AnythingLLM) expose an HTTP API and bind to a network interface configurable in settings. Standard pattern: run Ollama on a home server or VPS, run a UI on your laptop or phone pointing at the server's IP. Treat the API like any web service — bind to localhost behind a reverse proxy, or to a private network with proper authentication. Open WebUI ships with multi-user support out of the box.

Which tools support multi-user / team setups?

Open WebUI, LibreChat, h2oGPT, AnythingLLM (with admin features enabled), and Dify are designed for multi-user use, with role-based access and per-user conversation history. vLLM is the right serving layer underneath when concurrent inference matters — it batches requests across users for throughput unattainable on Ollama at concurrency above ~3.

How often does this directory get updated?

Every three months, with focused monthly updates in between. Monthly changes (a project goes inactive, a new tool gains meaningful share, a licence changes) get patched into the existing entry. Entirely new categories (like the Images & Video category) are added during the quarterly reviews to keep the structure stable. See the "Last updated" date at the top of this page for the most recent refresh. The "Sources" section above lists the community indexes used to spot-check what the ecosystem is actually doing between refreshes.

Can I do image generation locally without cloud calls?

Yes — Stable Diffusion, ComfyUI, Invoke AI, AUTOMATIC1111 WebUI, and others in the Images & Video category run entirely on local hardware. Stable Diffusion needs 6+ GB VRAM (RTX 3060, RTX 4060, or equivalent); Fooocus and other optimized UIs can run on cards with 4-6 GB. Real-ESRGAN upscales generated images; ControlNet adds spatial control (edges, poses, depth maps); AnimateDiff generates video from text. All run without sending data to external servers.

Browse the slides below or download as PDF for offline reference. Download Reference Card (PDF)

About this data

This directory is compiled with AI-assisted research from public sources (project READMEs, official sites, GitHub repositories). It is not exhaustive and may contain errors — hardware requirements, pricing, and platform support change frequently and some fields have not been independently verified yet.

Most entries carry no status badge: they are listed from public sources but not independently reviewed. A "Verified" badge means core facts (license, platforms, pricing) have been manually checked. A "PromptQuorum-tested" badge is only shown for tools PromptQuorum has actually installed and run — most entries do not have this badge yet.

The "PromptQuorum-tested" badge is reserved for tools we have personally installed and run — an untested tool never carries it, regardless of popularity or stars.

Spotted something wrong? Email hello@promptquorum.com and we will correct it.

← Back to Power Local LLM