Key Takeaways
- 46 tools, three kinds: inference engines (30), runtimes and managers (13), routers and gateways (4). Ollama appears in two groups because it is both an engine and a runtime.
- The table is generated from each tool's record and checked against its official README or site; a dash means "not stated in the documentation", never "no". For some well-known tools the documentation quoted here is silent on a given feature, so their cells show a dash.
- Every tool name in the table links to its own PromptQuorum review, which is where installation steps and limits are covered.
📍 In One Sentence
Running models locally involves three kinds of tool — inference engines, runtimes and managers, and routers and gateways — so the 46 tools in the PromptQuorum directory are compared within each kind, using a table generated from the same tool data as each tool's own review.
💬 In Plain Terms
An engine is the part that actually runs the model, a runtime or manager downloads and runs models for you, and a gateway routes requests between them. Comparing an engine with a gateway on GPU support makes no sense, so this guide compares like with like.
How We Compared
Each tool's facts — price, license, platforms, hardware needs and category-specific attributes — are stored once, in that tool's directory record. The comparison table below is generated from those records, and the tool's own review draws on the same record, so the two cannot state different values.
Category-specific attributes (for example OpenAI-compatible API or AMD GPU support) were taken from each project's official README or website and checked against the exact wording there. Where the documentation is silent, the table shows a dash rather than guessing; where a claim is qualified (experimental, planned, or available only through a separate project), the attribute is left out of the table and covered in the tool's review instead.
Only tools with their own PromptQuorum review are in the table. Tools listed only as API servers (h2oGPT, Tabby and OpenAI Edge TTS) are compared in their own categories. The comparison lists tools that run on your own hardware; it does not rank them, because the right one depends on your constraint.
Comparison Table
Choose a kind of tool below, then read across a row. Click a tool name to open its full PromptQuorum review.
| Tool | Price | License | Platforms | Runs | Hardware | Version | OpenAI-compatible API | NVIDIA GPU | Apple Silicon | AMD GPU | CPU inference | Multi-GPU / multi-node | Review | product link · disclosed |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| candle-vllm | Free | MIT | macOS, Windows, Linux | Local | Varies by model | v0.9.1 | Yes | Yes | Yes | — | — | Yes | Read review → | candle-vllm |
| claude-code-local | Free | MIT | macOS | Local | Varies by model | v0.3.0 | — | — | Yes | — | — | — | Read review → | claude-code-local |
| ExLlamaV2 | Free | MIT | Windows, Linux | Local | 8 GB VRAM | v0.3.2 | — | Yes | — | — | — | Yes | Read review → | ExLlamaV2 |
| exo | Free | Apache-2.0 | macOS, Linux | Local | Varies by model | — | Yes | — | Yes | — | Yes | Yes | Read review → | exo |
| KoboldCpp | Free | AGPL-3.0 | Windows, Linux, macOS | Local | Varies by model | v1.121 | Yes | Yes | Yes | Yes | Yes | — | Read review → | KoboldCpp |
| KServe | Free | Apache-2.0 | Linux | Local | CPU is enough | v0.20.0 | Yes | — | — | — | — | Yes | Read review → | KServe |
| llama.cpp | Free | MIT | macOS, Windows, Linux | Local | Varies by model | v0.4.1 | Yes | Yes | Yes | Yes | Yes | — | Read review → | llama.cpp |
| Llamafile | Free | Apache-2.0 | macOS, Windows, Linux | Local | Varies by model | 0.10.6 | — | — | — | — | — | — | Read review → | Llamafile |
| LMDeploy | Free | Apache-2.0 | Linux | Local | Varies by model | v0.17.0 | — | Yes | — | — | — | Yes | Read review → | LMDeploy |
| LocalAI | Free | MIT | macOS, Windows, Linux | Local | Varies by model | — | Yes | Yes | Yes | Yes | Yes | Yes | Read review → | LocalAI |
| LoRAX | Free | Apache-2.0 | Linux | Local | Varies by model | v0.12.1 | Yes | Yes | — | — | — | Yes | Read review → | LoRAX |
| Lucebox | Free | Apache-2.0 | Linux | Local | Varies by model | — | Yes | Yes | — | Yes | — | Yes | Read review → | Lucebox |
| MLC LLM | Free | Apache-2.0 | macOS, Windows, Linux, iOS, Android | Local | Varies by model | — | Yes | Yes | Yes | Yes | — | — | Read review → | MLC LLM |
| MLX-LM | Free | MIT | macOS | Local | Varies by model | — | — | — | Yes | — | — | Yes | Read review → | MLX-LM |
| mlx-serve | Free | MIT | macOS | Local | Varies by model | v26.9.4 | Yes | — | Yes | — | — | — | Read review → | mlx-serve |
| mlxcel | Free | Apache-2.0 | macOS, Linux | Local | Varies by model | v0.7.0 | Yes | Yes | Yes | — | — | Yes | Read review → | mlxcel |
| NVIDIA Dynamo | Free | Apache-2.0 | Linux | Local | Varies by model | v1.4–v1.6 | Yes | — | — | — | — | Yes | Read review → | NVIDIA Dynamo |
| Ollama | Free | MIT | macOS, Windows, Linux | Local | Varies by model | — | — | — | — | — | — | — | Read review → | Ollama |
| OlliteRT | Free | Apache-2.0 | Android | Local | CPU is enough | — | Yes | — | — | — | Yes | — | Read review → | OlliteRT |
| oMLX | Free | Apache-2.0 | macOS | Local | Varies by model | v0.6.4 | Yes | — | Yes | — | — | Yes | Read review → | oMLX |
| OpenLLM | Free | Apache-2.0 | — | Hybrid | Varies by model | — | Yes | — | — | — | — | — | Read review → | OpenLLM |
| Rapid-MLX | Free | Apache-2.0 | macOS | Local | 8 GB RAM | — | Yes | — | Yes | — | — | — | Read review → | Rapid-MLX |
| SGLang | Free | Apache-2.0 | Linux | Local | Varies by model | — | — | Yes | — | Yes | Yes | Yes | Read review → | SGLang |
| Shimmy | Free | Apache-2.0 | macOS, Windows, Linux | Local | Varies by model | — | Yes | Yes | Yes | Yes | — | — | Read review → | Shimmy |
| SwiftLM | Free | MIT | macOS, iOS | Local | Varies by model | — | Yes | — | Yes | — | — | — | Read review → | SwiftLM |
| TensorRT-LLM | Free | Apache-2.0 | Windows, Linux | Local | Varies by model | — | — | Yes | — | — | — | — | Read review → | TensorRT-LLM |
| text-generation-webui | Free | AGPL-3.0 | Windows, Linux, macOS | Local | Varies by model | — | Yes | Yes | Yes | Yes | Yes | — | Read review → | text-generation-webui |
| TurboFieldfare | Free | Apache-2.0 | macOS | Local | 2 GB RAM | — | — | — | Yes | — | — | — | Read review → | TurboFieldfare |
| vLLM | Free | Apache-2.0 | Linux | Local | Varies by model | — | Yes | Yes | — | Yes | Yes | Yes | Read review → | vLLM |
| vllm-mlx | Free | Apache-2.0 | macOS | Local | Varies by model | — | Yes | — | Yes | — | — | — | Read review → | vllm-mlx |
"—" means the project's own documentation does not state it, not that the feature is missing. Values come from each project's official README or site and are re-checked when a tool's review is updated.
Inference Engines: What Differs
- OpenAI-compatible API. candle-vllm, NVIDIA Dynamo, exo, KoboldCpp, KServe, llama.cpp, LocalAI, LoRAX, Lucebox, MLC LLM, mlx-serve, mlxcel, OlliteRT, oMLX, OpenLLM, Rapid-MLX, Shimmy, SwiftLM, text-generation-webui, vllm-mlx and vLLM document an OpenAI-compatible HTTP API, so apps written for the OpenAI API can point at them.
- NVIDIA GPUs. candle-vllm, ExLlamaV2, KoboldCpp, llama.cpp, LMDeploy, LocalAI, LoRAX, Lucebox, MLC LLM, mlxcel, SGLang, Shimmy, TensorRT-LLM, text-generation-webui and vLLM document NVIDIA GPU (CUDA) support.
- Apple Silicon. candle-vllm, claude-code-local, exo, KoboldCpp, llama.cpp, LocalAI, MLC LLM, MLX-LM, mlx-serve, mlxcel, oMLX, Rapid-MLX, Shimmy, SwiftLM, text-generation-webui, TurboFieldfare and vllm-mlx document Apple Silicon, Metal or MLX support.
- AMD GPUs. KoboldCpp, llama.cpp, LocalAI, Lucebox, MLC LLM, SGLang, Shimmy, text-generation-webui and vLLM document AMD GPU support.
- CPU inference. exo, KoboldCpp, llama.cpp, LocalAI, OlliteRT, SGLang, text-generation-webui and vLLM document running inference on a CPU.
- Multi-GPU and multi-node. candle-vllm, NVIDIA Dynamo, ExLlamaV2, exo, KServe, LMDeploy, LocalAI, LoRAX, Lucebox, MLX-LM, mlxcel, oMLX, SGLang and vLLM document multi-GPU, tensor-parallel or multi-node inference.
- License. Most engines here are Apache-2.0 (19 tools) or MIT (9); KoboldCpp and text-generation-webui are AGPL-3.0. Copyleft licenses attach conditions to distributing modified versions — see AI Tool Licenses Explained.
Runtimes and Managers: What Differs
- OpenAI-compatible API. Docker Model Runner, Foundry Local, GPUStack, Jan, Lemonade and Osaurus document an OpenAI-compatible API.
- Desktop app. GPT4All, Jan, LM Studio, Osaurus and Ypipe document an installable desktop app.
- Built-in model library. Foundry Local, GPUStack, Jan, Lemonade, LM Studio and Ollama document a built-in way to find and download models.
- Headless or server mode. DreamServer, GPUStack, Lemonade and Osaurus document running as a background service or server without the GUI.
- License and price. Four runtimes here are Apache-2.0 and four are MIT. LM Studio and Docker Model Runner are proprietary (LM Studio is free to use; Docker Model Runner is bundled with Docker Desktop), Msty is closed source with a free tier, and RunAnywhere and YPipe use their own or undocumented terms — check each review.
Routers and Gateways: What Differs
- OpenAI-compatible endpoint. AIClient2API and litellm document an OpenAI-compatible endpoint.
- Routing to local models. litellm documents local or self-hosted models (such as Ollama) among its supported providers.
- Fallback and load balancing. AIClient2API, ClawRouter and litellm document fallback, retry or load balancing across models or providers.
- License. Two of the four are MIT, one is GPL-3.0 and one is Apache-2.0.
What This Comparison Cannot Tell You
- It compares documented capabilities, not performance. It says nothing about tokens per second, memory use or latency — PromptQuorum has not measured them for these tools, and they depend heavily on your hardware and model.
- Dashes are gaps in the documentation we checked, not negative findings. Some tools may support a feature their README does not mention; this is most visible for a few widely used tools whose short READMEs state little (for example Ollama's engine features).
- Hardware support is what the documentation states, not a guarantee of a good experience on that hardware; read the tool's review for the real requirements.
- Tools change quickly. Each tool's review states the version it was checked against, and this guide is refreshed when a review is.
Frequently Asked Questions
What is the difference between an inference engine, a runtime and a gateway?
An inference engine executes the model on your hardware (for example llama.cpp or vLLM). A runtime or manager downloads models and runs them for you, often with a desktop app or a model library (for example Ollama or LM Studio). A router or gateway sits in front of one or more models or providers and forwards requests. They do different jobs, so they are compared separately.
What does a dash in the comparison table mean?
It means the project's own documentation does not state that attribute. It does not mean the feature is missing; check the tool's review or its repository.
Why is Ollama in two groups?
Ollama is both an inference engine and a runtime that manages models, so it is listed under both. Its README states few of the attributes compared here, so many of its cells show a dash.
Do any of these tools have an affiliate link?
No. PromptQuorum has no affiliate relationship with any tool in this comparison at the time of writing, and no link here earns a commission.
How often is this comparison updated?
It is refreshed twice a year and whenever one of the listed tools' reviews is updated, because the table is generated from the same data as those reviews.
Sources
- Each tool's official README or website, listed in that tool's PromptQuorum review (linked from the comparison table).
- PromptQuorum local AI app directory — the record each row of the table is generated from.
- AI Tool Licenses Explained — what the license families named above mean.