Skip to main content
PromptQuorum
Home/Power Local LLM/Local Inference Engines, Runtimes & Gateways Compared (2026): Run and Serve Models
Overview & Reference

Local Inference Engines, Runtimes & Gateways Compared (2026): Run and Serve Models

·10 min read·By Hans Kuepper · Founder of PromptQuorum, multi-model AI dispatch tool · PromptQuorum

The 46 local run-and-serve tools in the PromptQuorum directory split into three kinds that should be compared separately: inference engines (30 tools), runtimes and managers (13) and routers and gateways (4). Within engines, 15 document NVIDIA GPU support, 17 document Apple Silicon support and 21 document an OpenAI-compatible API; among runtimes, 6 document an OpenAI-compatible API. Use the comparison table below, and read each tool's own review before you install it.

Running a model on your own hardware involves three different kinds of tool — inference engines that execute the model, runtimes and managers that download and run models for you, and routers and gateways that sit in front of them — and no single feature list compares them fairly. This guide compares 46 free and freemium tools, one kind at a time, using a comparison table generated from the same data as each tool's own PromptQuorum review, so the table and the reviews cannot disagree.

This page contains links to third-party products for reference. PromptQuorum is not enrolled in any affiliate program — these are plain links that earn no commission. Clicking links and your next steps are entirely your own responsibility. These links do not represent any endorsement or verification by PromptQuorum.

Key Takeaways

  • 46 tools, three kinds: inference engines (30), runtimes and managers (13), routers and gateways (4). Ollama appears in two groups because it is both an engine and a runtime.
  • The table is generated from each tool's record and checked against its official README or site; a dash means "not stated in the documentation", never "no". For some well-known tools the documentation quoted here is silent on a given feature, so their cells show a dash.
  • Every tool name in the table links to its own PromptQuorum review, which is where installation steps and limits are covered.

📍 In One Sentence

Running models locally involves three kinds of tool — inference engines, runtimes and managers, and routers and gateways — so the 46 tools in the PromptQuorum directory are compared within each kind, using a table generated from the same tool data as each tool's own review.

💬 In Plain Terms

An engine is the part that actually runs the model, a runtime or manager downloads and runs models for you, and a gateway routes requests between them. Comparing an engine with a gateway on GPU support makes no sense, so this guide compares like with like.

How We Compared

Each tool's facts — price, license, platforms, hardware needs and category-specific attributes — are stored once, in that tool's directory record. The comparison table below is generated from those records, and the tool's own review draws on the same record, so the two cannot state different values.

Category-specific attributes (for example OpenAI-compatible API or AMD GPU support) were taken from each project's official README or website and checked against the exact wording there. Where the documentation is silent, the table shows a dash rather than guessing; where a claim is qualified (experimental, planned, or available only through a separate project), the attribute is left out of the table and covered in the tool's review instead.

Only tools with their own PromptQuorum review are in the table. Tools listed only as API servers (h2oGPT, Tabby and OpenAI Edge TTS) are compared in their own categories. The comparison lists tools that run on your own hardware; it does not rank them, because the right one depends on your constraint.

Comparison Table

Choose a kind of tool below, then read across a row. Click a tool name to open its full PromptQuorum review.

ToolPriceLicensePlatformsRunsHardwareVersionOpenAI-compatible APINVIDIA GPUApple SiliconAMD GPUCPU inferenceMulti-GPU / multi-nodeReviewproduct link · disclosed
candle-vllmFreeMITmacOS, Windows, LinuxLocalVaries by modelv0.9.1YesYesYesYesRead reviewcandle-vllm
claude-code-localFreeMITmacOSLocalVaries by modelv0.3.0YesRead reviewclaude-code-local
ExLlamaV2FreeMITWindows, LinuxLocal8 GB VRAMv0.3.2YesYesRead reviewExLlamaV2
exoFreeApache-2.0macOS, LinuxLocalVaries by modelYesYesYesYesRead reviewexo
KoboldCppFreeAGPL-3.0Windows, Linux, macOSLocalVaries by modelv1.121YesYesYesYesYesRead reviewKoboldCpp
KServeFreeApache-2.0LinuxLocalCPU is enoughv0.20.0YesYesRead reviewKServe
llama.cppFreeMITmacOS, Windows, LinuxLocalVaries by modelv0.4.1YesYesYesYesYesRead reviewllama.cpp
LlamafileFreeApache-2.0macOS, Windows, LinuxLocalVaries by model0.10.6Read reviewLlamafile
LMDeployFreeApache-2.0LinuxLocalVaries by modelv0.17.0YesYesRead reviewLMDeploy
LocalAIFreeMITmacOS, Windows, LinuxLocalVaries by modelYesYesYesYesYesYesRead reviewLocalAI
LoRAXFreeApache-2.0LinuxLocalVaries by modelv0.12.1YesYesYesRead reviewLoRAX
LuceboxFreeApache-2.0LinuxLocalVaries by modelYesYesYesYesRead reviewLucebox
MLC LLMFreeApache-2.0macOS, Windows, Linux, iOS, AndroidLocalVaries by modelYesYesYesYesRead reviewMLC LLM
MLX-LMFreeMITmacOSLocalVaries by modelYesYesRead reviewMLX-LM
mlx-serveFreeMITmacOSLocalVaries by modelv26.9.4YesYesRead reviewmlx-serve
mlxcelFreeApache-2.0macOS, LinuxLocalVaries by modelv0.7.0YesYesYesYesRead reviewmlxcel
NVIDIA DynamoFreeApache-2.0LinuxLocalVaries by modelv1.4–v1.6YesYesRead reviewNVIDIA Dynamo
OllamaFreeMITmacOS, Windows, LinuxLocalVaries by modelRead reviewOllama
OlliteRTFreeApache-2.0AndroidLocalCPU is enoughYesYesRead reviewOlliteRT
oMLXFreeApache-2.0macOSLocalVaries by modelv0.6.4YesYesYesRead reviewoMLX
OpenLLMFreeApache-2.0HybridVaries by modelYesRead reviewOpenLLM
Rapid-MLXFreeApache-2.0macOSLocal8 GB RAMYesYesRead reviewRapid-MLX
SGLangFreeApache-2.0LinuxLocalVaries by modelYesYesYesYesRead reviewSGLang
ShimmyFreeApache-2.0macOS, Windows, LinuxLocalVaries by modelYesYesYesYesRead reviewShimmy
SwiftLMFreeMITmacOS, iOSLocalVaries by modelYesYesRead reviewSwiftLM
TensorRT-LLMFreeApache-2.0Windows, LinuxLocalVaries by modelYesRead reviewTensorRT-LLM
text-generation-webuiFreeAGPL-3.0Windows, Linux, macOSLocalVaries by modelYesYesYesYesYesRead reviewtext-generation-webui
TurboFieldfareFreeApache-2.0macOSLocal2 GB RAMYesRead reviewTurboFieldfare
vLLMFreeApache-2.0LinuxLocalVaries by modelYesYesYesYesYesRead reviewvLLM
vllm-mlxFreeApache-2.0macOSLocalVaries by modelYesYesRead reviewvllm-mlx

"—" means the project's own documentation does not state it, not that the feature is missing. Values come from each project's official README or site and are re-checked when a tool's review is updated.

Inference Engines: What Differs

Runtimes and Managers: What Differs

Routers and Gateways: What Differs

  • OpenAI-compatible endpoint. AIClient2API and litellm document an OpenAI-compatible endpoint.
  • Routing to local models. litellm documents local or self-hosted models (such as Ollama) among its supported providers.
  • Fallback and load balancing. AIClient2API, ClawRouter and litellm document fallback, retry or load balancing across models or providers.
  • License. Two of the four are MIT, one is GPL-3.0 and one is Apache-2.0.

What This Comparison Cannot Tell You

  • It compares documented capabilities, not performance. It says nothing about tokens per second, memory use or latency — PromptQuorum has not measured them for these tools, and they depend heavily on your hardware and model.
  • Dashes are gaps in the documentation we checked, not negative findings. Some tools may support a feature their README does not mention; this is most visible for a few widely used tools whose short READMEs state little (for example Ollama's engine features).
  • Hardware support is what the documentation states, not a guarantee of a good experience on that hardware; read the tool's review for the real requirements.
  • Tools change quickly. Each tool's review states the version it was checked against, and this guide is refreshed when a review is.

Frequently Asked Questions

What is the difference between an inference engine, a runtime and a gateway?

An inference engine executes the model on your hardware (for example llama.cpp or vLLM). A runtime or manager downloads models and runs them for you, often with a desktop app or a model library (for example Ollama or LM Studio). A router or gateway sits in front of one or more models or providers and forwards requests. They do different jobs, so they are compared separately.

What does a dash in the comparison table mean?

It means the project's own documentation does not state that attribute. It does not mean the feature is missing; check the tool's review or its repository.

Why is Ollama in two groups?

Ollama is both an inference engine and a runtime that manages models, so it is listed under both. Its README states few of the attributes compared here, so many of its cells show a dash.

Do any of these tools have an affiliate link?

No. PromptQuorum has no affiliate relationship with any tool in this comparison at the time of writing, and no link here earns a commission.

How often is this comparison updated?

It is refreshed twice a year and whenever one of the listed tools' reviews is updated, because the table is generated from the same data as those reviews.

Sources

← Back to Power Local LLM