Skip to main content
PromptQuorumBuilt for humans. Structured for AI.
Home/Local LLMs/RTX 5090 vs RTX 4090 for Local LLM Inference
GPU Buying Guides

RTX 5090 vs RTX 4090 for Local LLM Inference

·6 min·By Hans Kuepper · Founder of PromptQuorum · Discovery engine for open-weight & open-source AI

This page contains links to third-party products for reference. PromptQuorum is not enrolled in any affiliate program — these are plain links that earn no commission. Clicking links and your next steps are entirely your own responsibility. These links do not represent any endorsement or verification by PromptQuorum.

RTX 5090 32GB on Amazonproduct link · disclosedRTX 4090 24GB on Amazonproduct link · disclosed

The RTX 4090 is discontinued (EOL) as of 2026 — used-market only, now priced around $2,000-2,600. The RTX 5090 is still in production but its street price has climbed to $4,300-$5,000+ under the 2026 GDDR7 shortage. RTX 5090 remains 30-50% faster on large models and has 32GB GDDR7 (vs 4090's 24GB GDDR6X), but the used 4090 is still the cheaper card per token generated.

The RTX 4090 is now discontinued (EOL) — every unit on the market is used, at prices that have climbed to ~$2,000-2,600. The RTX 5090 remains in production, but the 2026 GDDR7 memory shortage has pushed its street price to $4,300-$5,000+, more than double its $1,999 MSRP. For local LLMs, RTX 5090 is still 30-50% faster than RTX 4090 on 70B models and has 8GB more VRAM (32GB vs 24GB). If you already own a 4090, keep it — the upgrade math is worse than ever at current prices. If you need a 70B-capable card and don't own one, a used 4090 remains the cheaper way in per token generated, despite the EOL premium.

RTX 5090 vs RTX 4090 for Local LLM Inference

Key Takeaways

  • The RTX 4090 is discontinued (EOL) as of 2026 — every card on the market is used, at ~$2,000-2,600, up sharply from ~$999-1,299 earlier in the year.
  • The RTX 5090 is still in production, but the 2026 GDDR7 memory shortage has pushed street pricing to $4,300-$5,000+, more than double its $1,999 MSRP.
  • RTX 5090 is ~30-50% faster than RTX 4090 for local LLM inference on large models (70B Q4). On 7B-13B models the gap is smaller (~20%).
  • RTX 5090 has 32GB GDDR7 vs RTX 4090's 24GB GDDR6X — 8GB extra VRAM matters for 34B Q8 and large-context 70B Q4 runs.
  • A used RTX 4090 (~$56-72 per M tokens) is still cheaper per token generated than a new RTX 5090 at current street prices (~$83-96 per M tokens) — the EOL premium hasn't erased that gap.
  • For 7B-13B models: 4090 is overkill regardless of price. You'll hit CPU/cooling limits before maxing GPU.
  • For 70B models: 5090 shines. Can run 2-3 smaller 70B models in parallel or single 70B at higher batch sizes.
  • RTX 5080 often provides better value than 5090 for local LLMs unless you need dual-GPU setups, but it is subject to the same shortage-driven price inflation.

📍 In One Sentence

The RTX 4090 is discontinued and used-market only, while the RTX 5090 remains in production but costs more than double its MSRP due to a GDDR7 memory shortage; the 5090 is 30-50% faster on large models with 32GB VRAM versus the 4090's 24GB, but the used 4090 still generates tokens more cheaply per dollar.

💬 In Plain Terms

If you already have an RTX 4090, there's little reason to upgrade right now -- the 5090's price premium doesn't match its speed gain for most local LLM use. If you're buying fresh and need a 70B-capable card, a used 4090 remains the better value despite carrying a discontinued-card premium, since the 5090's own price has been pushed up by a memory shortage.

What Are the Raw Speed Differences?

RTX 5090: 21,760 CUDA cores, 1,457 TFLOPS, ~1,792 GB/s memory bandwidth, 32GB GDDR7. MSRP $1,999, but street pricing is $4,300-$5,000+ due to the GDDR7 memory shortage.

RTX 4090: 16,384 CUDA cores, 826 TFLOPS, ~1,008 GB/s memory bandwidth, 24GB GDDR6X. Discontinued (EOL) — NVIDIA has stopped production. Every unit for sale is used, at ~$2,000-2,600.

Real-world LLM inference (Llama 3.3 70B, Q4, batch=1): RTX 5090 scores ~50-55 tokens/sec, RTX 4090 scores ~36 tokens/sec. ~40-50% faster on 70B models. These per-card speeds are unchanged by the pricing shifts — only the cost to get them has moved.

For 7B models (memory-bound but smaller weights): RTX 5090 scores ~90 tokens/sec, RTX 4090 scores ~75 tokens/sec. ~20% faster. The gap is smaller at smaller model sizes.

RTX 5090 vs RTX 4090 Speed -- Real-world LLM inference
RTX 5090 vs RTX 4090 Speed -- Real-world LLM inference

Does VRAM Matter Between 4090 and 5090?

RTX 5090 has 32GB GDDR7. RTX 4090 has 24GB GDDR6X. The 8GB difference matters for specific model sizes — this VRAM math is unaffected by the 4090's discontinuation or either card's current price.

The VRAM advantage of the 5090 is real and meaningful: 34B Q8 (~28GB) fits comfortably in 32GB but is marginal in 24GB. 70B Q4 (~38-40GB) does NOT fit on either card alone — you need dual-GPU or Apple Silicon for that. For 7B-13B models, 24GB is more than enough.

Cost Per Token: Which Is Actually Cheaper?

  • Used RTX 4090: ~$2,000-2,600 (EOL, used-market only, up from ~$999-1,299 earlier in 2026). Achieves 36 tokens/sec on Llama 70B. Cost per token: ~$56-72 per M tokens.
  • RTX 5090 new: $4,300-5,000+ street price (MSRP $1,999, inflated by the 2026 GDDR7 shortage). Achieves ~52 tokens/sec on Llama 3.3 70B Q4. Cost per token: ~$83-96 per M tokens.
  • Verdict: the 4090 is still cheaper per token generated, even after its EOL price jump — because the 5090's shortage-driven premium has grown faster than the 4090's scarcity premium.
Cost Per Token: Which Is Cheaper? -- Price vs throughput on Llama 70B
Cost Per Token: Which Is Cheaper? -- Price vs throughput on Llama 70B

When Should You Actually Upgrade from 4090 to 5090?

Never upgrade for 7B-13B inference. 4090 is overkill for these regardless of price. You'll be CPU-bound or cooling-limited anyway.

If you already own a 4090: keep it. With the 5090 at $4,300-5,000+ street price, the upgrade math is worse than it has ever been. Only move up if you need >40 tok/s on 70B or the 32GB for 34B Q8 and the cost is not a blocker.

If you don't own either and need a 70B-capable card: compare current listings. A used 4090 (~$2,000-2,600, EOL/used-only) still beats a new 5090 (~$4,300-5,000+) on cost per token, at the price of 30-50% lower throughput.

Two-GPU alternative has changed: two used RTX 4090s now cost roughly $4,000-5,200 combined — close to a single 5090's street price — for ~65-72 tokens/sec on 70B Q4 combined (similar to or better than one 5090). The trade-offs are used-only sourcing (no warranty, EOL supply) and more setup complexity (tensor parallelism).

Common Assumptions About the 5090

  • Assuming the RTX 4090 is still sold new — it isn't. NVIDIA discontinued the RTX 40-series in 2026; every 4090 on the market is used, at prices well above where they sat earlier in the year.
  • Thinking the RTX 5090's $1,999 MSRP is what you'll actually pay —, street pricing is $4,300-$5,000+ due to the GDDR7 shortage.
  • Thinking 5090 is 2× faster than 4090 — it's 40-50% faster on 70B, only ~20% faster on 7B.
  • Thinking both cards have the same VRAM — they don't. 5090 has 32GB, 4090 has 24GB. The 8GB difference matters for 34B Q8 models.
  • Believing you need 5090 to run 70B models — you don't. 4090 runs Llama 3.3 70B Q4 at ~36 tokens/sec. That's sufficient for most local inference tasks, and it's the cheaper card per token even after its EOL price increase.

Frequently Asked Questions

Is the RTX 4090 still available new in 2026?

No. NVIDIA discontinued the RTX 40-series in 2026. Every RTX 4090 for sale now is used, and prices have climbed to ~$2,000-2,600 due to EOL scarcity plus knock-on demand from the RTX 50-series' own GDDR7 shortage.

Is RTX 5090 worth it for running Llama 3.3 70B?

Depends more on price than ever. 4090 gives you ~36 tok/s on 70B Q4, which is usable, for ~$2,000-2,600 used. 5090 gives ~52 tok/s — 40% faster — but street price is $4,300-$5,000+. Worth it if you're running 70B continuously and can't tolerate the slower card; a used 4090 is the pragmatic choice otherwise.

Should I buy RTX 5090 or two RTX 4090s?

Two used 4090s (~$4,000-5,200 total, EOL/used-only) beat one 5090 (~$4,300-5,000+ street) on raw 70B speed (~72 tok/s combined vs ~52). But two-GPU setups add complexity — tensor parallelism, driver overhead, used-only sourcing with no warranty. 5090 is simpler, still in production, and adds 32GB unified VRAM for 34B Q8 use cases.

Does RTX 5090 have better VRAM than 4090?

Yes. RTX 5090 has 32GB GDDR7, RTX 4090 has 24GB GDDR6X. The 8GB extra lets 5090 run 34B Q8 models (~28GB needed) without squeezing. Both cards cannot fit 70B Q4 (~38-40GB) without splitting across GPUs. This VRAM gap is unrelated to the 4090's discontinuation.

Is the RTX 5090 overkill for a 14B model?

Yes, for most use cases. A 14B Q4 model needs ~9GB VRAM and runs at ~55-65 tok/s on a 4090 — already fast. The 5090 would give ~65-75 tok/s, a marginal improvement. A used 4090 (or even a 3090) is far more cost-effective for 14B inference, especially at current 5090 street prices.

Does the RTX 5090 laptop GPU have 32GB VRAM?

No. The RTX 5090 Laptop GPU has 24GB GDDR7 (reduced from the desktop's 32GB due to power and thermal constraints). It's significantly slower than the desktop 5090 but still faster than a 4090 on the same task.

Will 5090 prices drop back to its $1,999 MSRP?

Only once the 2026 GDDR7 memory shortage eases — there's no fixed timeline. Street pricing sits at $4,300-$5,000+, more than double MSRP. The RTX 4090's own price history is a reminder that used pricing doesn't only trend downward: it fell from $1,499 (2022 launch) to a low of ~$999 used, then rebounded to ~$2,000-2,600 after its 2026 discontinuation.

Can I use RTX 5090 with a 750W power supply?

Barely. RTX 5090 draws 575W alone. Pair with a 850W or 1000W PSU to avoid voltage sag under load.

Is RTX 5080 a better value than 5090?

Often, yes — 5080 is roughly 80% of 5090's speed at a meaningfully lower price, though it is also subject to GDDR7 shortage-driven price inflation in 2026. For local LLMs, 5080 is usually the sweet spot unless you specifically need the 5090's 32GB.

How much faster is 5090 on multimodal models like Qwen-VL 70B?

Similar 20-25% lift. Multimodal compute is still memory-bound, so the bandwidth advantage of 5090 helps, but not dramatically.

Sources

  • NVIDIA RTX 5090 and 4090 official specifications: CUDA cores, TFLOPS, memory bandwidth
  • MLCommons MLPerf Inference Benchmark: Token generation speed on LLaMA 70B and Mistral models
  • TechPowerUp GPU Database: RTX 5090 vs. 4090 power consumption and memory bandwidth comparison

A Note on Third-Party Facts

This article references third-party AI models, benchmarks, prices, and licenses. The AI landscape changes rapidly. Benchmark scores, license terms, model names, and API prices can shift between the time of writing and the time you read this. Before making deployment or compliance decisions based on this article, verify current figures on each provider’s official source: Hugging Face model cards for licenses and benchmarks, provider websites for API pricing, and EUR-Lex for current GDPR and EU AI Act text.

Run PromptQuorum with a local LLM, your own API keys, or both — you pick the backend.

Download the PromptQuorum Beta →

← Back to Local LLMs