Key Takeaways
- The RTX 4090 is discontinued (EOL) as of 2026 — every card on the market is used, at ~$2,000-2,600, up sharply from ~$999-1,299 earlier in the year.
- The RTX 5090 is still in production, but the 2026 GDDR7 memory shortage has pushed street pricing to $4,300-$5,000+, more than double its $1,999 MSRP.
- RTX 5090 is ~30-50% faster than RTX 4090 for local LLM inference on large models (70B Q4). On 7B-13B models the gap is smaller (~20%).
- RTX 5090 has 32GB GDDR7 vs RTX 4090's 24GB GDDR6X — 8GB extra VRAM matters for 34B Q8 and large-context 70B Q4 runs.
- A used RTX 4090 (~$56-72 per M tokens) is still cheaper per token generated than a new RTX 5090 at current street prices (~$83-96 per M tokens) — the EOL premium hasn't erased that gap.
- For 7B-13B models: 4090 is overkill regardless of price. You'll hit CPU/cooling limits before maxing GPU.
- For 70B models: 5090 shines. Can run 2-3 smaller 70B models in parallel or single 70B at higher batch sizes.
- RTX 5080 often provides better value than 5090 for local LLMs unless you need dual-GPU setups, but it is subject to the same shortage-driven price inflation.
📍 In One Sentence
The RTX 4090 is discontinued and used-market only, while the RTX 5090 remains in production but costs more than double its MSRP due to a GDDR7 memory shortage; the 5090 is 30-50% faster on large models with 32GB VRAM versus the 4090's 24GB, but the used 4090 still generates tokens more cheaply per dollar.
💬 In Plain Terms
If you already have an RTX 4090, there's little reason to upgrade right now -- the 5090's price premium doesn't match its speed gain for most local LLM use. If you're buying fresh and need a 70B-capable card, a used 4090 remains the better value despite carrying a discontinued-card premium, since the 5090's own price has been pushed up by a memory shortage.
What Are the Raw Speed Differences?
RTX 5090: 21,760 CUDA cores, 1,457 TFLOPS, ~1,792 GB/s memory bandwidth, 32GB GDDR7. MSRP $1,999, but street pricing is $4,300-$5,000+ due to the GDDR7 memory shortage.
RTX 4090: 16,384 CUDA cores, 826 TFLOPS, ~1,008 GB/s memory bandwidth, 24GB GDDR6X. Discontinued (EOL) — NVIDIA has stopped production. Every unit for sale is used, at ~$2,000-2,600.
Real-world LLM inference (Llama 3.3 70B, Q4, batch=1): RTX 5090 scores ~50-55 tokens/sec, RTX 4090 scores ~36 tokens/sec. ~40-50% faster on 70B models. These per-card speeds are unchanged by the pricing shifts — only the cost to get them has moved.
For 7B models (memory-bound but smaller weights): RTX 5090 scores ~90 tokens/sec, RTX 4090 scores ~75 tokens/sec. ~20% faster. The gap is smaller at smaller model sizes.

Does VRAM Matter Between 4090 and 5090?
RTX 5090 has 32GB GDDR7. RTX 4090 has 24GB GDDR6X. The 8GB difference matters for specific model sizes — this VRAM math is unaffected by the 4090's discontinuation or either card's current price.
The VRAM advantage of the 5090 is real and meaningful: 34B Q8 (~28GB) fits comfortably in 32GB but is marginal in 24GB. 70B Q4 (~38-40GB) does NOT fit on either card alone — you need dual-GPU or Apple Silicon for that. For 7B-13B models, 24GB is more than enough.
Cost Per Token: Which Is Actually Cheaper?
- Used RTX 4090: ~$2,000-2,600 (EOL, used-market only, up from ~$999-1,299 earlier in 2026). Achieves 36 tokens/sec on Llama 70B. Cost per token: ~$56-72 per M tokens.
- RTX 5090 new: $4,300-5,000+ street price (MSRP $1,999, inflated by the 2026 GDDR7 shortage). Achieves ~52 tokens/sec on Llama 3.3 70B Q4. Cost per token: ~$83-96 per M tokens.
- Verdict: the 4090 is still cheaper per token generated, even after its EOL price jump — because the 5090's shortage-driven premium has grown faster than the 4090's scarcity premium.

When Should You Actually Upgrade from 4090 to 5090?
Never upgrade for 7B-13B inference. 4090 is overkill for these regardless of price. You'll be CPU-bound or cooling-limited anyway.
If you already own a 4090: keep it. With the 5090 at $4,300-5,000+ street price, the upgrade math is worse than it has ever been. Only move up if you need >40 tok/s on 70B or the 32GB for 34B Q8 and the cost is not a blocker.
If you don't own either and need a 70B-capable card: compare current listings. A used 4090 (~$2,000-2,600, EOL/used-only) still beats a new 5090 (~$4,300-5,000+) on cost per token, at the price of 30-50% lower throughput.
Two-GPU alternative has changed: two used RTX 4090s now cost roughly $4,000-5,200 combined — close to a single 5090's street price — for ~65-72 tokens/sec on 70B Q4 combined (similar to or better than one 5090). The trade-offs are used-only sourcing (no warranty, EOL supply) and more setup complexity (tensor parallelism).
Common Assumptions About the 5090
- Assuming the RTX 4090 is still sold new — it isn't. NVIDIA discontinued the RTX 40-series in 2026; every 4090 on the market is used, at prices well above where they sat earlier in the year.
- Thinking the RTX 5090's $1,999 MSRP is what you'll actually pay —, street pricing is $4,300-$5,000+ due to the GDDR7 shortage.
- Thinking 5090 is 2× faster than 4090 — it's 40-50% faster on 70B, only ~20% faster on 7B.
- Thinking both cards have the same VRAM — they don't. 5090 has 32GB, 4090 has 24GB. The 8GB difference matters for 34B Q8 models.
- Believing you need 5090 to run 70B models — you don't. 4090 runs Llama 3.3 70B Q4 at ~36 tokens/sec. That's sufficient for most local inference tasks, and it's the cheaper card per token even after its EOL price increase.
Frequently Asked Questions
Is the RTX 4090 still available new in 2026?
No. NVIDIA discontinued the RTX 40-series in 2026. Every RTX 4090 for sale now is used, and prices have climbed to ~$2,000-2,600 due to EOL scarcity plus knock-on demand from the RTX 50-series' own GDDR7 shortage.
Is RTX 5090 worth it for running Llama 3.3 70B?
Depends more on price than ever. 4090 gives you ~36 tok/s on 70B Q4, which is usable, for ~$2,000-2,600 used. 5090 gives ~52 tok/s — 40% faster — but street price is $4,300-$5,000+. Worth it if you're running 70B continuously and can't tolerate the slower card; a used 4090 is the pragmatic choice otherwise.
Should I buy RTX 5090 or two RTX 4090s?
Two used 4090s (~$4,000-5,200 total, EOL/used-only) beat one 5090 (~$4,300-5,000+ street) on raw 70B speed (~72 tok/s combined vs ~52). But two-GPU setups add complexity — tensor parallelism, driver overhead, used-only sourcing with no warranty. 5090 is simpler, still in production, and adds 32GB unified VRAM for 34B Q8 use cases.
Does RTX 5090 have better VRAM than 4090?
Yes. RTX 5090 has 32GB GDDR7, RTX 4090 has 24GB GDDR6X. The 8GB extra lets 5090 run 34B Q8 models (~28GB needed) without squeezing. Both cards cannot fit 70B Q4 (~38-40GB) without splitting across GPUs. This VRAM gap is unrelated to the 4090's discontinuation.
Is the RTX 5090 overkill for a 14B model?
Yes, for most use cases. A 14B Q4 model needs ~9GB VRAM and runs at ~55-65 tok/s on a 4090 — already fast. The 5090 would give ~65-75 tok/s, a marginal improvement. A used 4090 (or even a 3090) is far more cost-effective for 14B inference, especially at current 5090 street prices.
Does the RTX 5090 laptop GPU have 32GB VRAM?
No. The RTX 5090 Laptop GPU has 24GB GDDR7 (reduced from the desktop's 32GB due to power and thermal constraints). It's significantly slower than the desktop 5090 but still faster than a 4090 on the same task.
Will 5090 prices drop back to its $1,999 MSRP?
Only once the 2026 GDDR7 memory shortage eases — there's no fixed timeline. Street pricing sits at $4,300-$5,000+, more than double MSRP. The RTX 4090's own price history is a reminder that used pricing doesn't only trend downward: it fell from $1,499 (2022 launch) to a low of ~$999 used, then rebounded to ~$2,000-2,600 after its 2026 discontinuation.
Can I use RTX 5090 with a 750W power supply?
Barely. RTX 5090 draws 575W alone. Pair with a 850W or 1000W PSU to avoid voltage sag under load.
Is RTX 5080 a better value than 5090?
Often, yes — 5080 is roughly 80% of 5090's speed at a meaningfully lower price, though it is also subject to GDDR7 shortage-driven price inflation in 2026. For local LLMs, 5080 is usually the sweet spot unless you specifically need the 5090's 32GB.
How much faster is 5090 on multimodal models like Qwen-VL 70B?
Similar 20-25% lift. Multimodal compute is still memory-bound, so the bandwidth advantage of 5090 helps, but not dramatically.
Sources
- NVIDIA RTX 5090 and 4090 official specifications: CUDA cores, TFLOPS, memory bandwidth
- MLCommons MLPerf Inference Benchmark: Token generation speed on LLaMA 70B and Mistral models
- TechPowerUp GPU Database: RTX 5090 vs. 4090 power consumption and memory bandwidth comparison
