24GB · AI Score 8.2/100 · first-party measured on 12 AI workloads
8.2 AI Score Includes estimates
Every number on this page is first-party: NVIDIA Titan RTX was run on our pinned 12-workload AI suite on 2026-07-11, with under 0.5% run-to-run variance. On Llama 3.1 8B (Q4_K_M) NVIDIA Titan RTX delivers about 107.39 tokens/sec. Stepping up to Qwen3 32B it holds roughly 28.36 tok/s. Llama 3.3 70B does not fit. It needs roughly 42GB and this card has 24GB. For image generation, SDXL runs at 3.28 it/s, while FLUX.1-dev won't fit at BF16 (needs ~26GB). 6 of the 12 workloads won't fit on 24GB at the tested precision, Llama 3.3 70B, FLUX.1-dev, FLUX.1 Kontext, Qwen-Image-Edit and others. We publish those as hard gates rather than quietly dropping to a smaller quant.
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| MiniCPM5 2B | 225.16 tok/s | 189 W49°CQ4_K_M | ✓ Measured |
| LFM2.5 2.6B | 244.88 tok/s | 196 W48°CQ4_K_M | ✓ Measured |
| Granite 4.1 3B | 172.09 tok/s | 196 W54°CQ4_K_M | ✓ Measured |
| Nemotron 3 Nano 4B | 166.76 tok/s | 179 W51°CQ4_K_M | ✓ Measured |
| Qwen3 4B | 158.13 tok/s | 2.9 GB peak147 W67°C1.08 tok/WQ4_K_M | ✓ Measured |
| DeepSeek Coder 7B Instruct v1.5 | 122.74 tok/s | 209 W54°CQ4_K_M | ✓ Measured |
| Llama 3 8B | 108.26 tok/s | 203 W54°CQ4_K_M | ✓ Measured |
| Llama 3.1 8B | 107.39 tok/s | 4.8 GB peak159 W70°C0.68 tok/WQ4_K_M | ✓ Measured |
| Qwen3 8B | 104.36 tok/s | 211 W54°CQ4_K_M | ✓ Measured |
| Nemotron Nano 9B v2 | 80.66 tok/s | 209 W55°CQ4_K_M | ✓ Measured |
| Ornith 1.5 9B | 93.68 tok/s | 216 W53°CQ4_K_M | ✓ Measured |
| Gemma 4 12B | 66.58 tok/s | 219 W54°CQ4_K_M | ✓ Measured |
| Qwen2.5-Coder 14B | 59.75 tok/s | 8.5 GB peak152 W75°C0.39 tok/WQ4_K_M | ✓ Measured |
| Qwen3 14B | 62.47 tok/s | 223 W55°CQ4_K_M | ✓ Measured |
| Gemma 4 26B A4B | 119.86 tok/s | 162 W51°CQ4_K_M | ✓ Measured |
| Qwen3 30B A3B | 156.15 tok/s | 166 W52°CQ4_K_M | ✓ Measured |
| Qwen3 32B | 28.36 tok/s | 18.6 GB peak121 W77°C0.24 tok/WQ4_K_M | ✓ Measured |
| Llama 3.3 70B | ✕ Won't fit needs ~46 GB | VRAM-gated at this precision | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| FLUX.1 Kontext dev | ✕ Won't fit needs ~26 GB | VRAM-gated at this precision | ✓ Measured |
| Qwen-Image-Edit | ✕ Won't fit needs ~42 GB | VRAM-gated at this precision | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| Stable Diffusion XL | 6.56 images/min | 12.7 GB peak273 W81°C9.2 s/img | ✓ Measured |
| FLUX.1 dev | ✕ Won't fit needs ~26 GB | VRAM-gated at this precision | ✓ Measured |
| Workload | Result | Telemetry | Data |
|---|---|---|---|
| LTX-Video (distilled) | n/a frames/s | estimated | Est. |
| Wan 2.2 5B (720p) | n/a frames/s | estimated | Est. |
| Architecture | Turing (TU102) |
| CUDA cores | 4,608 |
| VRAM | 24GB GDDR6 |
| Memory bus | 384-bit |
| Memory bandwidth | 672 GB/s |
| Boost clock | 1,770 MHz |
| TDP | 280 W |
| Process | 12nm |
| Interface | PCIe 3.0 x16 |
| Release date | 2018-12-18 |
| Launch MSRP | $2,499 |
NVIDIA Titan RTX scores 8.2/100, #32 of 102. It ran 5 of 12; 6 exceeded its 24GB. Every figure here is our own measurement.
100% = this card, AI & Machine Learning headline metric (AI Score). #7 of 61 desktop cards in this vertical.
| GPU | Relative | % | AI Score |
|---|---|---|---|
| NVIDIA RTX 5880 Ada Generation | 179% | 14.7 | |
| NVIDIA RTX 5000 Ada Generation | 134% | 11 | |
| NVIDIA GeForce RTX 4090 | 129% | 10.6 | |
| NVIDIA GeForce RTX 3090 Ti | 104% | 8.5 | |
| NVIDIA Titan RTX | 100% | 8.2 | |
| NVIDIA GeForce RTX 3090 | 94% | 7.7 | |
| GeForce RTX 5080 | 63% | 5.2 | |
| NVIDIA GeForce RTX 4080 | 59% | 4.8 | |
| GeForce RTX 4080 Super | 57% | 4.7 |
Same card, other workloads: NVIDIA Titan RTX Gaming benchmarks
← All AI & Machine Learning GPU rankings
| Transistors | 18,600 million |
| Die size | 754 mm² |
| Process node | 12 nm |
| Fabricated by | TSMC |
| Transistor density | 24.7 million per mm² |
Denser than 69% of the 746 cards we have silicon data for. Density is the clearest measure of what a process node bought: a card that gained it without growing the die got its speed from the fab rather than the architecture.
Silicon figures from Wikipedia (CC BY-SA 4.0). Benchmarks on this page are our own. Compare every chip.
Whole-job timings, composed from our measured per-model results on this card.
| Workflow | Time | Energy | Basis |
|---|---|---|---|
| Full codebase review | 16.7 min | 42.48 Wh | measured |
Can't run: Product photo shoot (needs FLUX.1 Kontext dev), Photo restoration batch (needs FLUX.1 Kontext dev), Restore and enlarge photos (needs FLUX.1 Kontext dev), Character sheet, 12 poses (needs FLUX.1 dev), 6-panel comic page (needs FLUX.1 dev), Long-form article batch (needs Llama 3.3 70B).
This card is $2,499 to buy. The cheapest listed rate on Vast.ai is $0.242/hour, but that is the floor: we budget $0.290/hour, a 20% premium, because idle time, storage and unavailable cheap instances all land on the same bill. At that rate buying wins after 8,605 GPU-hours. Below it you are paying for idle silicon.
| How you would use it | GPU-hours a year | Rental cost a year | Time to break even |
|---|---|---|---|
| 2 hours a day, hobby | 730 | $212 | 11.8 years |
| 8 hours a day, working on it | 2,920 | $848 | 2.9 years |
| 24/7, always-on agent | 8,760 | $2,544 | 11.8 months |
At hobby usage this card is very unlikely to pay for itself before it is superseded. Rent it. Rental figures include a 20% premium over the cheapest listed rate. Ignores electricity, resale and the fact that a rented card can be a newer one tomorrow.
Cheapest of the RunPod and Vast on-demand rates we see, sampled daily. Spot and interruptible pricing runs lower.