Qwen3 8B · 57 GPUs measured first-party · llama.cpp Q4_K_M · Updated October 2026

What GPU Do You Need for Qwen3 8B?

Qwen3 8B is what we'd call the base tier for a real chat experience, the class of model the cheap API endpoints actually serve. We measured it on 57 GPUs with llama.cpp at Q4_K_M: 261 tok/s on the B300 at the top, and a ~6GB measured VRAM floor that puts it within reach of nearly every desktop card sold in the last five years.

Benchmarked weights: Qwen/Qwen3-8B-GGUF

Fastest we measured
NVIDIA B300

NVIDIA B300

261.4 tok/s on Qwen3 8B, the ceiling. Measured on our bench. 288GB of VRAM, $40,000 at launch.

Pros
  • 261.4 tok/s on Qwen3 8B
  • 288GB, clears the Qwen3 8B floor
  • Rentable by the hour rather than bought
Cons
  • 1400W board rating
  • Datacenter or workstation hardware, not a retail purchase
Best consumer card
NVIDIA GeForce RTX 5090

NVIDIA GeForce RTX 5090

243.9 tok/s on Qwen3 8B, fastest card you can buy at retail. Measured on our bench. 32GB of VRAM, $1,999 at launch.

Pros
  • 243.9 tok/s on Qwen3 8B
  • 32GB, clears the Qwen3 8B floor
Cons
  • 575W board rating
Cheapest card that runs it
NVIDIA GeForce GTX 1660

NVIDIA GeForce GTX 1660

33.74 tok/s on Qwen3 8B, lowest launch price that still fits. Measured on our bench. 6GB of VRAM, $219 at launch.

Pros
  • 33.74 tok/s on Qwen3 8B
  • 6GB, clears the Qwen3 8B floor
Cons
  • 120W board rating
Best value
NVIDIA GeForce RTX 5060

NVIDIA GeForce RTX 5060

77.26 tok/s on Qwen3 8B, most speed per dollar. Measured on our bench. 8GB of VRAM, $249 at launch. That is 310.3 tok/s per $1,000 of launch price.

Pros
  • 77.26 tok/s on Qwen3 8B
  • 8GB, clears the Qwen3 8B floor
Cons
  • 145W board rating
261.4tok/s
Fastest: NVIDIA B300
measured
43
Cards that run Qwen3 8B
of 43 we have data for
0
Cards that can't run it at all
published as hard gates, not omissions
616%
Fastest vs slowest that fits
261.4 vs 36.51 tok/s

What GPU Do You Need for Qwen3 8B?, tok/s by GPU

NVIDIA B300
261.4 tok/s
NVIDIA B200
253.4 tok/s
NVIDIA H200
248 tok/s
NVIDIA H100 80GB HBM3
244.2 tok/s
NVIDIA GeForce RTX 5090
243.9 tok/s
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
225.9 tok/s
NVIDIA H100 NVL
221.8 tok/s
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition
221.3 tok/s
NVIDIA H100 PCIe
185.9 tok/s
NVIDIA GeForce RTX 4090
164.3 tok/s
NVIDIA GeForce RTX 3090 Ti
156.5 tok/s
GeForce RTX 5080
153.9 tok/s
NVIDIA RTX 5880 Ada Generation
152 tok/s
NVIDIA A100 80GB SXM4
150 tok/s
NVIDIA A100 40GB PCIe
146.9 tok/s

Top 15 shown; 42 more cards in the full table below.

Measured on our own bench. A card absent from this chart has not been run on this model yet, or cannot fit it.

Efficiency: tok/s per 100W drawn

NVIDIA H200
194.8 tok/s / 100W
NVIDIA H100 80GB HBM3
140.11 tok/s / 100W
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition
114.39 tok/s / 100W
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
113.94 tok/s / 100W
NVIDIA H100 PCIe
106.52 tok/s / 100W
NVIDIA H100 NVL
100.41 tok/s / 100W
GeForce RTX 5060 Ti
98.12 tok/s / 100W
NVIDIA A100 80GB SXM4
97.95 tok/s / 100W
NVIDIA B300
95.02 tok/s / 100W
NVIDIA RTX PRO 4000 Blackwell
94.59 tok/s / 100W
NVIDIA GeForce RTX 5090
92.15 tok/s / 100W
NVIDIA A100 40GB PCIe
89.55 tok/s / 100W
NVIDIA A100 40GB SXM4
89.15 tok/s / 100W
NVIDIA B200
89.12 tok/s / 100W
GeForce RTX 5080
86.31 tok/s / 100W

Top 15 shown; 41 more cards in the full table below.

Power is the average pulled during the run, sampled at 1Hz. The fastest card is often not the one here, and for anything left running this is the number that shows up on the bill.

Value: tok/s per $1,000 of MSRP

NVIDIA GeForce RTX 5060
310.28 tok/s / $1k
NVIDIA GeForce RTX 3060 Ti
207.92 tok/s / $1k
NVIDIA GeForce GTX 1660 Super
205.98 tok/s / $1k
GeForce RTX 5070
199.69 tok/s / $1k
NVIDIA GeForce RTX 3060
194.04 tok/s / $1k
GeForce RTX 5070 Ti
185.94 tok/s / $1k
GeForce RTX 5060 Ti
178.62 tok/s / $1k
NVIDIA GeForce RTX 3080
172.89 tok/s / $1k
NVIDIA GeForce GTX 1660 Ti
172.69 tok/s / $1k
GeForce RTX 4060
171.61 tok/s / $1k
NVIDIA GeForce RTX 3070 Ti
170.95 tok/s / $1k
NVIDIA GeForce RTX 3070 Founders Edition
158.16 tok/s / $1k
NVIDIA GeForce GTX 1660
154.06 tok/s / $1k
GeForce RTX 5080
154.04 tok/s / $1k
NVIDIA GeForce RTX 4070
151.82 tok/s / $1k

Top 15 shown; 42 more cards in the full table below.

Launch price, not street price, so it ages. A speed leaderboard always crowns the most expensive card; this is the counterweight.

Qwen3 8B. Measured generation speed by GPU

NVIDIA B300261.4
NVIDIA B200253.4
NVIDIA H200248
NVIDIA H100 80GB HBM3244.2
NVIDIA GeForce RTX 5090243.9
NVIDIA RTX PRO 6000 Blackwell Workstation Edition225.9
NVIDIA H100 NVL221.8
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition221.3
NVIDIA H100 PCIe185.9
NVIDIA GeForce RTX 4090164.3
NVIDIA GeForce RTX 3090 Ti156.5
GeForce RTX 5080153.9
NVIDIA RTX 5880 Ada Generation152
NVIDIA A100 80GB SXM4150
NVIDIA A100 40GB PCIe146.9
NVIDIA A100 40GB SXM4146.2
NVIDIA GeForce RTX 3080 Ti140.7
GeForce RTX 5070 Ti139.3
NVIDIA GeForce RTX 3090137.6
NVIDIA L40S130.3
GeForce RTX 4080 Super126.4
NVIDIA GeForce RTX 4080123.7
NVIDIA GeForce RTX 3080120.8
NVIDIA RTX A6000120.2
NVIDIA GeForce RTX 4070 Ti Super117.4
NVIDIA RTX A5000114.8
GeForce RTX 5070109.6
NVIDIA Titan RTX104.4
NVIDIA GeForce RTX 3070 Ti102.4
NVIDIA RTX PRO 4000 Blackwell101.6
NVIDIA TITAN V98.34
NVIDIA GeForce RTX 407090.94
NVIDIA GeForce RTX 4070 Ti89.83
NVIDIA A10G84.09
NVIDIA GeForce RTX 3060 Ti82.96
NVIDIA Quadro RTX 800081.7
NVIDIA GeForce RTX 3070 Founders Edition78.92
NVIDIA GeForce RTX 506077.26
NVIDIA RTX 4500 Ada Generation76.96
GeForce RTX 5060 Ti76.63
NVIDIA RTX A400074.04
NVIDIA GeForce RTX 2070 SUPER73.65
NVIDIA Quadro RTX 500072.21
NVIDIA GeForce RTX 2070 (power capped)65.9
NVIDIA RTX 4000 (Ada Generation)65.39
NVIDIA GeForce RTX 306063.84
NVIDIA GeForce RTX 2060 Super (power capped)60.36
NVIDIA GeForce RTX 4060 Ti 16GB54.34
GeForce RTX 406051.31
GeForce GTX 1080 Ti49.08
NVIDIA L448.95
NVIDIA GeForce GTX 1660 Ti48.18
NVIDIA GeForce GTX 1660 Super47.17
NVIDIA TITAN Xp (power capped)45.11
NVIDIA T436.51
NVIDIA GeForce GTX 108034.97
NVIDIA GeForce GTX 166033.74
GPUtok/sPrompt t/stok/WAvg power
NVIDIA B300261.44934.70.95275.1 W
NVIDIA B200253.48894.80.89284.4 W
NVIDIA H2002488624.71.95127.3 W
NVIDIA H100 80GB HBM3244.28750.91.4174.3 W
NVIDIA GeForce RTX 5090243.913393.10.92264.7 W
NVIDIA RTX PRO 6000 Blackwell Workstation Edition225.911640.41.14198.3 W
NVIDIA H100 NVL221.88709.21220.9 W
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition221.311471.21.14193.5 W
NVIDIA H100 PCIe185.96740.31.07174.5 W
NVIDIA GeForce RTX 4090164.311239.70.77213.5 W
NVIDIA GeForce RTX 3090 Ti156.55887.60.54291.1 W
GeForce RTX 5080153.98115.30.86178.3 W
NVIDIA RTX 5880 Ada Generation1527913.50.8189.4 W
NVIDIA A100 80GB SXM41504362.80.98153.1 W
NVIDIA A100 40GB PCIe146.94599.70.9164.0 W
NVIDIA A100 40GB SXM4146.24200.70.89164.0 W
NVIDIA GeForce RTX 3080 Ti140.75087.60.6236.1 W
GeForce RTX 5070 Ti139.36568.80.84165.5 W
NVIDIA GeForce RTX 3090137.650360.56244.2 W
NVIDIA L40S130.39517.40.65200.3 W
GeForce RTX 4080 Super126.476220.77165.0 W
NVIDIA GeForce RTX 4080123.77580.80.72171.1 W
NVIDIA GeForce RTX 3080120.84500.70.59204.6 W
NVIDIA RTX A6000120.24711.70.62194.7 W
NVIDIA GeForce RTX 4070 Ti Super117.46570.30.73161.9 W
NVIDIA RTX A5000114.84112.40.65176.3 W
GeForce RTX 5070109.64988.90.78139.7 W
NVIDIA Titan RTX104.42812.40.5210.8 W
NVIDIA GeForce RTX 3070 Ti102.436140.49208.7 W
NVIDIA RTX PRO 4000 Blackwell101.64610.30.95107.4 W
NVIDIA TITAN V98.342634.30.73134.5 W
NVIDIA GeForce RTX 407090.944407.70.66138.5 W
NVIDIA GeForce RTX 4070 Ti89.835463.30.62144.5 W
NVIDIA A10G84.093142.40.72116.6 W
NVIDIA GeForce RTX 3060 Ti82.962679.80.61136.6 W
NVIDIA Quadro RTX 800081.72421.20.47172.1 W
NVIDIA GeForce RTX 3070 Founders Edition78.923090.20.5157.2 W
NVIDIA GeForce RTX 506077.263027.30.71109.3 W
NVIDIA RTX 4500 Ada Generation76.965339.10.67115.6 W
GeForce RTX 5060 Ti76.633223.40.9878.1 W
NVIDIA RTX A400074.042775.60.6123.3 W
NVIDIA GeForce RTX 2070 SUPER73.651936.30.45164.9 W
NVIDIA Quadro RTX 500072.212144.90.42173.1 W
NVIDIA GeForce RTX 2070 (power capped)65.91550.60.54122.3 W
NVIDIA RTX 4000 (Ada Generation)65.393885.90.64102.5 W
NVIDIA GeForce RTX 306063.842101.70.48132.1 W
NVIDIA GeForce RTX 2060 Super (power capped)60.361407.80.53112.9 W
NVIDIA GeForce RTX 4060 Ti 16GB54.343391.30.53102.3 W
GeForce RTX 406051.312401.4——
GeForce GTX 1080 Ti49.081005.30.2241.2 W
NVIDIA L448.952934.90.7862.8 W
NVIDIA GeForce GTX 1660 Ti48.18165.20.6475.0 W
NVIDIA GeForce GTX 1660 Super47.17144.90.5881.4 W
NVIDIA TITAN Xp (power capped)45.11797.30.31143.5 W
NVIDIA T436.511157.90.661.1 W
NVIDIA GeForce GTX 108034.97744.50.21170.5 W
NVIDIA GeForce GTX 166033.74144.90.4378.9 W

My take: this is the quality baseline. Below roughly this tier, local LLM use is pipeline work; at 8B it becomes something you'd actually let users talk to. If you're deciding what the minimum viable model for a product experience is, this is where I'd draw the line, 4B if you're squeezed, 8B if you want answers that consistently hold up. It's also the sweet spot for a first local setup: the ~6GB floor means an 8GB card runs it with room for context, and you don't need to think about quantization tricks or offloading.

What the numbers show. The B300 tops the chart at 261 tok/s but burns 275W doing it, 0.95 tok/W. The H200 lands 5% slower at 248 tok/s while drawing 127W, which is 1.95 tok/W and the best efficiency on the board. That's a pattern you'll see across our small-model results: Blackwell wins the headline, Hopper wins the power bill. At the affordable end, an Intel Arc A580 at $179 clears the VRAM floor, and even the ancient T4 still generates 37 tok/s, faster than most people read.

About Qwen3 8B. Qwen3 8B: from Qwen, 8.2B parameters, on Hugging Face since April 2025, Apache 2.0 licence. 14,968,230 downloads in the last 30 days and 6 community quantizations.

How it compares. H100 80GB HBM3: Qwen3 8B 244.2 tok/s, DeepSeek-R1-0528-Qwen3-8B 245.3 (8B), Qwen2.5-VL 7B Instruct 267.7 (8B), Llama-3.1-8B 261.8 (8B), Llama 3 8B 264.4 (8B). All 4 beat Qwen3 8B here.

Cost on a rented GPU. 1M generated tokens of Qwen3 8B: $0.16 on a RTX 3060 ($0.036/hr, 4.4 hours), $7.37 on a B300 ($6.94/hr, 64 min, 47.1x the cost).

Qwen3 8B: cost per 1M generated tokens on rented GPUs

NVIDIA GeForce RTX 3060$0.036/hr
NVIDIA GeForce RTX 3080$0.082/hr
NVIDIA GeForce RTX 3090$0.12/hr
NVIDIA GeForce RTX 3070 Founders Edition$0.075/hr
NVIDIA RTX A4000$0.078/hr
GeForce RTX 5070 Ti$0.15/hr
NVIDIA TITAN Xp (power capped)$0.049/hr
NVIDIA GeForce RTX 5060$0.090/hr
NVIDIA GeForce RTX 3080 Ti$0.18/hr
NVIDIA Quadro RTX 5000$0.096/hr
NVIDIA GeForce RTX 4070$0.12/hr
GeForce RTX 5080$0.21/hr
NVIDIA GeForce RTX 4070 Ti Super$0.16/hr
NVIDIA TITAN V$0.14/hr
NVIDIA RTX A5000$0.16/hr
NVIDIA Titan RTX$0.15/hr
NVIDIA GeForce RTX 5090$0.39/hr
NVIDIA GeForce RTX 4080$0.20/hr
NVIDIA GeForce RTX 3090 Ti$0.27/hr
GeForce RTX 5070$0.19/hr
GeForce RTX 4080 Super$0.22/hr
GeForce RTX 5060 Ti$0.14/hr
NVIDIA GeForce RTX 4070 Ti$0.18/hr
NVIDIA RTX PRO 4000 Blackwell$0.20/hr
NVIDIA GeForce RTX 4090$0.34/hr
NVIDIA RTX A6000$0.33/hr
NVIDIA A100 40GB PCIe$0.45/hr
NVIDIA RTX 4000 (Ada Generation)$0.20/hr
NVIDIA Quadro RTX 8000$0.26/hr
NVIDIA RTX 5880 Ada Generation$0.49/hr
NVIDIA A100 40GB SXM4$0.47/hr
NVIDIA T4$0.14/hr
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition$1.00/hr
NVIDIA RTX 4500 Ada Generation$0.36/hr
NVIDIA RTX PRO 6000 Blackwell Workstation Edition$1.08/hr
NVIDIA L40S$0.79/hr
NVIDIA A100 80GB SXM4$0.95/hr
NVIDIA H100 80GB HBM3$2.14/hr
NVIDIA L4$0.44/hr
NVIDIA H100 PCIe$1.94/hr
NVIDIA H100 NVL$2.59/hr
NVIDIA H200$3.59/hr
NVIDIA B200$5.98/hr
NVIDIA B300$6.94/hr
GPUCheapest rateSpeed (tok/s)Cost per 1M generated tokens
NVIDIA GeForce RTX 3060$0.036/hr63.84$0.16
NVIDIA GeForce RTX 3080$0.082/hr120.8$0.19
NVIDIA GeForce RTX 3090$0.12/hr137.6$0.25
NVIDIA GeForce RTX 3070 Founders Edition$0.075/hr78.92$0.26
NVIDIA RTX A4000$0.078/hr74.04$0.29
GeForce RTX 5070 Ti$0.15/hr139.3$0.30
NVIDIA TITAN Xp (power capped)$0.049/hr45.11$0.30
NVIDIA GeForce RTX 5060$0.090/hr77.26$0.32
NVIDIA GeForce RTX 3080 Ti$0.18/hr140.7$0.36
NVIDIA Quadro RTX 5000$0.096/hr72.21$0.37
NVIDIA GeForce RTX 4070$0.12/hr90.94$0.38
GeForce RTX 5080$0.21/hr153.9$0.38
NVIDIA GeForce RTX 4070 Ti Super$0.16/hr117.4$0.38
NVIDIA TITAN V$0.14/hr98.34$0.38
NVIDIA RTX A5000$0.16/hr114.8$0.39
NVIDIA Titan RTX$0.15/hr104.4$0.40
NVIDIA GeForce RTX 5090$0.39/hr243.9$0.44
NVIDIA GeForce RTX 4080$0.20/hr123.7$0.45
NVIDIA GeForce RTX 3090 Ti$0.27/hr156.5$0.48
GeForce RTX 5070$0.19/hr109.6$0.48
GeForce RTX 4080 Super$0.22/hr126.4$0.49
GeForce RTX 5060 Ti$0.14/hr76.63$0.49
NVIDIA GeForce RTX 4070 Ti$0.18/hr89.83$0.54
NVIDIA RTX PRO 4000 Blackwell$0.20/hr101.6$0.56
NVIDIA GeForce RTX 4090$0.34/hr164.3$0.57
NVIDIA RTX A6000$0.33/hr120.2$0.76
NVIDIA A100 40GB PCIe$0.45/hr146.9$0.85
NVIDIA RTX 4000 (Ada Generation)$0.20/hr65.39$0.85
NVIDIA Quadro RTX 8000$0.26/hr81.7$0.87
NVIDIA RTX 5880 Ada Generation$0.49/hr152$0.89
NVIDIA A100 40GB SXM4$0.47/hr146.2$0.90
NVIDIA T4$0.14/hr36.51$1.03
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition$1.00/hr221.3$1.26
NVIDIA RTX 4500 Ada Generation$0.36/hr76.96$1.31
NVIDIA RTX PRO 6000 Blackwell Workstation Edition$1.08/hr225.9$1.32
NVIDIA L40S$0.79/hr130.3$1.68
NVIDIA A100 80GB SXM4$0.95/hr150$1.75
NVIDIA H100 80GB HBM3$2.14/hr244.2$2.43
NVIDIA L4$0.44/hr48.95$2.50
NVIDIA H100 PCIe$1.94/hr185.9$2.89
NVIDIA H100 NVL$2.59/hr221.8$3.24
NVIDIA H200$3.59/hr248$4.02
NVIDIA B200$5.98/hr253.4$6.55
NVIDIA B300$6.94/hr261.4$7.37

Cheapest hourly rate we track on RunPod and Vast.ai, divided by the measured speed. Startup time and storage are extra.

Speed tiers for Qwen3 8B. 30+ tok/s: 57 (RTX 5090, RTX 4090, RTX 3090 Ti). 30 tok/s is roughly where replies outpace reading.

Reading your prompt. Before Qwen3 8B writes anything it reads the input: 13393.1 tok/s on the RTX 5090 (0.3s for a 4,000-token prompt), 11239.7 on the RTX 4090 (0.4s), 144.9 on the GTX 1660 Super (27.6s). Long documents and big code files feel this number more than the generation speed.

VRAM for Qwen3 8B. Measured peak 4.7GB, so 8GB is the smallest common card size; smallest card it ran on: GTX 1660 Ti (6GB). With long context: Q4_K_M 6GB (tested), Q2_K 4GB, Q3_K_M 5GB, Q5_K_M 7GB, Q6_K 9GB.

Power on Qwen3 8B. Most efficient: H100 80GB HBM3, 174W, 0.20 kWh per 1M generated tokens. Hungriest: RTX 3090 Ti, 291W, 0.52 kWh. At $0.15/kWh: $0.030 per 1M generated tokens.

Our verdict

Qwen3 8B at Q4_K_M: 261 tok/s on the B300, ~6GB floor, and the H200 doing 95% of the B300's speed at less than half the power. Our position: this is the base tier for user-facing chat, the model class the budget APIs serve, and the natural first model for anyone with an 8GB card.

FAQ

Can an 8GB card run Qwen3 8B?
Yes: we measured ~6GB peak at Q4_K_M, leaving headroom for context on an 8GB card. The $179 Intel Arc A580 clears the floor; so does any RTX 3050 or better.
Is Qwen3 8B good enough for a real product?
In our view it's the baseline where user-facing chat stops feeling like a toy. Below it (1.7B/4B) you're in pipeline-component territory; above it (14B/30B-A3B) you're buying quality with VRAM.
What's the most efficient GPU for Qwen3 8B?
The H200: 248 tok/s at 127W measured, 1.95 tok/W, double the efficiency of the B300 that beats it by only 5%. For sustained serving, that difference is your power bill.
Qwen3 8B or Qwen3 14B?
8B needs ~6GB and does 261 tok/s peak in our tests; 14B needs ~10GB and tops out at 165 tok/s. If your card has 12GB+, the 14B's quality is worth the speed loss; on an 8GB card the 8B is the honest fit.
How does it compare to Llama 3.1 8B?
Same weight class and similar throughput characteristics. Both are bandwidth-bound at Q4_K_M. Qwen3 has the newer training recipe; we benchmark both so you can compare tok/s per card directly on their pages.