Qwen3 1.7B · 11 GPUs measured first-party · llama.cpp Q4_K_M · Updated October 2026

What GPU Do You Need for Qwen3 1.7B?

Qwen3 1.7B sits in an awkward spot we think is worth being honest about: it's past the under-1B line where running on the user's device is realistic, but well short of the 4B tier where chat quality starts. We measured it on 11 GPUs (llama.cpp, Q4_K_M): 576 tok/s on the RTX PRO 6000 Blackwell at the top, 136 tok/s on a T4 at the bottom, ~3GB peak VRAM everywhere.

Benchmarked weights: Qwen/Qwen3-1.7B-GGUF

Fastest we measured
NVIDIA RTX PRO 6000 Blackwell Workstation Edition

NVIDIA RTX PRO 6000 Blackwell Workstation Edition

575.8 tok/s on Qwen3 1.7B, the ceiling. Measured on our bench. 96GB of VRAM, $8,565 at launch.

Pros
  • 575.8 tok/s on Qwen3 1.7B
  • 96GB, clears the Qwen3 1.7B floor
  • Rentable by the hour rather than bought
Cons
  • 600W board rating
  • Datacenter or workstation hardware, not a retail purchase
575.8tok/s
Fastest: NVIDIA RTX PRO 6000 Blackwell Workstation Edition
measured
11
Cards that run Qwen3 1.7B
of 11 we have data for
0
Cards that can't run it at all
published as hard gates, not omissions
311%
Fastest vs slowest that fits
575.8 vs 139.9 tok/s

What GPU Do You Need for Qwen3 1.7B?, tok/s by GPU

NVIDIA RTX PRO 6000 Blackwell Workstation Edition
575.8 tok/s
NVIDIA H200
563.5 tok/s
NVIDIA H100 80GB HBM3
556.5 tok/s
NVIDIA B300
555.1 tok/s
NVIDIA B200
551.8 tok/s
NVIDIA L40S
405.5 tok/s
NVIDIA A100 80GB SXM4
343.6 tok/s
NVIDIA A100 40GB SXM4
340.8 tok/s
NVIDIA A10G
267.1 tok/s
NVIDIA L4
177.5 tok/s
NVIDIA T4
139.9 tok/s

Measured on our own bench. A card absent from this chart has not been run on this model yet, or cannot fit it.

Efficiency: tok/s per 100W drawn

NVIDIA H200
561.27 tok/s / 100W
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
550.45 tok/s / 100W
NVIDIA L4
348.04 tok/s / 100W
NVIDIA A100 80GB SXM4
337.5 tok/s / 100W
NVIDIA H100 80GB HBM3
315.11 tok/s / 100W
NVIDIA A100 40GB SXM4
305.61 tok/s / 100W
NVIDIA L40S
283.58 tok/s / 100W
NVIDIA A10G
267.35 tok/s / 100W
NVIDIA T4
245.91 tok/s / 100W
NVIDIA B300
202.15 tok/s / 100W
NVIDIA B200
198.69 tok/s / 100W

Power is the average pulled during the run, sampled at 1Hz. The fastest card is often not the one here, and for anything left running this is the number that shows up on the bill.

Value: tok/s per $1,000 of MSRP

NVIDIA A10G
95.39 tok/s / $1k
NVIDIA L4
71 tok/s / $1k
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
67.22 tok/s / $1k
NVIDIA T4
60.86 tok/s / $1k
NVIDIA L40S
54.07 tok/s / $1k
NVIDIA A100 40GB SXM4
28.4 tok/s / $1k
NVIDIA A100 80GB SXM4
20.21 tok/s / $1k
NVIDIA H100 80GB HBM3
18.55 tok/s / $1k
NVIDIA H200
18.18 tok/s / $1k
NVIDIA B300
13.88 tok/s / $1k
NVIDIA B200
13.79 tok/s / $1k

Launch price, not street price, so it ages. A speed leaderboard always crowns the most expensive card; this is the counterweight.

Qwen3 1.7B. Measured generation speed by GPU

NVIDIA RTX PRO 6000 Blackwell Workstation Edition575.8
NVIDIA H200563.5
NVIDIA H100 80GB HBM3556.5
NVIDIA B300555.1
NVIDIA B200551.8
NVIDIA L40S405.5
NVIDIA A100 80GB SXM4343.6
NVIDIA A100 40GB SXM4340.8
NVIDIA A10G267.1
NVIDIA L4177.5
NVIDIA T4139.9
GPUtok/sPrompt t/stok/WAvg power
NVIDIA RTX PRO 6000 Blackwell Workstation Edition575.827571.15.5104.6 W
NVIDIA H200563.521109.95.61100.4 W
NVIDIA H100 80GB HBM3556.522702.33.15176.6 W
NVIDIA B300555.115682.52.02274.6 W
NVIDIA B200551.824383.11.99277.7 W
NVIDIA L40S405.526431.52.84143.0 W
NVIDIA A100 80GB SXM4343.611762.33.38101.8 W
NVIDIA A100 40GB SXM4340.810340.53.06111.5 W
NVIDIA A10G267.110871.12.6799.9 W
NVIDIA L4177.5111453.4851.0 W
NVIDIA T4139.94593.72.4656.9 W

Where a 1.7B model actually earns its keep. My rule of thumb: over 1B parameters you can no longer assume the user's machine can run it, the compute floor of a cheap device kills the experience even when the VRAM technically fits. So 1.7B isn't a 'runs everywhere' model and it isn't a chat model either. Where it fits is inside pipelines: turning images descriptions into tags, normalizing scraped text, pre-classifying requests before a bigger model sees them. Work where it's a component in a workflow, not the thing the user talks to. For that, 576 tok/s on a single card means one GPU can feed a very large pipeline.

Reading the results. The top three cards land within 4% of each other (576, 564, 556 tok/s), at 1.7B the model still can't stress big silicon, so H100-class hardware buys you almost nothing over the workstation card. Note the efficiency column: the H200 does 5.61 tok/W at just 100W, and even the L4, a 72W card, clears 176 tok/s. If this model is your batch workhorse, a small efficient card at near-idle power is the right tool; renting anything bigger is paying for headroom the model can't use.

About Qwen3 1.7B. Qwen3 1.7B: from Qwen, 2.0B parameters, on Hugging Face since April 2025, Apache 2.0 licence. 3,897,790 downloads in the last 30 days and 2 community quantizations.

How it compares. H100 80GB HBM3: Qwen3 1.7B 556.5 tok/s, DeepSeek-R1 Distill 1.5B 537.3 (2B), Qwen2.5-1.5B 537.2 (2B), Qwen2.5-Coder-1.5B 538.4 (2B), Qwen2-1.5B 536.5 (2B). Qwen3 1.7B beats all 4 here.

Cost on a rented GPU. 1M generated tokens of Qwen3 1.7B: $0.27 on a T4 ($0.14/hr, 119 min), $0.52 on a RTX PRO 6000 Blackwell Workstation Edition ($1.08/hr, 29 min, 1.9x the cost).

Qwen3 1.7B: cost per 1M generated tokens on rented GPUs

NVIDIA T4$0.14/hr
NVIDIA A100 40GB SXM4$0.47/hr
NVIDIA RTX PRO 6000 Blackwell Workstation Edition$1.08/hr
NVIDIA L40S$0.79/hr
NVIDIA L4$0.44/hr
NVIDIA A100 80GB SXM4$0.95/hr
NVIDIA H100 80GB HBM3$2.14/hr
NVIDIA H200$3.59/hr
NVIDIA B200$5.98/hr
NVIDIA B300$6.94/hr
GPUCheapest rateSpeed (tok/s)Cost per 1M generated tokens
NVIDIA T4$0.14/hr139.9$0.27
NVIDIA A100 40GB SXM4$0.47/hr340.8$0.38
NVIDIA RTX PRO 6000 Blackwell Workstation Edition$1.08/hr575.8$0.52
NVIDIA L40S$0.79/hr405.5$0.54
NVIDIA L4$0.44/hr177.5$0.69
NVIDIA A100 80GB SXM4$0.95/hr343.6$0.77
NVIDIA H100 80GB HBM3$2.14/hr556.5$1.07
NVIDIA H200$3.59/hr563.5$1.77
NVIDIA B200$5.98/hr551.8$3.01
NVIDIA B300$6.94/hr555.1$3.47

Cheapest hourly rate we track on RunPod and Vast.ai, divided by the measured speed. Startup time and storage are extra.

Speed tiers for Qwen3 1.7B. 30+ tok/s: 11 (RTX PRO 6000 Blackwell Workstation Edition, H200, H100 80GB HBM3). 30 tok/s is roughly where replies outpace reading.

Reading your prompt. Before Qwen3 1.7B writes anything it reads the input: 27571.1 tok/s on the RTX PRO 6000 Blackwell Workstation Edition (0.1s for a 4,000-token prompt), 4593.7 on the T4 (0.9s). Long documents and big code files feel this number more than the generation speed.

VRAM for Qwen3 1.7B. Measured peak 1.8GB, so 8GB is the smallest common card size; smallest card it ran on: T4 (16GB). With long context: Q4_K_M 2GB (tested), Q2_K 2GB, Q3_K_M 2GB, Q5_K_M 3GB, Q6_K 3GB.

Power on Qwen3 1.7B. Most efficient: L4, 51W, 79.8 Wh per 1M generated tokens. Hungriest: B200, 278W, 0.14 kWh. At $0.15/kWh: $0.012 per 1M generated tokens.

Our verdict

Qwen3 1.7B: 576 tok/s peak, ~3GB floor, and near-identical results across the top of the chart because the model can't stress big hardware. Our take: too big to assume it runs on user devices, too small for user-facing chat, but as the cheap text-processing stage inside a pipeline, one mid-range GPU running this model does an enormous amount of work.

FAQ

What is Qwen3 1.7B actually good for?
Pipeline work: tagging, extraction, normalization, pre-classification, stages where a model processes text inside a workflow rather than chatting with a user. For user-facing chat, the 4B and 8B tiers are where quality starts.
How much VRAM does Qwen3 1.7B need?
~3GB measured peak at Q4_K_M. Any 4GB card fits it; a $139 GTX 1050 Ti clears the floor on capacity.
Can it run client-side like Qwen3 0.6B?
Not reliably. Once you cross 1B parameters, low-end devices have the memory but not the compute, generation drops below usable speed on weak hardware. If 'runs on the user's machine' is the requirement, stay under 1B.
Do I need a datacenter GPU for it?
No. The chart is unusually flat at the top (top three cards within 4%), because the model can't saturate big silicon. An efficient small card like the L4 (176 tok/s at ~50W measured draw) is the economical choice for sustained batch work.
Qwen3 1.7B or DeepSeek R1 Distill Qwen 1.5B?
Same size class, different personalities: the R1 distill spends tokens on reasoning traces, Qwen3 1.7B answers directly. For pipeline throughput you usually want direct answers, reasoning traces triple your token bill for marginal gain at this scale.