LoRA fine-tuning · 4 models · 11 GPUs measured first-party · Updated October 2026

Best GPU for LoRA fine-tuning

Which graphics card to use for LoRA fine-tuning, from first-party measurements of Qwen2.5 1.5B LoRA, SmolLM2 1.7B LoRA, TinyLlama 1.1B LoRA and more on 11 GPUs.

Fastest we measured
NVIDIA B300

NVIDIA B300

22351.1 train tok/s on Qwen2.5 1.5B LoRA, the ceiling. Measured on our bench. 288GB of VRAM, $40,000 at launch.

Pros
  • 22351.1 train tok/s on Qwen2.5 1.5B LoRA
  • 288GB, clears the Qwen2.5 1.5B LoRA floor
  • Rentable by the hour rather than bought
Cons
  • 1400W board rating
  • Datacenter or workstation hardware, not a retail purchase
Cheapest card that runs it
NVIDIA RTX PRO 6000 Blackwell Workstation Edition

NVIDIA RTX PRO 6000 Blackwell Workstation Edition

15753.9 train tok/s on Qwen2.5 1.5B LoRA, lowest launch price that still fits. Measured on our bench. 96GB of VRAM, $8,565 at launch.

Pros
  • 15753.9 train tok/s on Qwen2.5 1.5B LoRA
  • 96GB, clears the Qwen2.5 1.5B LoRA floor
  • Rentable by the hour rather than bought
Cons
  • 600W board rating
  • Datacenter or workstation hardware, not a retail purchase
4
Models measured
Qwen2.5 1.5B LoRA, SmolLM2 1.7B LoRA, TinyLlama 1.1B LoRA and more
11
GPUs measured
first-party runs, not spec-sheet estimates
26938.2train tok/s
Fastest: NVIDIA B300
on TinyLlama 1.1B LoRA
4GB
Lightest model's VRAM need
measured peak, +5% headroom

LoRA fine-tuning teaches an existing language model new behaviour by training a small set of extra weights instead of the whole model. It is how most people customise a model on their own data, and unlike chatting with a model it keeps the GPU at full load for the whole run, often for hours.

We measured 4 models for LoRA fine-tuning on 11 GPUs. Speed is training tokens processed per second. Every number below is a first-party run on our own harness; cards absent from a model's chart have not been run on it yet.

Qwen2.5 1.5B LoRA: train tok/s by GPU

NVIDIA B300
22351.1 train tok/s
NVIDIA H100 80GB HBM3
16034.9 train tok/s
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
15753.9 train tok/s
NVIDIA B200
15322 train tok/s
NVIDIA H200
14931.7 train tok/s
NVIDIA L40S
12120.1 train tok/s
NVIDIA A100 80GB SXM4
8252.8 train tok/s
NVIDIA A100 40GB SXM4
6349.7 train tok/s
NVIDIA A10G
4214.4 train tok/s
NVIDIA L4
3720.9 train tok/s
NVIDIA T4
1434.4 train tok/s

SmolLM2 1.7B LoRA: train tok/s by GPU

NVIDIA B300
26114.3 train tok/s
NVIDIA H100 80GB HBM3
18910.3 train tok/s
NVIDIA B200
18064.7 train tok/s
NVIDIA H200
17650.4 train tok/s
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
16255.4 train tok/s
NVIDIA L40S
12052.1 train tok/s
NVIDIA A100 80GB SXM4
9750.8 train tok/s
NVIDIA A100 40GB SXM4
6938.8 train tok/s
NVIDIA A10G
3668.3 train tok/s
NVIDIA L4
3240.2 train tok/s
NVIDIA T4
1561.9 train tok/s

TinyLlama 1.1B LoRA: train tok/s by GPU

NVIDIA B300
26938.2 train tok/s
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
18556.4 train tok/s
NVIDIA B200
17993.5 train tok/s
NVIDIA H100 80GB HBM3
17522 train tok/s
NVIDIA L40S
16408.1 train tok/s
NVIDIA H200
16222.4 train tok/s
NVIDIA A100 80GB SXM4
9553.3 train tok/s
NVIDIA A100 40GB SXM4
7471 train tok/s
NVIDIA A10G
5499.6 train tok/s
NVIDIA L4
4756 train tok/s
NVIDIA T4
1520.1 train tok/s

Qwen2.5 7B LoRA: train tok/s by GPU

NVIDIA B300
14323.1 train tok/s
NVIDIA B200
13722 train tok/s
NVIDIA H200
8855.4 train tok/s
NVIDIA H100 80GB HBM3
8514.3 train tok/s
NVIDIA RTX PRO 6000 Blackwell Workstation Edition
6041.8 train tok/s
NVIDIA A100 80GB SXM4
3900.6 train tok/s
NVIDIA L40S
3736 train tok/s
NVIDIA A100 40GB SXM4
3670.8 train tok/s
NVIDIA A10G
1239 train tok/s
NVIDIA L4
1063.9 train tok/s

Which models fit which card, for LoRA fine-tuning

Qwen2.5 1.5B LoRA8.2GB
SmolLM2 1.7B LoRA8.2GB
TinyLlama 1.1B LoRA4GB
Qwen2.5 7B LoRA20.6GB
ModelVRAM used8GB card12GB card16GB card24GB card32GB cardLicence
Qwen2.5 1.5B LoRA8.2GBNoYesYesYesYesApache-2.0
SmolLM2 1.7B LoRA8.2GBNoYesYesYesYesApache-2.0
TinyLlama 1.1B LoRA4GBYesYesYesYesYesApache-2.0
Qwen2.5 7B LoRA20.6GBNoNoNoYesYesApache-2.0

From the lowest VRAM peak we measured for each model, plus 5% headroom. 'No' means it did not fit in that much memory at our settings, not that no setting ever could.

Every GPU x every LoRA fine-tuning model (train tok/s)

NVIDIA B30022351.1
NVIDIA H100 80GB HBM316034.9
NVIDIA RTX PRO 6000 Blackwell Workstation Edition15753.9
NVIDIA B20015322.0
NVIDIA H20014931.7
NVIDIA L40S12120.1
NVIDIA A100 80GB SXM48252.8
NVIDIA A100 40GB SXM46349.7
NVIDIA A10G4214.4
NVIDIA L43720.9
NVIDIA T41434.4
GPUQwen2.5 1.5BSmolLM2 1.7BTinyLlama 1.1BQwen2.5 7B
NVIDIA B30022351.126114.326938.214323.1
NVIDIA H100 80GB HBM316034.918910.317522.08514.3
NVIDIA RTX PRO 6000 Blackwell Workstation Edition15753.916255.418556.46041.8
NVIDIA B20015322.018064.717993.513722.0
NVIDIA H20014931.717650.416222.48855.4
NVIDIA L40S12120.112052.116408.13736.0
NVIDIA A100 80GB SXM48252.89750.89553.33900.6
NVIDIA A100 40GB SXM46349.76938.87471.03670.8
NVIDIA A10G4214.43668.35499.61239.0
NVIDIA L43720.93240.24756.01063.9
NVIDIA T41434.41561.91520.1—

— = not measured on that card yet.

What the numbers show.

Qwen2.5 1.5B LoRA: fastest on the NVIDIA B300 at 22351.1 train tok/s, 15.58x the slowest card we measured (NVIDIA T4); it used about 8.2GB of VRAM.

SmolLM2 1.7B LoRA: fastest on the NVIDIA B300 at 26114.3 train tok/s, 16.72x the slowest card we measured (NVIDIA T4); it used about 8.2GB of VRAM.

TinyLlama 1.1B LoRA: fastest on the NVIDIA B300 at 26938.2 train tok/s, 17.72x the slowest card we measured (NVIDIA T4); it used about 4GB of VRAM.

Qwen2.5 7B LoRA: fastest on the NVIDIA B300 at 14323.1 train tok/s, 13.46x the slowest card we measured (NVIDIA L4); it used about 20.6GB of VRAM.

Which model to pick. On the same card, the NVIDIA B300, TinyLlama 1.1B LoRA runs at 26938.2 train tok/s in about 4GB; SmolLM2 1.7B LoRA runs at 26114.3 train tok/s in about 8.2GB; Qwen2.5 1.5B LoRA runs at 22351.1 train tok/s in about 8.2GB; Qwen2.5 7B LoRA runs at 14323.1 train tok/s in about 20.6GB. TinyLlama 1.1B LoRA gets through the work 1.9x as fast as Qwen2.5 7B LoRA, so the model you choose moves the speed as much as the card does.

How it compares. B300: Qwen2.5 1.5B LoRA 22351.1 train tok/s, SmolLM2 1.7B LoRA 26114.3, TinyLlama 1.1B LoRA 26938.2, Qwen2.5 7B LoRA 14323.1. 2 of 3 beat Qwen2.5 1.5B LoRA here.

Cost on a rented GPU. 100M training tokens of Qwen2.5 1.5B LoRA: $1.81 on a L40S ($0.79/hr, 2.3 hours), $8.62 on a B300 ($6.94/hr, 75 min, 4.8x the cost).

Qwen2.5 1.5B LoRA: cost per 100M training tokens on rented GPUs

NVIDIA L40S$0.79/hr
NVIDIA RTX PRO 6000 Blackwell Workstation Edition$1.08/hr
NVIDIA A100 40GB SXM4$0.47/hr
NVIDIA T4$0.14/hr
NVIDIA A100 80GB SXM4$0.95/hr
NVIDIA L4$0.44/hr
NVIDIA H100 80GB HBM3$2.14/hr
NVIDIA H200$3.59/hr
NVIDIA B300$6.94/hr
NVIDIA B200$5.98/hr
GPUCheapest rateSpeed (train tok/s)Cost per 100M training tokens
NVIDIA L40S$0.79/hr12120.1$1.81
NVIDIA RTX PRO 6000 Blackwell Workstation Edition$1.08/hr15753.9$1.90
NVIDIA A100 40GB SXM4$0.47/hr6349.7$2.06
NVIDIA T4$0.14/hr1434.4$2.63
NVIDIA A100 80GB SXM4$0.95/hr8252.8$3.19
NVIDIA L4$0.44/hr3720.9$3.28
NVIDIA H100 80GB HBM3$2.14/hr16034.9$3.70
NVIDIA H200$3.59/hr14931.7$6.68
NVIDIA B300$6.94/hr22351.1$8.62
NVIDIA B200$5.98/hr15322$10.84

Cheapest hourly rate we track on RunPod and Vast.ai, divided by the measured speed. Startup time and storage are extra.

Speed tiers for Qwen2.5 1.5B LoRA. 10000+ train tok/s: 6 (B300, H100 80GB HBM3, RTX PRO 6000 Blackwell Workstation Edition); 3000-10000 train tok/s: 4 (A100 80GB SXM4, A100 40GB SXM4, A10G); under 3000 train tok/s: 1 (T4). 10,000 train tok/s gets through 100M tokens in under three hours.

VRAM for Qwen2.5 1.5B LoRA. Measured peak 7.8GB, so 12GB is the smallest common card size; smallest card it ran on: T4 (16GB).

Power on Qwen2.5 1.5B LoRA. Most efficient: B300, 371W, 0.46 kWh per 100M training tokens. Hungriest: B200, 388W, 0.70 kWh. At $0.15/kWh: $0.069 per 100M training tokens.

Our verdict

For LoRA fine-tuning, the NVIDIA B300 is the fastest card we measured. We have not measured a consumer card on this job yet; the picks above are datacenter and workstation hardware. Check the fit table before buying: VRAM, not speed, is what rules a card out.

FAQ

What is the fastest GPU for LoRA fine-tuning?
In our runs, the NVIDIA B300 at 26938.2 train tok/s on TinyLlama 1.1B LoRA. We measured 4 models on 11 GPUs for this page.
How much VRAM do I need for LoRA fine-tuning?
The lightest model here, TinyLlama 1.1B LoRA, used about 4GB. The table above shows which models fit 8, 12, 16, 24 and 32GB cards, from measured peaks.
Are these numbers measured or estimated?
Measured. Every number on this page is a first-party run on our own harness, with power and VRAM sampled during the run. Cards we have not run yet are simply absent, not filled in.
What GPU do I need to run Qwen2.5 1.5B LoRA?
About 8GB. Smallest card that ran it: NVIDIA T4 (16GB).
How much does it cost to run Qwen2.5 1.5B LoRA in the cloud?
$1.81 per 100M training tokens on a NVIDIA L40S at $0.79/hr, cheapest of 10 rentable cards we measured.
Can I run Qwen2.5 1.5B LoRA on a 12GB, 16GB or 24GB card?
It used 7.8GB at the precision we tested. 12GB: yes; 16GB: yes; 24GB: yes.
Is the H100 80GB HBM3 or the A100 80GB SXM4 faster for Qwen2.5 1.5B LoRA?
The H100 80GB HBM3: 16034.9 vs 8252.8 train tok/s, 94% faster on our bench.

How we test

Each model runs a fixed workload on every card: a warmup, then timed runs with power, temperature and VRAM sampled every half second through NVML. Models run at the precision and settings from their model card. Datacenter cards run on Modal; consumer cards on rented machines. Non-commercially licensed models are not part of this page.