LoRA fine-tuning · 4 models · 11 GPUs measured first-party · Updated October 2026
Which graphics card to use for LoRA fine-tuning, from first-party measurements of Qwen2.5 1.5B LoRA, SmolLM2 1.7B LoRA, TinyLlama 1.1B LoRA and more on 11 GPUs.

22351.1 train tok/s on Qwen2.5 1.5B LoRA, the ceiling. Measured on our bench. 288GB of VRAM, $40,000 at launch.

15753.9 train tok/s on Qwen2.5 1.5B LoRA, lowest launch price that still fits. Measured on our bench. 96GB of VRAM, $8,565 at launch.
LoRA fine-tuning teaches an existing language model new behaviour by training a small set of extra weights instead of the whole model. It is how most people customise a model on their own data, and unlike chatting with a model it keeps the GPU at full load for the whole run, often for hours.
We measured 4 models for LoRA fine-tuning on 11 GPUs. Speed is training tokens processed per second. Every number below is a first-party run on our own harness; cards absent from a model's chart have not been run on it yet.
Qwen2.5 1.5B LoRA: train tok/s by GPU
SmolLM2 1.7B LoRA: train tok/s by GPU
TinyLlama 1.1B LoRA: train tok/s by GPU
Qwen2.5 7B LoRA: train tok/s by GPU
Which models fit which card, for LoRA fine-tuning
| Model | VRAM used | 8GB card | 12GB card | 16GB card | 24GB card | 32GB card | Licence |
|---|---|---|---|---|---|---|---|
| Qwen2.5 1.5B LoRA | 8.2GB | No | Yes | Yes | Yes | Yes | Apache-2.0 |
| SmolLM2 1.7B LoRA | 8.2GB | No | Yes | Yes | Yes | Yes | Apache-2.0 |
| TinyLlama 1.1B LoRA | 4GB | Yes | Yes | Yes | Yes | Yes | Apache-2.0 |
| Qwen2.5 7B LoRA | 20.6GB | No | No | No | Yes | Yes | Apache-2.0 |
From the lowest VRAM peak we measured for each model, plus 5% headroom. 'No' means it did not fit in that much memory at our settings, not that no setting ever could.
Every GPU x every LoRA fine-tuning model (train tok/s)
| GPU | Qwen2.5 1.5B | SmolLM2 1.7B | TinyLlama 1.1B | Qwen2.5 7B |
|---|---|---|---|---|
| NVIDIA B300 | 22351.1 | 26114.3 | 26938.2 | 14323.1 |
| NVIDIA H100 80GB HBM3 | 16034.9 | 18910.3 | 17522.0 | 8514.3 |
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | 15753.9 | 16255.4 | 18556.4 | 6041.8 |
| NVIDIA B200 | 15322.0 | 18064.7 | 17993.5 | 13722.0 |
| NVIDIA H200 | 14931.7 | 17650.4 | 16222.4 | 8855.4 |
| NVIDIA L40S | 12120.1 | 12052.1 | 16408.1 | 3736.0 |
| NVIDIA A100 80GB SXM4 | 8252.8 | 9750.8 | 9553.3 | 3900.6 |
| NVIDIA A100 40GB SXM4 | 6349.7 | 6938.8 | 7471.0 | 3670.8 |
| NVIDIA A10G | 4214.4 | 3668.3 | 5499.6 | 1239.0 |
| NVIDIA L4 | 3720.9 | 3240.2 | 4756.0 | 1063.9 |
| NVIDIA T4 | 1434.4 | 1561.9 | 1520.1 | — |
— = not measured on that card yet.
What the numbers show.
Qwen2.5 1.5B LoRA: fastest on the NVIDIA B300 at 22351.1 train tok/s, 15.58x the slowest card we measured (NVIDIA T4); it used about 8.2GB of VRAM.
SmolLM2 1.7B LoRA: fastest on the NVIDIA B300 at 26114.3 train tok/s, 16.72x the slowest card we measured (NVIDIA T4); it used about 8.2GB of VRAM.
TinyLlama 1.1B LoRA: fastest on the NVIDIA B300 at 26938.2 train tok/s, 17.72x the slowest card we measured (NVIDIA T4); it used about 4GB of VRAM.
Qwen2.5 7B LoRA: fastest on the NVIDIA B300 at 14323.1 train tok/s, 13.46x the slowest card we measured (NVIDIA L4); it used about 20.6GB of VRAM.
Which model to pick. On the same card, the NVIDIA B300, TinyLlama 1.1B LoRA runs at 26938.2 train tok/s in about 4GB; SmolLM2 1.7B LoRA runs at 26114.3 train tok/s in about 8.2GB; Qwen2.5 1.5B LoRA runs at 22351.1 train tok/s in about 8.2GB; Qwen2.5 7B LoRA runs at 14323.1 train tok/s in about 20.6GB. TinyLlama 1.1B LoRA gets through the work 1.9x as fast as Qwen2.5 7B LoRA, so the model you choose moves the speed as much as the card does.
How it compares. B300: Qwen2.5 1.5B LoRA 22351.1 train tok/s, SmolLM2 1.7B LoRA 26114.3, TinyLlama 1.1B LoRA 26938.2, Qwen2.5 7B LoRA 14323.1. 2 of 3 beat Qwen2.5 1.5B LoRA here.
Cost on a rented GPU. 100M training tokens of Qwen2.5 1.5B LoRA: $1.81 on a L40S ($0.79/hr, 2.3 hours), $8.62 on a B300 ($6.94/hr, 75 min, 4.8x the cost).
Qwen2.5 1.5B LoRA: cost per 100M training tokens on rented GPUs
| GPU | Cheapest rate | Speed (train tok/s) | Cost per 100M training tokens |
|---|---|---|---|
| NVIDIA L40S | $0.79/hr | 12120.1 | $1.81 |
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | $1.08/hr | 15753.9 | $1.90 |
| NVIDIA A100 40GB SXM4 | $0.47/hr | 6349.7 | $2.06 |
| NVIDIA T4 | $0.14/hr | 1434.4 | $2.63 |
| NVIDIA A100 80GB SXM4 | $0.95/hr | 8252.8 | $3.19 |
| NVIDIA L4 | $0.44/hr | 3720.9 | $3.28 |
| NVIDIA H100 80GB HBM3 | $2.14/hr | 16034.9 | $3.70 |
| NVIDIA H200 | $3.59/hr | 14931.7 | $6.68 |
| NVIDIA B300 | $6.94/hr | 22351.1 | $8.62 |
| NVIDIA B200 | $5.98/hr | 15322 | $10.84 |
Cheapest hourly rate we track on RunPod and Vast.ai, divided by the measured speed. Startup time and storage are extra.
Speed tiers for Qwen2.5 1.5B LoRA. 10000+ train tok/s: 6 (B300, H100 80GB HBM3, RTX PRO 6000 Blackwell Workstation Edition); 3000-10000 train tok/s: 4 (A100 80GB SXM4, A100 40GB SXM4, A10G); under 3000 train tok/s: 1 (T4). 10,000 train tok/s gets through 100M tokens in under three hours.
VRAM for Qwen2.5 1.5B LoRA. Measured peak 7.8GB, so 12GB is the smallest common card size; smallest card it ran on: T4 (16GB).
Power on Qwen2.5 1.5B LoRA. Most efficient: B300, 371W, 0.46 kWh per 100M training tokens. Hungriest: B200, 388W, 0.70 kWh. At $0.15/kWh: $0.069 per 100M training tokens.
For LoRA fine-tuning, the NVIDIA B300 is the fastest card we measured. We have not measured a consumer card on this job yet; the picks above are datacenter and workstation hardware. Check the fit table before buying: VRAM, not speed, is what rules a card out.
Each model runs a fixed workload on every card: a warmup, then timed runs with power, temperature and VRAM sampled every half second through NVML. Models run at the precision and settings from their model card. Datacenter cards run on Modal; consumer cards on rented machines. Non-commercially licensed models are not part of this page.