Qwen2.5 7B LoRA needs roughly 20GB of VRAM at our tested precision. A model that doesn't fit doesn't run slowly, it doesn't run. Here's every card from our bench, sorted by measured speed.
Weights: Qwen/Qwen2.5-7B
| GPU | Speed | VRAM | Basis |
|---|---|---|---|
| NVIDIA B300 | 14323.1 train tok/s | 288GB | ✓ Measured |
| NVIDIA B200 | 13722 train tok/s | 192GB | ✓ Measured |
| NVIDIA H200 | 8855.4 train tok/s | 141GB | ✓ Measured |
| NVIDIA H100 80GB HBM3 | 8514.3 train tok/s | 80GB | ✓ Measured |
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition | 6041.8 train tok/s | 96GB | ✓ Measured |
| NVIDIA A100 80GB SXM4 | 3900.6 train tok/s | 80GB | ✓ Measured |
| NVIDIA L40S | 3736 train tok/s | 48GB | ✓ Measured |
| NVIDIA A100 40GB SXM4 | 3670.8 train tok/s | 40GB | Est. |
| NVIDIA A10G | 1239 train tok/s | 24GB | ✓ Measured |
| NVIDIA L4 | 1063.9 train tok/s | 24GB | ✓ Measured |
These cards' VRAM is below the requirement at tested precision, published as data, not hidden.
| GPU | VRAM |
|---|---|
| NVIDIA T4 | 16GB |
Qwen2.5 7B LoRA needs roughly 20GB of VRAM at our tested precision. 10 of the GPUs on our bench run it; 1 don't have the VRAM for it.
NVIDIA B300 is the fastest card we've measured running Qwen2.5 7B LoRA, at 14323.1 train tok/s.
NVIDIA L4 ($2,500 MSRP) is the cheapest card on our bench that runs Qwen2.5 7B LoRA, at 1063.9 train tok/s.