TRELLIS.2 Image-to-3D · 8 GPUs measured first-party · image-to-3D · Updated October 2026
TRELLIS.2 is the best open image-to-3D model we've run: one picture in, a textured GLB out. We timed it the way you'd use it for a batch of game assets: a worker that's already warm, at the model's default 1024 setting, all the way to a web-ready GLB (100K faces, 1024px textures). An H100 or H200 makes one every 40 seconds, about 90 an hour. It runs on gaming cards too: 63 seconds per asset on an RTX 4090 (about 57 an hour) and 120 on an RTX 3090 (about 30 an hour), warm, at the default setting, against 117 on an A10G and 153 on an L4.
Benchmarked weights: microsoft/TRELLIS.2-4B

90.1 assets/hour on TRELLIS.2 Image-to-3D, the ceiling. Measured on our bench. 141GB of VRAM, $31,000 at launch.

57.2 assets/hour on TRELLIS.2 Image-to-3D, fastest card you can buy at retail. Measured on our bench. 24GB of VRAM, $1,599 at launch.

30.1 assets/hour on TRELLIS.2 Image-to-3D, lowest launch price that still fits. Measured on our bench. 24GB of VRAM, $1,499 at launch.
What GPU Do You Need for TRELLIS.2 Image-to-3D?, assets/hour by GPU
Measured on our own bench. A card absent from this chart has not been run on this model yet, or cannot fit it.
Value: assets/hour per $1,000 of MSRP
Launch price, not street price, so it ages. A speed leaderboard always crowns the most expensive card; this is the counterweight.
TRELLIS.2 Image-to-3D: measured speed per finished asset, by GPU
| GPU | Assets/hour | s per asset | Peak VRAM | Avg power |
|---|---|---|---|---|
| NVIDIA H200 | 90.1 | 39.97 | 13.8GB | — |
| NVIDIA H100 80GB HBM3 | 87.9 | 40.94 | 13.8GB | — |
| NVIDIA L40S | 66.7 | 53.96 | 13.9GB | — |
| NVIDIA GeForce RTX 4090 | 57.2 | 62.9 | 14.1GB | — |
| NVIDIA A100 80GB SXM4 | 43.4 | 82.97 | 14.4GB | — |
| NVIDIA A10G | 30.8 | 117.02 | 14.1GB | — |
| NVIDIA GeForce RTX 3090 | 30.1 | 119.63 | 13.7GB | — |
| NVIDIA L4 | 23.6 | 152.64 | 14.1GB | — |
Warm beats many. The slow part of a fresh TRELLIS.2 worker isn't the model: it's the sparse-convolution kernels tuning themselves for each new mesh size. On a fresh H100 the shape decode took 25-34 seconds for the first assets and under 1 second once the worker had seen about ten. So one warm GPU working through a queue beats many cold ones, which each pay the 2-minute load and the tuning again: sending seven parts of one asset to seven fresh GPUs pays that warm-up seven times.
Planning numbers. Warm, default setting, to a finished GLB: H200 40.0 seconds per asset, H100 40.9, L40S 54.0, A100 83.0. The H200's extra memory doesn't help, and the A100 falls behind the newer L40S. The GLB export is about a third of every asset (9-26 seconds, mostly CPU work, about the same on every card), which is why running it alongside the next generation pays. At the cheapest rates we track, a finished asset costs about 1.2 cents on an L40S, 2.2 on an A100, 2.4 on an H100 (1.7 with the export overlapped) and 4.0 on an H200.
About TRELLIS.2 Image-to-3D. TRELLIS.2 Image-to-3D: from microsoft, 4.0B parameters, on Hugging Face since December 2025, MIT licence. 1,903,026 downloads in the last 30 days and 39 community quantizations.
How it compares. RTX 4090: TRELLIS.2 Image-to-3D 57.2 assets/hour, Stable Fast 3D 9113.9 (1B), Hunyuan3D 2.0 12.27, Hunyuan3D 2mini 19.81, Shap-E (image to 3D) 1309.1. 2 of 4 beat TRELLIS.2 Image-to-3D here. These aren't like-for-like: TripoSR returns a rough vertex-coloured mesh and TripoSG bare geometry, while both TRELLIS numbers are timed to a finished, textured GLB.
Cost on a rented GPU. 1,000 assets of TRELLIS.2 Image-to-3D: $4.05 on a RTX 3090 ($0.12/hr, 33.2 hours), $39.84 on a H200 ($3.59/hr, 11.1 hours, 9.8x the cost).
TRELLIS.2 Image-to-3D: cost per 1,000 assets on rented GPUs
| GPU | Cheapest rate | Speed (assets/hour) | Cost per 1,000 assets |
|---|---|---|---|
| NVIDIA GeForce RTX 3090 | $0.12/hr | 30.1 | $4.05 |
| NVIDIA GeForce RTX 4090 | $0.34/hr | 57.2 | $5.87 |
| NVIDIA L40S | $0.79/hr | 66.7 | $11.84 |
| NVIDIA L4 | $0.44/hr | 23.6 | $18.64 |
| NVIDIA A100 80GB SXM4 | $0.95/hr | 43.4 | $21.82 |
| NVIDIA H100 80GB HBM3 | $2.14/hr | 87.9 | $24.30 |
| NVIDIA H200 | $3.59/hr | 90.1 | $39.84 |
Cheapest hourly rate we track on RunPod and Vast.ai, divided by the measured speed. Startup time and storage are extra.
Speed tiers for TRELLIS.2 Image-to-3D. 60-360 assets/hour: 3 (H200, H100 80GB HBM3, L40S); under 60 assets/hour: 5 (RTX 4090, RTX 3090). 360 assets/hour is ten seconds per mesh.
VRAM for TRELLIS.2 Image-to-3D. Measured peak 13.7GB, so 16GB is the smallest common card size; smallest card it ran on: RTX 4090 (24GB). That peak is PyTorch's own; the GLB exporter's mesh tools allocate on top of it and ran out of memory on a 24GB A10G until we freed PyTorch's cache first. Microsoft lists 24GB as the minimum, and that's the size to plan for.
For a batch of game assets, rent one L40S or H100 and keep it running. The L40S is the cheapest per asset we measured (about 1.2 cents), the H100 the fastest per hour you'd pay for. The H200's extra memory buys nothing here. Draft at 512, pick, then make finals at the default setting.
Model microsoft/TRELLIS.2-4B, 12 sampling steps per stage, five game-asset pictures (rifle, pistol, magazine, crate, tree) as transparent cut-outs, so the non-commercial background remover is never used. Each card first makes all five untimed (kernel tuning happens here), then the same five are timed one after another: image prep, generation, decode and GLB export (o_voxel to_glb, 100K faces, 1024px textures). Peak VRAM is the most PyTorch allocated during the timed pass. Separate runs measured cold start, the 512 and 1536 settings, drafts in one pass and an export running alongside generation. Datacenter cards on Modal, October 2026.