TRELLIS.2 Image-to-3D · 8 GPUs measured first-party · image-to-3D · Updated October 2026

What GPU Do You Need for TRELLIS.2 Image-to-3D?

TRELLIS.2 is the best open image-to-3D model we've run: one picture in, a textured GLB out. We timed it the way you'd use it for a batch of game assets: a worker that's already warm, at the model's default 1024 setting, all the way to a web-ready GLB (100K faces, 1024px textures). An H100 or H200 makes one every 40 seconds, about 90 an hour. It runs on gaming cards too: 63 seconds per asset on an RTX 4090 (about 57 an hour) and 120 on an RTX 3090 (about 30 an hour), warm, at the default setting, against 117 on an A10G and 153 on an L4.

Benchmarked weights: microsoft/TRELLIS.2-4B

Fastest we measured
NVIDIA H200

NVIDIA H200

90.1 assets/hour on TRELLIS.2 Image-to-3D, the ceiling. Measured on our bench. 141GB of VRAM, $31,000 at launch.

Pros
  • 90.1 assets/hour on TRELLIS.2 Image-to-3D
  • 141GB, clears the TRELLIS.2 Image-to-3D floor
  • Rentable by the hour rather than bought
Cons
  • 700W board rating
  • Datacenter or workstation hardware, not a retail purchase
Best consumer card
NVIDIA GeForce RTX 4090

NVIDIA GeForce RTX 4090

57.2 assets/hour on TRELLIS.2 Image-to-3D, fastest card you can buy at retail. Measured on our bench. 24GB of VRAM, $1,599 at launch.

Pros
  • 57.2 assets/hour on TRELLIS.2 Image-to-3D
  • 24GB, clears the TRELLIS.2 Image-to-3D floor
Cons
  • 450W board rating
Cheapest card that runs it
NVIDIA GeForce RTX 3090

NVIDIA GeForce RTX 3090

30.1 assets/hour on TRELLIS.2 Image-to-3D, lowest launch price that still fits. Measured on our bench. 24GB of VRAM, $1,499 at launch.

Pros
  • 30.1 assets/hour on TRELLIS.2 Image-to-3D
  • 24GB, clears the TRELLIS.2 Image-to-3D floor
Cons
  • 350W board rating
90.1assets/hour
Fastest: NVIDIA H200
measured, warm, to a finished GLB
8
GPUs measured
first-party, warm worker
~15GB
VRAM needed
measured peak at the default setting, +5% headroom
194s → 41s
First asset vs warm, H100
kernel tuning on a fresh worker

What GPU Do You Need for TRELLIS.2 Image-to-3D?, assets/hour by GPU

NVIDIA H200
90.1 assets/hour
NVIDIA H100 80GB HBM3
87.9 assets/hour
NVIDIA L40S
66.7 assets/hour
NVIDIA GeForce RTX 4090
57.2 assets/hour
NVIDIA A100 80GB SXM4
43.4 assets/hour
NVIDIA A10G
30.8 assets/hour
NVIDIA GeForce RTX 3090
30.1 assets/hour
NVIDIA L4
23.6 assets/hour

Measured on our own bench. A card absent from this chart has not been run on this model yet, or cannot fit it.

Value: assets/hour per $1,000 of MSRP

NVIDIA GeForce RTX 4090
35.77 assets/hour / $1k
NVIDIA GeForce RTX 3090
20.08 assets/hour / $1k
NVIDIA A10G
11 assets/hour / $1k
NVIDIA L4
9.44 assets/hour / $1k
NVIDIA L40S
8.89 assets/hour / $1k
NVIDIA H100 80GB HBM3
2.93 assets/hour / $1k
NVIDIA H200
2.91 assets/hour / $1k
NVIDIA A100 80GB SXM4
2.55 assets/hour / $1k

Launch price, not street price, so it ages. A speed leaderboard always crowns the most expensive card; this is the counterweight.

TRELLIS.2 Image-to-3D: measured speed per finished asset, by GPU

NVIDIA H20090.1
NVIDIA H100 80GB HBM387.9
NVIDIA L40S66.7
NVIDIA GeForce RTX 409057.2
NVIDIA A100 80GB SXM443.4
NVIDIA A10G30.8
NVIDIA GeForce RTX 309030.1
NVIDIA L423.6
GPUAssets/hours per assetPeak VRAMAvg power
NVIDIA H20090.139.9713.8GB—
NVIDIA H100 80GB HBM387.940.9413.8GB—
NVIDIA L40S66.753.9613.9GB—
NVIDIA GeForce RTX 409057.262.914.1GB—
NVIDIA A100 80GB SXM443.482.9714.4GB—
NVIDIA A10G30.8117.0214.1GB—
NVIDIA GeForce RTX 309030.1119.6313.7GB—
NVIDIA L423.6152.6414.1GB—

Warm beats many. The slow part of a fresh TRELLIS.2 worker isn't the model: it's the sparse-convolution kernels tuning themselves for each new mesh size. On a fresh H100 the shape decode took 25-34 seconds for the first assets and under 1 second once the worker had seen about ten. So one warm GPU working through a queue beats many cold ones, which each pay the 2-minute load and the tuning again: sending seven parts of one asset to seven fresh GPUs pays that warm-up seven times.

Planning numbers. Warm, default setting, to a finished GLB: H200 40.0 seconds per asset, H100 40.9, L40S 54.0, A100 83.0. The H200's extra memory doesn't help, and the A100 falls behind the newer L40S. The GLB export is about a third of every asset (9-26 seconds, mostly CPU work, about the same on every card), which is why running it alongside the next generation pays. At the cheapest rates we track, a finished asset costs about 1.2 cents on an L40S, 2.2 on an A100, 2.4 on an H100 (1.7 with the export overlapped) and 4.0 on an H200.

About TRELLIS.2 Image-to-3D. TRELLIS.2 Image-to-3D: from microsoft, 4.0B parameters, on Hugging Face since December 2025, MIT licence. 1,903,026 downloads in the last 30 days and 39 community quantizations.

How it compares. RTX 4090: TRELLIS.2 Image-to-3D 57.2 assets/hour, Stable Fast 3D 9113.9 (1B), Hunyuan3D 2.0 12.27, Hunyuan3D 2mini 19.81, Shap-E (image to 3D) 1309.1. 2 of 4 beat TRELLIS.2 Image-to-3D here. These aren't like-for-like: TripoSR returns a rough vertex-coloured mesh and TripoSG bare geometry, while both TRELLIS numbers are timed to a finished, textured GLB.

Cost on a rented GPU. 1,000 assets of TRELLIS.2 Image-to-3D: $4.05 on a RTX 3090 ($0.12/hr, 33.2 hours), $39.84 on a H200 ($3.59/hr, 11.1 hours, 9.8x the cost).

TRELLIS.2 Image-to-3D: cost per 1,000 assets on rented GPUs

NVIDIA GeForce RTX 3090$0.12/hr
NVIDIA GeForce RTX 4090$0.34/hr
NVIDIA L40S$0.79/hr
NVIDIA L4$0.44/hr
NVIDIA A100 80GB SXM4$0.95/hr
NVIDIA H100 80GB HBM3$2.14/hr
NVIDIA H200$3.59/hr
GPUCheapest rateSpeed (assets/hour)Cost per 1,000 assets
NVIDIA GeForce RTX 3090$0.12/hr30.1$4.05
NVIDIA GeForce RTX 4090$0.34/hr57.2$5.87
NVIDIA L40S$0.79/hr66.7$11.84
NVIDIA L4$0.44/hr23.6$18.64
NVIDIA A100 80GB SXM4$0.95/hr43.4$21.82
NVIDIA H100 80GB HBM3$2.14/hr87.9$24.30
NVIDIA H200$3.59/hr90.1$39.84

Cheapest hourly rate we track on RunPod and Vast.ai, divided by the measured speed. Startup time and storage are extra.

Speed tiers for TRELLIS.2 Image-to-3D. 60-360 assets/hour: 3 (H200, H100 80GB HBM3, L40S); under 60 assets/hour: 5 (RTX 4090, RTX 3090). 360 assets/hour is ten seconds per mesh.

VRAM for TRELLIS.2 Image-to-3D. Measured peak 13.7GB, so 16GB is the smallest common card size; smallest card it ran on: RTX 4090 (24GB). That peak is PyTorch's own; the GLB exporter's mesh tools allocate on top of it and ran out of memory on a 24GB A10G until we freed PyTorch's cache first. Microsoft lists 24GB as the minimum, and that's the size to plan for.

Our verdict

For a batch of game assets, rent one L40S or H100 and keep it running. The L40S is the cheapest per asset we measured (about 1.2 cents), the H100 the fastest per hour you'd pay for. The H200's extra memory buys nothing here. Draft at 512, pick, then make finals at the default setting.

FAQ

What GPU do you need for TRELLIS.2?
At the default 1024 setting it peaked at about 14GB in our runs, and 18-20GB at 1536. Microsoft lists 24GB as the minimum. It runs on gaming cards too: 63 seconds per asset on an RTX 4090 (about 57 an hour) and 120 on an RTX 3090 (about 30 an hour), warm, at the default setting, against 117 on an A10G and 153 on an L4. The cards worth renting for a batch are the L40S (cheapest per asset) and the H100 (fastest).
How long does one asset take?
On a warm H100: 11-41 seconds to generate depending on the shape (a magazine is quick, a tree slow), plus 9-24 seconds to export the GLB. 41 seconds on average, 88 an hour. The very first asset on a fresh worker took 194 seconds, because TRELLIS.2 tunes its GPU kernels for each new shape size.
How do I make a batch of assets fastest?
Keep one worker running instead of starting many: each fresh one spends 2 minutes loading and its first few assets tuning kernels, so a fresh H100 took about 8 minutes for its first 5 assets. Once warm, run the GLB export on a second thread while the next asset generates: in our 12-asset queue test that took an H100 to 123 assets an hour, against 88 one after another.
Does the 512 setting help?
As a draft, yes: about 6 seconds to generate on an H100 or H200 once warm, and 15 seconds with the GLB. Look at the drafts, keep the shapes you like, then make the finals at 1024. Asking for 4 drafts in one pass cost 1.25x the time of one on an H100 (2.8x on an L40S).
Is TRELLIS.2 worth it over v1?
For anything people look at closely, yes: the shapes and textures are clearly better. Measured to a finished GLB on the same H100, v1 is about 1.9x faster (167 against 88 an hour) and needs about 12GB, so it still has a place for background props.
Why are these numbers different from before?
Our July run timed the 1536 setting once on a cold worker, so most of what it measured was kernel tuning, and it multiplied one asset by ten to get a batch time. Since October 2026 we time the default setting on a warm worker over five real game assets, each to a finished GLB.

How we test

Model microsoft/TRELLIS.2-4B, 12 sampling steps per stage, five game-asset pictures (rifle, pistol, magazine, crate, tree) as transparent cut-outs, so the non-commercial background remover is never used. Each card first makes all five untimed (kernel tuning happens here), then the same five are timed one after another: image prep, generation, decode and GLB export (o_voxel to_glb, 100K faces, 1024px textures). Peak VRAM is the most PyTorch allocated during the timed pass. Separate runs measured cold start, the 512 and 1536 settings, drafts in one pass and an export running alongside generation. Datacenter cards on Modal, October 2026.