Image-to-3D · measured on our own runs · October 2026 · Updated October 2026
Text prompt to picture, picture to textured 3D model: that's the open-source pipeline that works for game assets right now, and every step can run on one rented GPU. We timed each step on six datacenter cards with real game-asset prompts (a rifle, a pistol, a magazine, a crate and a tree) to answer one question: how fast, and how cheaply, can you make a whole kit? The short answer is that keeping one GPU busy matters far more than which GPU you rent.
The pipeline. 1) Make a picture of each asset with a fast image model. We used Z-Image Turbo (Apache-2.0) with prompts like "a sci-fi assault rifle, game asset, single object isolated on a plain white background, three-quarter view, soft studio lighting". 2) Cut the object out. Because we asked for a plain white background, a simple flood fill did it with no model at all. That matters: TRELLIS.2's built-in remover, RMBG-2.0, is licensed for non-commercial use only, and TRELLIS.2 skips it when the picture already has a transparent background. 3) Make quick 3D drafts and keep the shapes you like. 4) Make the finals and export each one as a GLB a game engine can load.
TRELLIS.2, seconds per finished asset (warm worker, default setting, to GLB)
Mean over five game assets, each from picture to a web-ready GLB (100K faces, 1024px textures). Lower is faster.
Mistake 1: starting lots of fresh GPUs. It's tempting to send each asset (or each part of one) to its own GPU. A fresh TRELLIS.2 worker spends about 2 minutes loading and then tunes its GPU kernels for each new shape size: on an H100 the first asset took 194 seconds and the shape decode took 25-34 seconds per asset until the worker had seen about ten, then under 1 second. Every fresh GPU pays that again. One warm GPU working through a queue is both faster per asset and cheaper.
Mistake 2: ignoring the export. Turning TRELLIS.2's output into a GLB (simplifying the mesh, baking the textures) took 9-26 seconds per asset on every card we tried, about a third of each asset's time, and it's mostly CPU work, so a faster GPU doesn't shorten it. Run the export on a second thread while the next asset generates: on a warm H100 that took a 12-asset queue from 88 assets an hour to 123. With the older TRELLIS v1 it's even more lopsided: 3-7 seconds to generate, and 70-85% of each asset's time in the export.
Mistake 3: making finals before you've picked a shape. TRELLIS.2's 512 setting makes a rough draft in about 6 seconds on a warm H100 or H200 (8 on an L40S), against 11-41 seconds for a final at the default 1024 setting. Make drafts, keep the shapes you like, and only make finals of those. You can even ask for 4 drafts in one pass: on an H100 that cost 1.25x the time of one.
Mistake 4: paying for memory you don't use. At its default setting TRELLIS.2 peaked at about 14GB in our runs (18-20GB at its 1536 setting), so the H200's 141GB buys nothing: it was no faster than the H100 (40.0 against 40.9 seconds) at 1.7x the rental price. The L40S, at 54 seconds an asset for $0.79 an hour, is the cheapest per asset. The older A100 is the one to skip: 83 seconds, slower than the cheaper L40S.
A 100-asset kit on one rented GPU, measured step times
Measured in $0.79/hr. Longer is faster.
| Step | L40S ($0.79/hr) | H100 ($2.14/hr) |
|---|---|---|
| Load and warm up TRELLIS.2 (5 assets) | 9 min | 9 min |
| 200 pictures, Z-Image Turbo, 1024px (2 per asset) | 24 min (7.2s each) | 9 min (2.7s each) |
| 100 drafts, TRELLIS.2 512 | 14 min (8.4s each) | 10 min (6.1s each) |
| 100 finals to GLB, TRELLIS.2 default | 90 min (54s each) | 68 min (41s each); 49 min with export overlapped |
| Total GPU time | about 2.3 hours | about 1.6 hours; 1.3 with overlap |
| Cost at the cheapest rates we track | about $1.80 | about $3.40; $2.75 with overlap |
Step times are our measured means on each card. Model loading for the image model and the time you spend choosing drafts are extra; batch the choosing so the GPU isn't idle while you look.
What about my own GPU? Z-Image Turbo makes 7.4 pictures a minute on an RTX 4090 and 10.7 on an RTX 5090, so step 1 is quick at home. TRELLIS.2 runs at home on a 24GB card: 63 seconds per asset on an RTX 4090 (about 57 an hour) and 120 on an RTX 3090 (about 30 an hour), against 54 on a rented L40S. Microsoft lists 24GB as its minimum, and that holds: our 4090 and 3090 runs peaked at about 14GB while generating, but on a 16GB RTX A4000 the GLB export ran out of memory. TRELLIS v1 needs about 12GB and makes a textured GLB in 52 seconds even on an L4, so it's the one to try first on a 16GB card.
Generate pictures with a plain white background, cut them out without a model, draft at 512, then make finals at TRELLIS.2's default setting on one warm GPU with the export running alongside. A 100-asset kit is about 2.3 hours and $1.80 on an L40S, or 1.3 hours on an H100.
Z-Image Turbo speeds are our measured 1024px image runs. TRELLIS.2 (microsoft/TRELLIS.2-4B) runs at 12 sampling steps per stage on five game-asset pictures given as transparent cut-outs: each card makes all five untimed first, then the same five are timed one after another, picture to GLB (o_voxel to_glb, 100K faces, 1024px textures). The cold-start, queue and overlapped-export figures come from separate H100 runs of 12 assets. Datacenter cards on Modal; rental rates are the cheapest we track on RunPod and Vast.ai, October 2026.