All benchmarks

September 21, 2026 · RTX 4070 Ti SUPER 16GB + RTX 5090 32GB

Qwen Image 2.1: INT4 + FP4

I tested INT4 on a 16GB 4070 Ti SUPER and FP4 on a 32GB 5090. On the 5090, FP4 generates 1024 × 1024 images in 4.76 seconds at 25 steps, versus 10.18 seconds for BF16. Text and local edits hold up well, but faces, hands and layouts still change. I wouldn’t call it lossless.

4.76s
RTX 5090 FP4 · 25 steps
2.14×
5090 speedup vs BF16 · 25 steps
4.54 GiB
FP4 transformer tensors
120 images
Browse both hardware tests

Look at the images

Switch scenarios and step counts here; open any image at full size.

At 25 steps FP4 gives the ladle to the woman in mustard instead of the silver-haired man. At 40 steps the man holds it.

BF16 transformer reference · 40 steps
Six people, six jobs — BF16 transformer reference, 40 steps
FP4 rank 128 · 40 steps
Six people, six jobs — FP4 rank 128, 40 steps
Exact prompt · seed 21011 · 1024 × 1024

Documentary photograph for a community cooking program, wide eye-level composition in a bright, practical teaching kitchen. Exactly six adults, all visible from at least the waist up. At the left counter, a woman with short black hair in a green apron chops carrots on a wooden board, with one hand holding the knife and the other safely curled on the carrot. At center, a silver-haired man in a white apron holds a ladle over a large soup pot. Immediately to his right, a woman in a mustard cardigan passes a rectangular tray of four bread rolls to a younger man in a navy shirt, who receives it with both hands; their hands meet opposite edges of the same tray. At the far right, a man in a red apron rinses a colander of leafy greens at the sink. Behind the central counter, a woman wearing glasses and a purple sweater reads a paper recipe card. Natural varied expressions, clear individual faces, plausible arms and hands attached to their owners, no extra people, realistic morning window light and stainless-steel reflections. No readable signage. The image should feel like a real coordinated cooking session, not a posed group portrait.

All 12 scenarios from my RTX 5090 test. Both models remain GPU-resident, sharing the NF4 encoder and BF16 VAE. Edit inputs are the same frozen sources used in the earlier test. Full-size comparisons are lossless WebP. Same-seed runs can vary.

What I quantized

BF16 is 16-bit floating point. I compressed the original image transformer using 4-bit INT4 on Ada or native FP4 on Blackwell, with BF16 rank 128 branches. The stored transformer tensors are about two-thirds smaller.

TransformerOriginal BF16INT4 · AdaFP4 · Blackwell
Tensor storage13.25 GiB4.34 GiB4.54 GiB

These sizes cover transformer tensors, not the complete pipeline or VRAM. Every BF16 comparison here keeps the same NF4 text/vision encoder and BF16 VAE as its quantized counterpart. It is not a fully BF16 pipeline comparison.

RTX 5090: FP4 vs BF16

I ran both versions on the same 5090 with the entire pipeline in GPU memory. FP4 is 2.14× faster for generation at 25 steps and 2.17× at 40.

WorkloadStepsBF16 referenceFP4 rank 128Speedup
Generation · median of 82510.18s4.76s2.14×
Generation · median of 84016.06s7.40s2.17×
One-reference edit · median of 32512.02s6.43s1.87×
One-reference edit · median of 34018.64s9.81s1.90×
Two-reference portrait · one run2513.94s8.20s1.70×
Two-reference portrait · one run4021.36s12.34s1.73×

1024 × 1024 engine inference, eager execution, CFG 1, per-request prefix KV caching, no prompt/reference cache hits. Excludes warmup, PNG encoding, upload, queue and network time. Ratios compare matched timings, not equal image quality.

Highest sampled board memory: FP4 18.74 GiB (19,192 MiB), BF16 27.38 GiB (28,038 MiB). These are observed peaks, not guaranteed limits.

RunPod serverless: actual requests

I also sent the same 24 cases to an RTX 5090 serverless endpoint. These are the requests after the first startup, with the server’s default prompt/reference caches enabled and the platform’s queue delays included.

WorkloadStepsExecutionClient wall time
Generation · median of 8256.94s8.75s
Generation · median of 8409.58s11.25s
One-reference edit · median of 32510.08s11.67s
One-reference edit · median of 34013.08s14.97s
Two-reference portrait · one run2511.79s14.01s
Two-reference portrait · one run4016.17s18.32s

Execution includes image encoding and upload. Client wall time also includes queueing, polling and network time. Two workers served these requests; queue delays reached 22.13 seconds, so these are not guaranteed warm-worker timings.

One cold-start example: 349.15 seconds queued, 7.50 seconds executing, 357.57 seconds from the client. The endpoint could scale to zero between requests.

RTX 4070 Ti SUPER: INT4 vs BF16

On the 16GB card, I streamed the BF16 transformer from CPU memory. INT4 generation is 1.77× faster at 25 steps and 1.87× at 40. This older test uses a larger prompt set and different offloading, so it is not a controlled GPU-to-GPU speed comparison.

WorkloadStepsBF16 referenceINT4 rank 128Speedup
Generation · median of 142527.51s15.50s1.77×
Generation · median of 144042.76s22.82s1.87×
One-reference edit · median of 62531.98s19.48s1.64×
One-reference edit · median of 64048.70s28.43s1.71×
Two-reference portrait · one run2536.79s23.76s1.55×
Two-reference portrait · one run4055.28s34.35s1.61×

Serial 1024 × 1024 engine inference, model loaded; excludes startup and file I/O. No prompt/reference LRU reuse; per-request prefix KV retained. INT4 sampled board memory peaked at 11.71 GiB (11,994 MiB).

Where it still falls short

The BF16 references also make mistakes: keys land on the chair, diagrams miss required objects and a cup relocation fails. Quantization can change faces, poses and object placement on top of that. Forty steps is the official starting recommendation; it doesn’t reliably fix counts or relationships.

The same seed can change details. A repeated 5090 FP4 kitchen image changed hand and ladle placement with application caches off. BF16 was byte-identical in that one check. INT4 also failed strict repeatability in the earlier test; these checks do not establish general determinism.
Test setup and scope

The 4070 test covers 21 scenarios at 25 and 40 steps across BF16, calibrated rank 128 and experimental selective rank 512: 126 images. The 5090 test covers 12 of those scenarios with resident BF16 and FP4 at both step counts: 48 images. The gallery shows those 12 scenarios from both tests, 120 images in total. Edit inputs are identical frozen source files across tests.

I kept INT4 rank 128 over the larger selective rank 512 experiment: lower denoiser error didn’t translate into better images. On the original 18-image set, rank 128 improved median SSIM from 0.7608 to 0.8508 versus my first conversion. SSIM measures resemblance, not a percentage of quality.

These are diagnostic comparisons; earlier results informed calibration. I haven’t benchmarked native 2K, compared head-to-head with Klein, or established production readiness.

Qwen Research License: research and evaluation use; commercial use needs a separate license.

Download weights ↗

Official Qwen Image 2.1 pipeline documentation ↗