September 21, 2026 · RTX 4070 Ti SUPER 16GB + RTX 5090 32GB
Qwen Image 2.1: INT4 + FP4
I tested INT4 on a 16GB 4070 Ti SUPER and FP4 on a 32GB 5090. On the 5090, FP4 generates 1024 × 1024 images in 4.76 seconds at 25 steps, versus 10.18 seconds for BF16. Text and local edits hold up well, but faces, hands and layouts still change. I wouldn’t call it lossless.
Look at the images
Switch scenarios and step counts here; open any image at full size.
At 25 steps FP4 gives the ladle to the woman in mustard instead of the silver-haired man. At 40 steps the man holds it.
Exact prompt · seed 21011 · 1024 × 1024
Documentary photograph for a community cooking program, wide eye-level composition in a bright, practical teaching kitchen. Exactly six adults, all visible from at least the waist up. At the left counter, a woman with short black hair in a green apron chops carrots on a wooden board, with one hand holding the knife and the other safely curled on the carrot. At center, a silver-haired man in a white apron holds a ladle over a large soup pot. Immediately to his right, a woman in a mustard cardigan passes a rectangular tray of four bread rolls to a younger man in a navy shirt, who receives it with both hands; their hands meet opposite edges of the same tray. At the far right, a man in a red apron rinses a colander of leafy greens at the sink. Behind the central counter, a woman wearing glasses and a purple sweater reads a paper recipe card. Natural varied expressions, clear individual faces, plausible arms and hands attached to their owners, no extra people, realistic morning window light and stainless-steel reflections. No readable signage. The image should feel like a real coordinated cooking session, not a posed group portrait.
All 12 scenarios from my RTX 5090 test. Both models remain GPU-resident, sharing the NF4 encoder and BF16 VAE. Edit inputs are the same frozen sources used in the earlier test. Full-size comparisons are lossless WebP. Same-seed runs can vary.
What I quantized
BF16 is 16-bit floating point. I compressed the original image transformer using 4-bit INT4 on Ada or native FP4 on Blackwell, with BF16 rank 128 branches. The stored transformer tensors are about two-thirds smaller.
| Transformer | Original BF16 | INT4 · Ada | FP4 · Blackwell |
|---|---|---|---|
| Tensor storage | 13.25 GiB | 4.34 GiB | 4.54 GiB |
These sizes cover transformer tensors, not the complete pipeline or VRAM. Every BF16 comparison here keeps the same NF4 text/vision encoder and BF16 VAE as its quantized counterpart. It is not a fully BF16 pipeline comparison.
RTX 5090: FP4 vs BF16
I ran both versions on the same 5090 with the entire pipeline in GPU memory. FP4 is 2.14× faster for generation at 25 steps and 2.17× at 40.
| Workload | Steps | BF16 reference | FP4 rank 128 | Speedup |
|---|---|---|---|---|
| Generation · median of 8 | 25 | 10.18s | 4.76s | 2.14× |
| Generation · median of 8 | 40 | 16.06s | 7.40s | 2.17× |
| One-reference edit · median of 3 | 25 | 12.02s | 6.43s | 1.87× |
| One-reference edit · median of 3 | 40 | 18.64s | 9.81s | 1.90× |
| Two-reference portrait · one run | 25 | 13.94s | 8.20s | 1.70× |
| Two-reference portrait · one run | 40 | 21.36s | 12.34s | 1.73× |
1024 × 1024 engine inference, eager execution, CFG 1, per-request prefix KV caching, no prompt/reference cache hits. Excludes warmup, PNG encoding, upload, queue and network time. Ratios compare matched timings, not equal image quality.
Highest sampled board memory: FP4 18.74 GiB (19,192 MiB), BF16 27.38 GiB (28,038 MiB). These are observed peaks, not guaranteed limits.
RunPod serverless: actual requests
I also sent the same 24 cases to an RTX 5090 serverless endpoint. These are the requests after the first startup, with the server’s default prompt/reference caches enabled and the platform’s queue delays included.
| Workload | Steps | Execution | Client wall time |
|---|---|---|---|
| Generation · median of 8 | 25 | 6.94s | 8.75s |
| Generation · median of 8 | 40 | 9.58s | 11.25s |
| One-reference edit · median of 3 | 25 | 10.08s | 11.67s |
| One-reference edit · median of 3 | 40 | 13.08s | 14.97s |
| Two-reference portrait · one run | 25 | 11.79s | 14.01s |
| Two-reference portrait · one run | 40 | 16.17s | 18.32s |
Execution includes image encoding and upload. Client wall time also includes queueing, polling and network time. Two workers served these requests; queue delays reached 22.13 seconds, so these are not guaranteed warm-worker timings.
One cold-start example: 349.15 seconds queued, 7.50 seconds executing, 357.57 seconds from the client. The endpoint could scale to zero between requests.
RTX 4070 Ti SUPER: INT4 vs BF16
On the 16GB card, I streamed the BF16 transformer from CPU memory. INT4 generation is 1.77× faster at 25 steps and 1.87× at 40. This older test uses a larger prompt set and different offloading, so it is not a controlled GPU-to-GPU speed comparison.
| Workload | Steps | BF16 reference | INT4 rank 128 | Speedup |
|---|---|---|---|---|
| Generation · median of 14 | 25 | 27.51s | 15.50s | 1.77× |
| Generation · median of 14 | 40 | 42.76s | 22.82s | 1.87× |
| One-reference edit · median of 6 | 25 | 31.98s | 19.48s | 1.64× |
| One-reference edit · median of 6 | 40 | 48.70s | 28.43s | 1.71× |
| Two-reference portrait · one run | 25 | 36.79s | 23.76s | 1.55× |
| Two-reference portrait · one run | 40 | 55.28s | 34.35s | 1.61× |
Serial 1024 × 1024 engine inference, model loaded; excludes startup and file I/O. No prompt/reference LRU reuse; per-request prefix KV retained. INT4 sampled board memory peaked at 11.71 GiB (11,994 MiB).
Where it still falls short
The BF16 references also make mistakes: keys land on the chair, diagrams miss required objects and a cup relocation fails. Quantization can change faces, poses and object placement on top of that. Forty steps is the official starting recommendation; it doesn’t reliably fix counts or relationships.
Test setup and scope
The 4070 test covers 21 scenarios at 25 and 40 steps across BF16, calibrated rank 128 and experimental selective rank 512: 126 images. The 5090 test covers 12 of those scenarios with resident BF16 and FP4 at both step counts: 48 images. The gallery shows those 12 scenarios from both tests, 120 images in total. Edit inputs are identical frozen source files across tests.
I kept INT4 rank 128 over the larger selective rank 512 experiment: lower denoiser error didn’t translate into better images. On the original 18-image set, rank 128 improved median SSIM from 0.7608 to 0.8508 versus my first conversion. SSIM measures resemblance, not a percentage of quality.
These are diagnostic comparisons; earlier results informed calibration. I haven’t benchmarked native 2K, compared head-to-head with Klein, or established production readiness.
Qwen Research License: research and evaluation use; commercial use needs a separate license.

