20 prompts, one run per model, no cherry-picking. Qwen-Image-2.1 ran locally on an RTX 3060 12 GB with qwen-image-local (int8 weights, 40 steps, native 2K). gpt-image-2 ran in the cloud through ChatGPT. Edits use real photos from Wikimedia Commons.
Score
Category
Qwen
Tie
gpt-image-2
Photography
0
1
5
Creativity
0
0
6
Photo editing
1
7
0
Total
1
8
11
Short version: for editing photos, Qwen-Image-2.1 is on par with gpt-image-2 and it is free, private and unlimited on your own card. For text-to-image it follows long prompts less closely and knows less about the world (it drew a Chinese gourd for "mate" twice).
Time per image
Run
Average
Qwen · text→image · native 2K, 40 steps
7.2 min
Qwen · edit · 1536 px, 40 steps
5.3 min
Qwen · text→image · 1024², 20 steps
0.6 min
gpt-image-2 · cloud, 3 in parallel
1.2 min
Peak VRAM during a native 2K generation: 11.9 of 12.0 GB. Each model (text encoder 8.9 GB, image model 6.9 GB, VAE 0.6 GB) is loaded fully into VRAM when its stage runs; nothing is split between CPU and GPU.
VRAM profiles
Same prompts on the 12 GB card, then with ComfyUI told to leave 4.3 GB free (--reserve-vram 4.3). That cap is soft: peak usage (desktop included) still went above 8 GB with both profiles, so this does not prove the app runs on a real 8 GB card. What it does show: the compact profile lowers peak VRAM by about 2.5 GB at the same speed. Reproduce with bash bench/tiers/run.sh.
Setup
Test
Total
s/step
Peak VRAM
12 GB card · balanced
Text→image 1024², 20 steps
32 s
1.36
11.4 GB
12 GB card · balanced
Text→image 1.6 MP, 30 steps
81 s
2.53
11.4 GB
12 GB card · balanced
Photo edit 1280 px, 30 steps
133 s
3.76
11.2 GB
Soft 8 GB cap · compact
Text→image 1024², 20 steps
33 s
1.37
8.9 GB
Soft 8 GB cap · compact
Text→image 1.6 MP, 30 steps
82 s
2.54
8.8 GB
Soft 8 GB cap · compact
Photo edit 1280 px, 30 steps
141 s
3.99
9.5 GB
Soft 8 GB cap · balanced
Text→image 1024², 20 steps
34 s
1.42
8.9 GB
Soft 8 GB cap · balanced
Text→image 1.6 MP, 30 steps
84 s
2.61
9.1 GB
Soft 8 GB cap · balanced
Photo edit 1280 px, 30 steps
142 s
4.03
9.1 GB
All 20 tests
Photography · text-to-image 3:4
Portrait: Patagonian fisherman
gpt-image-2 winsPrompt
Close-up editorial portrait of a 70-year-old Patagonian fisherman with deep wrinkles, weathered sun-damaged skin and salt-and-pepper stubble, wearing a navy wool beanie and a yellow oilskin jacket, standing at a windy harbor in Puerto Madryn at golden hour. Shot on an 85mm lens at f/1.8, shallow depth of field, soft rim light, visible skin pores, natural color grading.
Qwen-Image-2.1 · local RTX 30607.5 mingpt-image-2 · cloud1.3 min
Qwen renders extreme skin detail at 2K but overdoes it (HDR-looking wrinkles) and the background is a plain sea. GPT builds a believable harbor with natural skin.
Photography · text-to-image 3:2
Street: San Telmo after rain
TiePrompt
Street photograph of a narrow cobblestone street in San Telmo, Buenos Aires, at night after rain. The wet cobblestones glisten, a black-and-yellow taxi passes with slight motion blur, old colonial facades with wrought-iron balconies, warm sodium streetlights, a couple sharing an umbrella in the distance. 35mm film look, Kodak Portra 800 grain.
Qwen-Image-2.1 · local RTX 30607.4 mingpt-image-2 · cloud1.2 min
Qwen looks like a real documentary street photo but ignores two details (couple 'in the distance', warm sodium light). GPT follows the prompt more closely but reads like a postcard.
Photography · text-to-image 4:3
Food: medialunas and cortado
gpt-image-2 winsPrompt
Overhead food photography: a rustic wooden table with a white ceramic plate holding three glossy medialunas (Argentine croissants), a cortado coffee in a small glass, a folded newspaper, scattered crumbs, morning window light from the left casting soft shadows, 50mm lens, high detail.
Qwen-Image-2.1 · local RTX 30607.4 mingpt-image-2 · cloud1.0 min
GPT knows an Argentine medialuna is small, curved and glazed, and the newspaper is legible. Qwen draws large French croissants, gibberish newsprint and an espresso instead of a cortado.
Photography · text-to-image 3:2
Hands: luthier carving a scroll
gpt-image-2 winsPrompt
Macro photograph of the hands of an elderly luthier carving the scroll of a violin with a small gouge, thin wood shavings curling off, fingers with calluses and visible veins, a workshop full of hanging violins out of focus in the background, warm tungsten light.
Qwen-Image-2.1 · local RTX 30607.4 mingpt-image-2 · cloud1.2 min
Both hands are anatomically fine. Qwen carves an already varnished violin with an odd tool; GPT shows raw maple, a real gouge and plausible shavings.
Photography · text-to-image 16:9
Landscape: Mount Fitz Roy
gpt-image-2 winsPrompt
Landscape photograph of Mount Fitz Roy at sunrise seen from Laguna de los Tres, pink alpenglow on the granite peaks, glacier below, perfectly still turquoise lake with a mirror reflection, two hikers in red jackets small in the foreground for scale, ultra wide 16mm lens, crisp detail.
Qwen-Image-2.1 · local RTX 30607.2 mingpt-image-2 · cloud0.9 min
The prompt asks for a still lake with a mirror reflection; only GPT delivers it. Qwen's Fitz Roy is recognizable but the lake barely reflects.
Photography · text-to-image 1:1
Product: watch on wet slate
gpt-image-2 winsPrompt
Studio product photograph of a matte black mechanical wristwatch with orange accents lying on a slab of wet dark slate, water droplets on the surface, a single dramatic softbox from the top right, pure black background, crisp reflections on the sapphire crystal, commercial advertising quality.
Qwen-Image-2.1 · local RTX 30607.0 mingpt-image-2 · cloud1.1 min
GPT gives ad-ready framing with droplets on the crystal. Qwen leaves the softbox in frame and the watch small, though the wet slate texture is excellent.
Creativity · text-to-image 2:3
Surreal: lighthouse of books
gpt-image-2 winsPrompt
A lighthouse built from towering stacks of old books standing on a rock in an ocean of black ink, its beam made of glowing floating letters, paper birds flying around it, rendered as a detailed ligne claire comic illustration with a limited palette of cream, deep navy and coral.
Qwen-Image-2.1 · local RTX 30607.4 mingpt-image-2 · cloud1.1 min
GPT hits every element: paper birds, ink sea with floating pages, coral palette, a beam of letters. Qwen draws real seagulls and a normal sea.
Creativity · text-to-image 2:3
Poster with Spanish text
gpt-image-2 winsPrompt
Vintage 1950s Argentine travel poster, screen-print style. Big bold title at the top: "MAR DEL PLATA". Subtitle below it: "Veraneá en la Perla del Atlántico". Illustration of the Rambla with its two stone sea lion statues, striped beach umbrellas and a woman in a retro swimsuit. Small text at the bottom: "Ferrocarril Sud — Temporada 1954".
Qwen-Image-2.1 · local RTX 30607.4 mingpt-image-2 · cloud1.3 min
Both render every word perfectly, accents included. GPT also knows the real Rambla (Hotel Provincial, stone sea lions on pedestals) and has more poster punch.
Creativity · text-to-image 1:1
Infographic: how to make mate
gpt-image-2 winsPrompt
A clean flat-design infographic titled "Cómo preparar mate" showing exactly 6 numbered steps in a 2x3 grid, each with a simple icon and a short caption: 1 "Llenar 3/4 con yerba", 2 "Inclinar el mate", 3 "Agua tibia en el hueco", 4 "Colocar la bombilla", 5 "Agua a 75°C", 6 "Cebar y compartir". Warm green and cream palette.
Qwen-Image-2.1 · local RTX 30607.1 mingpt-image-2 · cloud1.4 min
GPT: perfect small text and correct mate icons. Qwen misspells small captions ('Inclinar enj mate') and draws cups and a wine glass instead of a mate gourd.
Creativity · text-to-image 16:9
Character: capybara astronaut
gpt-image-2 winsPrompt
A 3D animated movie still of a chubby capybara astronaut floating inside a cozy spaceship cabin, holding a mate gourd, small potted plants drifting in zero gravity, round portholes showing planet Earth, soft volumetric light, expressive eyes, feature-animation stylization.
Qwen-Image-2.1 · local RTX 30607.2 mingpt-image-2 · cloud1.1 min
Both are feature-animation quality. Qwen reads 'mate gourd' literally and hands the capybara a raw calabash; GPT has it sipping mate through a bombilla.
Creativity · text-to-image 1:1
Counting: 3 apples, 2 pears, 1 banana
gpt-image-2 winsPrompt
Top-down photo of exactly three red apples, two green pears and one banana arranged on a blue ceramic plate, with a silver fork to the left of the plate and a folded yellow napkin to the right, on a white marble table.
Qwen-Image-2.1 · local RTX 30607.0 mingpt-image-2 · cloud1.0 min
GPT gets every count and position right. Qwen draws 2 apples and puts the banana off the plate.
Creativity · text-to-image 1:1
Transparent sticker (RGBA)
gpt-image-2 winsPrompt
This is an RGBA format image with transparency. A cute cartoon sticker of a mate gourd with a smiling face and a metal bombilla, thick white die-cut outline. The image has an alpha channel and a transparent background.
Qwen-Image-2.1 · local RTX 30607.0 mingpt-image-2 · cloud1.1 min
Both return real alpha. Qwen again draws an East Asian hyōtan gourd instead of a mate.
Photo editing · edit
Edit: sunny street → snowy night
TiePrompt
Change the scene to a snowy winter night: snow on the balconies, awnings, trees and street, snowflakes falling, warm lights glowing in the windows. Keep the buildings, signs, people and composition exactly the same.
Input photoMbaro01 · CC BY-SA 4.0Qwen-Image-2.1 · local RTX 30605.1 mingpt-image-2 · cloud1.5 min
Both keep the MARTINEZ sign, people, flower stand and architecture. Qwen is slightly more faithful to the source; GPT adds more of the requested warm light.
Photo editing · edit
Edit: change clothes, add glasses
TiePrompt
Change her denim shirt to a red wool turtleneck sweater and add round tortoiseshell glasses. Keep her face, identity, hair, expression, pose, phone and the background exactly the same.
Input photoAnthony Ginsbrook aginsbrook · CC0Qwen-Image-2.1 · local RTX 30604.9 mingpt-image-2 · cloud1.1 min
Identity, hair, expression, phone and background preserved in both. Professional in both.
Photo editing · edit
Edit: rewrite a shop sign
TiePrompt
Edit the text on the round sign: replace "SOLAR ROAST" with "LA PORTEÑA" and replace "COFFEE" with "CAFÉ", using the same gold serif lettering and the same sign design. Keep the rest of the photo unchanged.
Input photoSolarroast · CC BY-SA 4.0Qwen-Image-2.1 · local RTX 30605.0 mingpt-image-2 · cloud1.3 min
Both replace SOLAR ROAST COFFEE with LA PORTEÑA CAFÉ, Ñ and accent included, in the same gold serif, leaving the logo and the rest of the photo untouched.
Photo editing · edit
Edit: remove all people
TiePrompt
Remove all the people from the sidewalk tables and from inside the café, leaving the chairs and tables empty. Keep everything else identical.
Input photoHenrique Félix henriquefelix · CC0Qwen-Image-2.1 · local RTX 30604.9 mingpt-image-2 · cloud1.3 min
Both remove everyone without ghosts. Qwen keeps the original table layout more exactly; GPT fills the interior with more life.
Photo editing · edit
Edit: photo → watercolor
TiePrompt
Turn this photo into a hand-painted watercolor illustration in the style of classic 1990s Japanese animated films, soft painterly grass and warm light. Keep the composition and the dog's pose.
Input photoDavid Whelan · CC0Qwen-Image-2.1 · local RTX 30604.9 mingpt-image-2 · cloud1.2 min
Both keep composition and pose. Qwen leans toward animation colors, GPT toward loose watercolor texture.
Photo editing · edit
Edit: combine two photos
TiePrompt
Place the golden retriever from <image2> sitting on the sidewalk next to the green bicycle in <image1>, matching the lighting and perspective of <image1>, with a realistic contact shadow. Keep the bicycle and the wall unchanged.
Input photowww.Pixel.la Free Stock Photos · CC0Input photoDavid Whelan · CC0Qwen-Image-2.1 · local RTX 30607.2 mingpt-image-2 · cloud1.2 min
Both place the dog from photo 2 next to the bicycle with a contact shadow and matching light, bicycle unchanged.
Photo editing · edit
Edit: remove background
Qwen winsPrompt
Remove the background, and output a PNG image
Input photowww.Pixel.la Free Stock Photos · CC0Qwen-Image-2.1 · local RTX 30605.0 mingpt-image-2 · cloud1.7 min
Qwen cuts out the exact bicycle at its original position and scale, which is what you need for compositing. GPT redraws it larger and changes details of the rear rack.
Photo editing · edit
Edit: remove an object
TiePrompt
Remove the rifle hanging on the left of the wall and fill in the plaster wall texture naturally. Keep all the picture frames exactly the same.
Input photoWilfredor · CC0Qwen-Image-2.1 · local RTX 30605.1 mingpt-image-2 · cloud1.4 min
Both remove the rifle and rebuild the plaster; every frame is intact in both.
FAQ
What AI image model runs on a 12 GB GPU?
Qwen-Image-2.1 (7B, open weights, September 2026) runs fully in VRAM on a 12 GB card such as an RTX 3060 when you use the int8 weights and load the text encoder, the image model and the VAE one at a time. qwen-image-local does that for you: about 35 seconds for a 1024×1024 image and about 7 minutes at native 2K.
Can Qwen-Image-2.1 run on an 8 GB GPU?
Not verified yet. The compact profile swaps the text encoder for a 6.3 GB w4a8 version and cuts peak VRAM by about 2.5 GB at the same speed, so it should fit 10 GB cards. We could only emulate 8 GB with a soft cap, and peak usage still went above 8 GB, so treat 8 GB cards as untested.
Is Qwen-Image-2.1 as good as gpt-image-2?
For photo editing, yes in our tests: 7 ties and 1 win out of 8 edits. For text-to-image, no: gpt-image-2 won 11 of 12, mostly on prompt adherence and world knowledge.
Is it free to use commercially?
The app is MIT. The Qwen-Image-2.1 weights are under the Qwen Research License, which does not allow commercial use without a separate license from Alibaba.