qwen-image-local banner, generated and then edited by Qwen-Image-2.1 on an RTX 3060

Qwen-Image-2.1 on a 12 GB GPU vs gpt-image-2

20 prompts, one run per model, no cherry-picking. Qwen-Image-2.1 ran locally on an RTX 3060 12 GB with qwen-image-local (int8 weights, 40 steps, native 2K). gpt-image-2 ran in the cloud through ChatGPT. Edits use real photos from Wikimedia Commons.

Score

CategoryQwenTiegpt-image-2
Photography015
Creativity006
Photo editing170
Total1811

Short version: for editing photos, Qwen-Image-2.1 is on par with gpt-image-2 and it is free, private and unlimited on your own card. For text-to-image it follows long prompts less closely and knows less about the world (it drew a Chinese gourd for "mate" twice).

Time per image

RunAverage
Qwen · text→image · native 2K, 40 steps7.2 min
Qwen · edit · 1536 px, 40 steps5.3 min
Qwen · text→image · 1024², 20 steps0.6 min
gpt-image-2 · cloud, 3 in parallel1.2 min

Peak VRAM during a native 2K generation: 11.9 of 12.0 GB. Each model (text encoder 8.9 GB, image model 6.9 GB, VAE 0.6 GB) is loaded fully into VRAM when its stage runs; nothing is split between CPU and GPU.

VRAM profiles

Same prompts on the 12 GB card, then with ComfyUI told to leave 4.3 GB free (--reserve-vram 4.3). That cap is soft: peak usage (desktop included) still went above 8 GB with both profiles, so this does not prove the app runs on a real 8 GB card. What it does show: the compact profile lowers peak VRAM by about 2.5 GB at the same speed. Reproduce with bash bench/tiers/run.sh.

SetupTestTotals/stepPeak VRAM
12 GB card · balancedText→image 1024², 20 steps32 s1.3611.4 GB
12 GB card · balancedText→image 1.6 MP, 30 steps81 s2.5311.4 GB
12 GB card · balancedPhoto edit 1280 px, 30 steps133 s3.7611.2 GB
Soft 8 GB cap · compactText→image 1024², 20 steps33 s1.378.9 GB
Soft 8 GB cap · compactText→image 1.6 MP, 30 steps82 s2.548.8 GB
Soft 8 GB cap · compactPhoto edit 1280 px, 30 steps141 s3.999.5 GB
Soft 8 GB cap · balancedText→image 1024², 20 steps34 s1.428.9 GB
Soft 8 GB cap · balancedText→image 1.6 MP, 30 steps84 s2.619.1 GB
Soft 8 GB cap · balancedPhoto edit 1280 px, 30 steps142 s4.039.1 GB

All 20 tests

Photography · text-to-image 3:4

Portrait: Patagonian fisherman

gpt-image-2 wins
Prompt

Close-up editorial portrait of a 70-year-old Patagonian fisherman with deep wrinkles, weathered sun-damaged skin and salt-and-pepper stubble, wearing a navy wool beanie and a yellow oilskin jacket, standing at a windy harbor in Puerto Madryn at golden hour. Shot on an 85mm lens at f/1.8, shallow depth of field, soft rim light, visible skin pores, natural color grading.

Qwen-Image-2.1 · local RTX 30607.5 min
Portrait: Patagonian fisherman, generated by Qwen-Image-2.1
gpt-image-2 · cloud1.3 min
Portrait: Patagonian fisherman, generated by gpt-image-2

Qwen renders extreme skin detail at 2K but overdoes it (HDR-looking wrinkles) and the background is a plain sea. GPT builds a believable harbor with natural skin.

Photography · text-to-image 3:2

Street: San Telmo after rain

Tie
Prompt

Street photograph of a narrow cobblestone street in San Telmo, Buenos Aires, at night after rain. The wet cobblestones glisten, a black-and-yellow taxi passes with slight motion blur, old colonial facades with wrought-iron balconies, warm sodium streetlights, a couple sharing an umbrella in the distance. 35mm film look, Kodak Portra 800 grain.

Qwen-Image-2.1 · local RTX 30607.4 min
Street: San Telmo after rain, generated by Qwen-Image-2.1
gpt-image-2 · cloud1.2 min
Street: San Telmo after rain, generated by gpt-image-2

Qwen looks like a real documentary street photo but ignores two details (couple 'in the distance', warm sodium light). GPT follows the prompt more closely but reads like a postcard.

Photography · text-to-image 4:3

Food: medialunas and cortado

gpt-image-2 wins
Prompt

Overhead food photography: a rustic wooden table with a white ceramic plate holding three glossy medialunas (Argentine croissants), a cortado coffee in a small glass, a folded newspaper, scattered crumbs, morning window light from the left casting soft shadows, 50mm lens, high detail.

Qwen-Image-2.1 · local RTX 30607.4 min
Food: medialunas and cortado, generated by Qwen-Image-2.1
gpt-image-2 · cloud1.0 min
Food: medialunas and cortado, generated by gpt-image-2

GPT knows an Argentine medialuna is small, curved and glazed, and the newspaper is legible. Qwen draws large French croissants, gibberish newsprint and an espresso instead of a cortado.

Photography · text-to-image 3:2

Hands: luthier carving a scroll

gpt-image-2 wins
Prompt

Macro photograph of the hands of an elderly luthier carving the scroll of a violin with a small gouge, thin wood shavings curling off, fingers with calluses and visible veins, a workshop full of hanging violins out of focus in the background, warm tungsten light.

Qwen-Image-2.1 · local RTX 30607.4 min
Hands: luthier carving a scroll, generated by Qwen-Image-2.1
gpt-image-2 · cloud1.2 min
Hands: luthier carving a scroll, generated by gpt-image-2

Both hands are anatomically fine. Qwen carves an already varnished violin with an odd tool; GPT shows raw maple, a real gouge and plausible shavings.

Photography · text-to-image 16:9

Landscape: Mount Fitz Roy

gpt-image-2 wins
Prompt

Landscape photograph of Mount Fitz Roy at sunrise seen from Laguna de los Tres, pink alpenglow on the granite peaks, glacier below, perfectly still turquoise lake with a mirror reflection, two hikers in red jackets small in the foreground for scale, ultra wide 16mm lens, crisp detail.

Qwen-Image-2.1 · local RTX 30607.2 min
Landscape: Mount Fitz Roy, generated by Qwen-Image-2.1
gpt-image-2 · cloud0.9 min
Landscape: Mount Fitz Roy, generated by gpt-image-2

The prompt asks for a still lake with a mirror reflection; only GPT delivers it. Qwen's Fitz Roy is recognizable but the lake barely reflects.

Photography · text-to-image 1:1

Product: watch on wet slate

gpt-image-2 wins
Prompt

Studio product photograph of a matte black mechanical wristwatch with orange accents lying on a slab of wet dark slate, water droplets on the surface, a single dramatic softbox from the top right, pure black background, crisp reflections on the sapphire crystal, commercial advertising quality.

Qwen-Image-2.1 · local RTX 30607.0 min
Product: watch on wet slate, generated by Qwen-Image-2.1
gpt-image-2 · cloud1.1 min
Product: watch on wet slate, generated by gpt-image-2

GPT gives ad-ready framing with droplets on the crystal. Qwen leaves the softbox in frame and the watch small, though the wet slate texture is excellent.

Creativity · text-to-image 2:3

Surreal: lighthouse of books

gpt-image-2 wins
Prompt

A lighthouse built from towering stacks of old books standing on a rock in an ocean of black ink, its beam made of glowing floating letters, paper birds flying around it, rendered as a detailed ligne claire comic illustration with a limited palette of cream, deep navy and coral.

Qwen-Image-2.1 · local RTX 30607.4 min
Surreal: lighthouse of books, generated by Qwen-Image-2.1
gpt-image-2 · cloud1.1 min
Surreal: lighthouse of books, generated by gpt-image-2

GPT hits every element: paper birds, ink sea with floating pages, coral palette, a beam of letters. Qwen draws real seagulls and a normal sea.

Creativity · text-to-image 2:3

Poster with Spanish text

gpt-image-2 wins
Prompt

Vintage 1950s Argentine travel poster, screen-print style. Big bold title at the top: "MAR DEL PLATA". Subtitle below it: "Veraneá en la Perla del Atlántico". Illustration of the Rambla with its two stone sea lion statues, striped beach umbrellas and a woman in a retro swimsuit. Small text at the bottom: "Ferrocarril Sud — Temporada 1954".

Qwen-Image-2.1 · local RTX 30607.4 min
Poster with Spanish text, generated by Qwen-Image-2.1
gpt-image-2 · cloud1.3 min
Poster with Spanish text, generated by gpt-image-2

Both render every word perfectly, accents included. GPT also knows the real Rambla (Hotel Provincial, stone sea lions on pedestals) and has more poster punch.

Creativity · text-to-image 1:1

Infographic: how to make mate

gpt-image-2 wins
Prompt

A clean flat-design infographic titled "Cómo preparar mate" showing exactly 6 numbered steps in a 2x3 grid, each with a simple icon and a short caption: 1 "Llenar 3/4 con yerba", 2 "Inclinar el mate", 3 "Agua tibia en el hueco", 4 "Colocar la bombilla", 5 "Agua a 75°C", 6 "Cebar y compartir". Warm green and cream palette.

Qwen-Image-2.1 · local RTX 30607.1 min
Infographic: how to make mate, generated by Qwen-Image-2.1
gpt-image-2 · cloud1.4 min
Infographic: how to make mate, generated by gpt-image-2

GPT: perfect small text and correct mate icons. Qwen misspells small captions ('Inclinar enj mate') and draws cups and a wine glass instead of a mate gourd.

Creativity · text-to-image 16:9

Character: capybara astronaut

gpt-image-2 wins
Prompt

A 3D animated movie still of a chubby capybara astronaut floating inside a cozy spaceship cabin, holding a mate gourd, small potted plants drifting in zero gravity, round portholes showing planet Earth, soft volumetric light, expressive eyes, feature-animation stylization.

Qwen-Image-2.1 · local RTX 30607.2 min
Character: capybara astronaut, generated by Qwen-Image-2.1
gpt-image-2 · cloud1.1 min
Character: capybara astronaut, generated by gpt-image-2

Both are feature-animation quality. Qwen reads 'mate gourd' literally and hands the capybara a raw calabash; GPT has it sipping mate through a bombilla.

Creativity · text-to-image 1:1

Counting: 3 apples, 2 pears, 1 banana

gpt-image-2 wins
Prompt

Top-down photo of exactly three red apples, two green pears and one banana arranged on a blue ceramic plate, with a silver fork to the left of the plate and a folded yellow napkin to the right, on a white marble table.

Qwen-Image-2.1 · local RTX 30607.0 min
Counting: 3 apples, 2 pears, 1 banana, generated by Qwen-Image-2.1
gpt-image-2 · cloud1.0 min
Counting: 3 apples, 2 pears, 1 banana, generated by gpt-image-2

GPT gets every count and position right. Qwen draws 2 apples and puts the banana off the plate.

Creativity · text-to-image 1:1

Transparent sticker (RGBA)

gpt-image-2 wins
Prompt

This is an RGBA format image with transparency. A cute cartoon sticker of a mate gourd with a smiling face and a metal bombilla, thick white die-cut outline. The image has an alpha channel and a transparent background.

Qwen-Image-2.1 · local RTX 30607.0 min
Transparent sticker (RGBA), generated by Qwen-Image-2.1
gpt-image-2 · cloud1.1 min
Transparent sticker (RGBA), generated by gpt-image-2

Both return real alpha. Qwen again draws an East Asian hyōtan gourd instead of a mate.

Photo editing · edit

Edit: sunny street → snowy night

Tie
Prompt

Change the scene to a snowy winter night: snow on the balconies, awnings, trees and street, snowflakes falling, warm lights glowing in the windows. Keep the buildings, signs, people and composition exactly the same.

Input photoMbaro01 · CC BY-SA 4.0
Input photo for Edit: sunny street → snowy night by Mbaro01
Qwen-Image-2.1 · local RTX 30605.1 min
Edit: sunny street → snowy night, generated by Qwen-Image-2.1
gpt-image-2 · cloud1.5 min
Edit: sunny street → snowy night, generated by gpt-image-2

Both keep the MARTINEZ sign, people, flower stand and architecture. Qwen is slightly more faithful to the source; GPT adds more of the requested warm light.

Photo editing · edit

Edit: change clothes, add glasses

Tie
Prompt

Change her denim shirt to a red wool turtleneck sweater and add round tortoiseshell glasses. Keep her face, identity, hair, expression, pose, phone and the background exactly the same.

Input photoAnthony Ginsbrook aginsbrook · CC0
Input photo for Edit: change clothes, add glasses by Anthony Ginsbrook aginsbrook
Qwen-Image-2.1 · local RTX 30604.9 min
Edit: change clothes, add glasses, generated by Qwen-Image-2.1
gpt-image-2 · cloud1.1 min
Edit: change clothes, add glasses, generated by gpt-image-2

Identity, hair, expression, phone and background preserved in both. Professional in both.

Photo editing · edit

Edit: rewrite a shop sign

Tie
Prompt

Edit the text on the round sign: replace "SOLAR ROAST" with "LA PORTEÑA" and replace "COFFEE" with "CAFÉ", using the same gold serif lettering and the same sign design. Keep the rest of the photo unchanged.

Input photoSolarroast · CC BY-SA 4.0
Input photo for Edit: rewrite a shop sign by Solarroast
Qwen-Image-2.1 · local RTX 30605.0 min
Edit: rewrite a shop sign, generated by Qwen-Image-2.1
gpt-image-2 · cloud1.3 min
Edit: rewrite a shop sign, generated by gpt-image-2

Both replace SOLAR ROAST COFFEE with LA PORTEÑA CAFÉ, Ñ and accent included, in the same gold serif, leaving the logo and the rest of the photo untouched.

Photo editing · edit

Edit: remove all people

Tie
Prompt

Remove all the people from the sidewalk tables and from inside the café, leaving the chairs and tables empty. Keep everything else identical.

Input photoHenrique Félix henriquefelix · CC0
Input photo for Edit: remove all people by Henrique Félix henriquefelix
Qwen-Image-2.1 · local RTX 30604.9 min
Edit: remove all people, generated by Qwen-Image-2.1
gpt-image-2 · cloud1.3 min
Edit: remove all people, generated by gpt-image-2

Both remove everyone without ghosts. Qwen keeps the original table layout more exactly; GPT fills the interior with more life.

Photo editing · edit

Edit: photo → watercolor

Tie
Prompt

Turn this photo into a hand-painted watercolor illustration in the style of classic 1990s Japanese animated films, soft painterly grass and warm light. Keep the composition and the dog's pose.

Input photoDavid Whelan · CC0
Input photo for Edit: photo → watercolor by David Whelan
Qwen-Image-2.1 · local RTX 30604.9 min
Edit: photo → watercolor, generated by Qwen-Image-2.1
gpt-image-2 · cloud1.2 min
Edit: photo → watercolor, generated by gpt-image-2

Both keep composition and pose. Qwen leans toward animation colors, GPT toward loose watercolor texture.

Photo editing · edit

Edit: combine two photos

Tie
Prompt

Place the golden retriever from <image2> sitting on the sidewalk next to the green bicycle in <image1>, matching the lighting and perspective of <image1>, with a realistic contact shadow. Keep the bicycle and the wall unchanged.

Input photowww.Pixel.la Free Stock Photos · CC0
Input photo for Edit: combine two photos by www.Pixel.la Free Stock Photos
Input photoDavid Whelan · CC0
Input photo for Edit: combine two photos by David Whelan
Qwen-Image-2.1 · local RTX 30607.2 min
Edit: combine two photos, generated by Qwen-Image-2.1
gpt-image-2 · cloud1.2 min
Edit: combine two photos, generated by gpt-image-2

Both place the dog from photo 2 next to the bicycle with a contact shadow and matching light, bicycle unchanged.

Photo editing · edit

Edit: remove background

Qwen wins
Prompt

Remove the background, and output a PNG image

Input photowww.Pixel.la Free Stock Photos · CC0
Input photo for Edit: remove background by www.Pixel.la Free Stock Photos
Qwen-Image-2.1 · local RTX 30605.0 min
Edit: remove background, generated by Qwen-Image-2.1
gpt-image-2 · cloud1.7 min
Edit: remove background, generated by gpt-image-2

Qwen cuts out the exact bicycle at its original position and scale, which is what you need for compositing. GPT redraws it larger and changes details of the rear rack.

Photo editing · edit

Edit: remove an object

Tie
Prompt

Remove the rifle hanging on the left of the wall and fill in the plaster wall texture naturally. Keep all the picture frames exactly the same.

Input photoWilfredor · CC0
Input photo for Edit: remove an object by Wilfredor
Qwen-Image-2.1 · local RTX 30605.1 min
Edit: remove an object, generated by Qwen-Image-2.1
gpt-image-2 · cloud1.4 min
Edit: remove an object, generated by gpt-image-2

Both remove the rifle and rebuild the plaster; every frame is intact in both.

FAQ

What AI image model runs on a 12 GB GPU?

Qwen-Image-2.1 (7B, open weights, September 2026) runs fully in VRAM on a 12 GB card such as an RTX 3060 when you use the int8 weights and load the text encoder, the image model and the VAE one at a time. qwen-image-local does that for you: about 35 seconds for a 1024×1024 image and about 7 minutes at native 2K.

Can Qwen-Image-2.1 run on an 8 GB GPU?

Not verified yet. The compact profile swaps the text encoder for a 6.3 GB w4a8 version and cuts peak VRAM by about 2.5 GB at the same speed, so it should fit 10 GB cards. We could only emulate 8 GB with a soft cap, and peak usage still went above 8 GB, so treat 8 GB cards as untested.

Is Qwen-Image-2.1 as good as gpt-image-2?

For photo editing, yes in our tests: 7 ties and 1 win out of 8 edits. For text-to-image, no: gpt-image-2 won 11 of 12, mostly on prompt adherence and world knowledge.

Is it free to use commercially?

The app is MIT. The Qwen-Image-2.1 weights are under the Qwen Research License, which does not allow commercial use without a separate license from Alibaba.