A practitioner's breakdown of using GPT Image 2.5 across 100+ images to produce thumbnails, ads, brand decks, and product photos at scale.
Adapted from @TheMattBerman# How to Run a Whole Campaign From One Photo With GPT Image 2.5
we ship hundreds of pieces of content a week.
almost 90% of it needs approval before it goes out, and it's never one kind of thing. thumbnails, ads, brand decks, product photos.
gpt image 2.5 JUST dropped. i spent the week running it on workflows that actually ship and make money.
before we get into the breakdown, take the prompts. i packaged all 10 production prompts, the negative prompt rules, and the pre-ship QA checklist into a free swipe file:
👉 go.bigplayers.co/make-everything/ (https://go.bigplayers.co/make-everything/)
i ran three models (flare, sunburst, and the older gpt-image 2) across more than 100 images with identical prompts. i wanted to see what works, what fails, and what you can actually use in your business today.
here's the tldr: realism is done. every model in this test made photos that absolutely pass as human created.
i've hired dozens and dozens of graphic designers over the years. gpt image 2.5 is beating 99% of them.
that slide above came out of a single prompt.
it handled hierarchy, typography, layout, and client decision making in one shot.
the model operated as an art director.
to run a real campaign with this tech, use the anchor method:
1. freeze one real photo. never let the model guess what your product looks like. every asset descends from that anchor.
1. never leave a blank in your prompt. if you ask for a logo and don't name the brand, the ai guesses. and sometimes it stamps a competitor's trademark on your ad.
1. route by the job. flare is your art director for decks and fast ads. sunburst is your researcher for maps and factual graphics.
here's the proof, the numbers, and what you can build with it today.
## 1. start with the formats you already run
i started with five static formats we run every week:
- the us vs them split screen
- the radial stat callout
- the ugc mirror selfie
- the native iphone notes app screenshot
- the kitchen whiteboard
flare is the clear winner on the whiteboard. and honestly, that surprised me.
it understood the window joke. those are actual app windows with title bars and controls, not just a pile of rectangles. the photo feels the most natural, and the product looks great.
don't always pick the strongest model. pick l the one that gets the job done.
sunburst wins this one for me. the headline hits first, the box looks real, and the numbers have room to breathe. it feels like a finished ad. flare's close, but sunburst put the whole thing together better.
this one's close. i'd give flare the edge. the tired face, messy sink, and morning light feel like somebody actually took these photos four weeks apart. sunburst looks great too, just a little cleaner.
all three understood the before and after. but flare made me believe it a little more.
if you run paid traffic, this reprices your creative testing. you can test 20 net-new visual angles in an hour instead of waiting two weeks for a designer.
prompt rule for weekly formats: give the model the exact layout role. don't say "make a cool ad." say: "A square split-screen comparison ad. Left side: our clean solution. Right side: the messy alternative. Exactly four lines of copy per side, kept strictly to their half."
## 2. thumbnails are solved
i took this public domain pic of Albert Einstein and turned him into the hottest YouTuber of 2026.
flare nailed the MrBeast version. It used the three element rule (the pointing, giant laptop & red arrow). most importantly, you know exactly what you're supposed to look at. it looks like an upload with 2 million views.
for the tech reviewer style sunburst wins (but flare is close). bigger face, cleaner headline, and the robot coming out of the laptop gives it some life. gpt image 2 made Einstein tiny and buried the title inside the screen.
then the dramatic version. flare and sunburst both look pretty damn good. i'd still pick sunburst. the expression, red lighting, and deleted files callout make you want to know what happened.
i like the OpenAI logo here. the model added it unprompted but it works to pull in the viewer with a prominent brand they recognize before reading anything.
## 3. three brand identities and the deck to pitch them
this was the centerpiece of the entire project. and fun to make.
my wife owns an art gallery in New Orleans. she's throwing a party with an Absinthe brand and i wanted to throw some themes, visual styles, and names at them.
gpt image 2 come up with the initial concepts.
these are pretty dang good already. posters, shirts, menus, tickets, even what the room could look like. enough to see whether an idea has legs.
then i gave those concepts to gpt image 2.5 and asked it to turn them into a deck we could present.
look at how well it did. this is just flare.
better visuals. clean copy. and it actually understood how to present the ideas.
the big image gets your attention. the smaller pieces show you how the theme works across the event. the names and descriptions sit where you'd expect them. nothing feels crammed in just because it had to fit.
then it built a decision slide to help us choose.
that's what impressed me. it understood what i was trying to do with the deck. i wanted to give them something they could react to and pick from. flare got it.
to test this on a brand i don't own, i ran the same workflow on Duolingo from a single screenshot with the green hex code on it.
the same reference went into six separate generations (API so no shared chat history). you can still recognize Duo in every one.
but look at sunburst's bench version in the bottom row. i asked to see him from behind, and it put face like details on his back. gpt image 2 and flare handled that angle better.
## 4. one real photo is enough for a whole campaign
i pulled one product image off the Roots of Fight store: a Brooklyn Dodgers Jackie Robinson hoodie.
that photo is the anchor.
never let the model guess your fabric, wash, or product details. It’ll hallucinate. for brand work, every single ad comes from that one frozen image.
look at what came out of that one photo:
the hoodie reproduced perfectly across all models:
- the royal blue washed fleece
- the white script Brooklyn "B" with the red "42" on the left chest
- the sleeve patch on the correct right arm
- the flat white drawstring. the kangaroo pocket
- the hem tag. nailed it.
each headline nailed the text verbatim.
here's the exact prompt i used for the bleacher ad. copy and paste this structure:
(look closely at the bottom right corner of that first panel on the bleachers. see that logo? hold that thought)
## 5. product shots without a product shoot
sticking with that exact same hoodie, i wanted to see if we could produce ecommerce catalog shots (PDP) without booking a studio or hiring talent.
a typical catalog shoot starts at $3,500 to $5,000 (by the time you pay the photographer, the studio, the model, and retouching). often more.
the front shot is an easy win. all three models put the real hoodie on a real looking model in a clean studio setting with even commercial lighting. you could put any of them on Shopify right now.
then i asked for the back of the hoodie.
none of the models were ever shown what the back of this hoodie looks like.
this test revealed how different models handle missing information:
- gpt image 2 and sunburst left the back completely blank.
- flare took the front graphic (the script B and red 42) and duplicated it onto the back.
now look at the real product in the fourth panel. the real hoodie has an arched "BROOKLYN DODGERS" print, a massive red "42", "Robinson" in script, and Roots of Fight/MLB logos.
an empty back misrepresents a product that has graphics. but a fake design is even more dangerous. it’s rendered so cleanly that a reviewer could miss it, ship it, and not even realize. So check product shots carefully or better yet build a gate that forces the model to compare against the real SKU before shipping.
the rule: you cannot generate the back of a physical product from a photo of the front. give the model both sides.
## 6. ugc for a brand that never filmed anything
next, i wanted UGC for paid social.
using Jackie Robinson hoodie I prompted for a casual bathroom mirror selfie shot on a phone, with unflattering overhead lighting and a real iphone look.
all three passed the scroll test. on TikTok or Instagram, you'd scroll right past without thinking "AI."
but check the reflection.
none of the three models reversed the script "B" or the number "42" in the mirror. the text reads normal. in a real mirror reflection that text would read backwards. none of the image models understood mirror physics. i saw this same failure on a completely different products across my tests .
gpt image 2 also botched the caption: "hooodie" with three o's. flare and sunburst both rendered the text exactly as written.
## 7. charts, maps and explainers
most brands also need technical assets: diagrams, feature cards, infographics product explainers.
1. The B2B SaaS Explainer (Three Way Tie)
i asked for a dark mode neobrutalist infographic for StealAds explaining our three step engine.
all three models nailed it. every headline, every monospace label, connecting arrows, and the exact website URL came out perfectly. they even nailed the attached logo mark without distorting the font. any of the three could run.
when you give the models clear branding and a simple workflow, model choice becomes a matter of taste. visually, i prefer sunburst here.
2. The DC Metro Map (Sunburst Crushed)
then i asked for an accurate transit map of the Washington DC Metro system, showing all six lines and major transfer hubs.
sunburst crushed this test. it listed the exact real world terminal names for all six lines.
flare made up gibberish station names ("Moely Roym", "Pentagon Aest"). gpt image 2 drew a pretty map, but with far less geographic detail.
## 8. the revision round
every agency owner and founder knows this pain: you get a great concept, and then the client asks for two changes.
i tested what happens when you feed a finished render back into the model for two consecutive rounds of revisions.
Revision 1: "Change ONLY the headline to: 'Your competitors' ads are winning. Now you know why.' Keep everything else identical."
prompt rule for weekly formats: tton colour from green to magenta #FF2E93 and change the button text to 'See a live teardown'. Keep everything else identical."
here's how the models handled the feedback:
- flare was the clear winner. it applied both changes, kept the exact subhead, preserved the magenta hairline accent, and held the layout locked across all three generations.
- sunburst changed the words and color correctly, but the layout began to drift. the headline scaled up, the margins moved, and the logo shifted.
- gpt-image 2 failed in the worst possible way. it made the changes, but broke the company name into two words: "Steal Ads".
that's the error that slips into production. by round four of edits, nobody re-reads the company name.
## 9. the mistakes that'll save you months
look back at the bleacher ad for the Jackie Robinson hoodie in section 4.
the prompt said: "Small brand wordmark bottom right."
i didn't name the brand. i left a blank.
here's what happened:
- gpt image 2 stamped "RUSSELL ATHLETIC / 19 U.S.A. 02" with a lion crest. it slapped an actual competing sportswear company on a Roots of Fight ad.
- flare invented a fictional brand called "Rugged".
- sunburst read the tiny woven neck and hem tags inside the source photo, figured out the brand was Roots of Fight, and generated "ROF".
sunburst showed incredible reasoning there. but i'd rather give it the actual logo than make it guess which company i'm selling.
the rule:an ai model never leaves a space empty. if you ask for a badge, a logo, or a website without naming it, the model fills it. and sometimes it fills it with your competitor's trademark.
name every element in your brief.
## 10. flare vs sunburst
OpenAI released two versions of 2.5: flare and sunburst.
most people treat them like minor speed variations. they're completely different models with opposite failure modes.
→ Factual diagrams, maps, transit lines: sunburst
Only model that got all six DC Metro terminus stations right. flare garbled station names.
→ Executive pitch decks, presentation slides: flare/sunburst
Built the 5 slide brand identity deck with locked margins and zero coordinate drift.
→ Removing objects from a photo: flare
Kept the laptop lighting when a subject was removed; other models left an empty glow.
→ Multi-round revisions: flare
Held typography, buttons, and margins across two chained revision rounds.
→ Posing characters or mascots from behind: g2 or flare
Handled Duo from behind cleanly. sunburst drew his face on the back of his head.
→ Smart brand inference from photos: sunburst
Deduced the "ROF" mark from tiny woven tags on the hoodie.
→ Mirror selfies and reflections: none of them
Every model failed to reverse chest lettering in mirror shots.
speed is also a massive factor. flare clocked in at 15 to 19 seconds per generation. sunburst took 26 to 32 seconds.
flare is roughly 1.7x faster. for 80% of routine ad production, YouTube thumbnails, and slide design, flare is your workhorse. save sunburst for technical maps, diagrams, and complex reasoning.
## 11. get the prompts
here's the checklist (go.bigplayers.co/make-everything/) we run before any generated asset ships to a client or an ad account: (https://go.bigplayers.co/make-everything/)
1. the anchor check: did every product angle originate from a frozen real-world photo?
1. the blank check: did we explicitly name every logo, badge, and URL in the prompt?
1. the physical check: did we verify reflections, lighting sources, and the unseen sides of products?
1. the revision check: did we re read the static copy (like the brand name) after every feedback round?
the folks winning with this arent arguing benchmarks or if ai content is good enough (it is).
they pick one anchor photo. they close every blank in the prompt. build a workflow and eval system to scale and they run a production engine that ships 50 clean creative assets every week.
go big, Matt
P.S. Which format are you testing first?