AI Ads That Don't Look AI-Generated: The Evidence and the Recipe
AI-generated ads match human-made ads on click-through once campaign controls go in, and the thing that actually costs you clicks is looking AI-generated, whichever way the ad was made.
A working paper covering 369 million matched ad impressions puts AI-generated and human-made ads at parity on click-through once campaign controls go in. Underneath sits a pattern the authors flag without modelling: ads that read as AI underperform whichever way they were made, so the cell beating the human baseline is AI creative nobody clocks as AI.
Every number below carries its source in the same block, either a link or a line saying where I read it and that the tally is my own. The de-slop pass, the prompt grammar and the QA checklist come out of pipelines I have shipped and never A/B tested against click-through.
Do AI-generated ads work?
In matched field data, yes, at parity. Exner, Hartmann, Netzer and Zhang covered 16.4 billion impressions and 116 million clicks across 305,121 display and native ads, then narrowed to 4,633 sibling ads matched on advertiser, objective and landing page. AI images returned 0.76% raw click-through against 0.65% across 369,533,326 impressions, and campaign fixed effects take the gap out of significance.
These are static images on display and native inventory, where the creative baseline is weak, so the cues should transfer to vertical video while the magnitudes should not. The largest field RCTs on generative AI in retail put sales effects at 0% to 16.3%, concentrated where the baseline was poor.
The four cells that matter
Human raters scored 1,751 of the images blind for perceived artificiality, and people are bad at it: 24.87% of genuinely human images landed in the likely-or-definitely-AI bucket, and 58.92% of genuinely AI images read as not-sure or human.
| The image was | It read as | Mean CTR |
|---|---|---|
| AI-generated | Human-made, or unclear | 0.79% |
| Human-made | Human-made, or unclear | 0.67% |
| AI-generated | AI-generated | 0.62% |
| Human-made | AI-generated | 0.55% |
Unadjusted cell means from the Exner, Hartmann, Netzer and Zhang paper, with no cell sizes or confidence intervals, perceived artificiality never randomly assigned, and no peer review as of September 2026. The bottom row should bother you, because nothing generative touched it and the ad still comes last.
Why HBR and Ipsos found the opposite
In May 2026, Ipsos and two Syracuse Newhouse researchers tested 20 ads across 10 brands with 3,000 US consumers and found human-made ads over-indexing the benchmark by 11 points while AI-made ads under-indexed by five, which Harvard Business Review wrote up as AI ads performing worse even when customers cannot tell them apart.
They measure different things. Peruta told phys.org the ads "started from the same brief. The only thing that changed was who made them," setting produced campaigns from ten brands against reconstructions built for the study, so the variable is craft tier wearing an AI label, scored as predicted effect on a panel rather than clicks in a live auction. In the same write-up, only about 13% of Ipsos viewers were confident an AI ad was AI, and as many suspected the human ads.
The numbers everyone repeats, and where they came from
Search this topic and the cluster appears: 12% to 23% higher CTR on Meta, 21% lower CPA, 3.4x ROAS against 3.1x, 89% cheaper per asset, 14.3 variants against 3.7, on dozens of pages with no attribution. All of it traces back to one content-marketing page by Lapis, an AI ad generation company, whose five cited sources are Taboola at roughly 500 million impressions, Soku.ai at 2,500+ campaigns, Social Operator at 1,200+, Genesis/Synthesia at 800+ and DigitalApplied at 1,500+, so four vendors and an SEO publisher.

Taboola is the one academically grounded name there, and the paper written on its data is the Exner work above, while what Lapis reports as the Taboola study is a 500-million-impression headline test about text rather than imagery. Lapis says its figures come from "real-world campaigns with real budgets" rather than a controlled A/B test, a caveat repeated nowhere downstream.
Meta's own lifts are single digits
The platform with the most to gain publishes three figures on its own Advantage+ creative page, checked 10 September 2026, with no methodology, sample or period attached:
- 2% to 3% lift in conversions from background generation on catalog ads, per Meta
- 2% lift in conversions on Facebook Reels from video expansion, same page
- 13% more conversions using related media, same page
The often-quoted 22% ROAS number comes from Meta's engineering blog, where the full sentence changes it: "when advertisers who did not previously use Advantage+ creative turned on its AI-driven targeting features, they experienced a 22% increase in ROAS from our ads." A new-adopter figure about targeting, then, and the same post only estimates 7% for image generation.
The vendor figures in this article
Ask of any ad statistic who published it, whether they sell the thing the number flatters, and where the citation chain ends, since a vendor sitting at the end of one is a rumour with footnotes. Run that test on my own tables and it draws blood. Motion sells creative analytics whose pitch is that testing volume matters, and the volume argument below is sourced entirely to it, with no way to audit the data from outside. invideo and Poppify sell generation credits and supply the yield figures I use, and Icon sells the human UGC its own price is compared against.
How to make AI ads that don't look AI-generated: the de-slop pass
AI slop is creative carrying the cues consumers read as machine-made, worth 0.12 to 0.17 percentage points of click-through in the matched data, unadjusted. Every "make your AI ads look real" guide recommends shallow depth of field, an 85mm look and a cinematic grade, worth holding against the three features that moved perceived artificiality:
- Intense colour saturation signals AI generation to consumers, the one cue with a clean production fix.
- Images displaying text read as more likely AI-generated, which the authors put down to diffusion models rendering text badly, so keep copy out of the frame at generation time.
- High aesthetic scores and medium-to-large faces predict lower perceived artificiality, and the AI images in the dataset had both.
The recipe, as preserve, add, avoid
I run this image-to-image pass on every storyboard frame before it becomes video, changing micro-realism only. Preserve, explicitly:
- each face's shape, width and proportions, one to one
- framing, composition and camera distance
- pose, gesture, product identity and label placement
Image models love to slim faces on a second pass, and a re-sculpted face is its own tell. Add:
- pore-level skin with fine vellus hair
- real material detail in fabric, plastic, metal and paper
- even daytime light with gentle highlight roll-off
- faint true sensor noise in shadow
- deep focus, with the background sharp
- the look of a flat, unedited phone photo
Avoid:
- waxy or poreless skin and beauty-filter smoothing
- over-saturation, HDR glow, bloom and halos
- oversharpening and the teal-orange grade
- shallow depth of field, bokeh, the cinematic DSLR look generally
- any text or watermark in the frame


What the evidence does not say, and what I do anyway
Only desaturating and keeping text out of frame come from the measurement, and part of the rest contradicts the paper, since blur in its feature table is small and not significant while higher aesthetic quality predicts lower perceived artificiality. I strip the grade anyway because a shallow-focus look pushes the face smaller in frame, and face size is the strongest human-reading cue the study measures. Nothing links this pass to click-through in any A/B test I can point at. Captions go on in post, burned from a word-level transcript, as the Seedance 2.5 guide sets out.
What the de-slop pass cannot do
The pass targets human perception and does nothing about machine detection, which is where platform labels come from. Meta labels ad images created or significantly edited with AI tools, TikTok reads C2PA Content Credentials and says it can label AI content whatever the creator declared, YouTube auto-applies its label to C2PA metadata, SynthID is built to survive cropping, filters, frame-rate changes and compression, which describes a realism pass, and ByteDance's Dreamina and Seedance surfaces ship watermarks, C2PA credentials and AI labels.
If you are planning around a label, budget the 31.5% click-through cost an NYU Stern finding puts on one, via PPC Land reporting an IAB framework. YouTube says disclosure "won't limit a video's audience or impact its eligibility to earn money," and TikTok says the same, so a label costs reader response while delivery carries on.
What is a good hook rate, and when should you kill an ad?
Hook rate is 3-second video plays over impressions, no platform publishes an official benchmark, and every figure in circulation is a practitioner aggregate. Ad Library's is the most transparent: healthy cold-traffic hook rate 25% to 35%, best-in-class above 45%, under 15% a kill signal, with 15-second hold rate healthy at 15% to 25%, kill below 10%, best-in-class 30%+. Sepia Lab's widely repeated TikTok average of 30.7% traces back to eleven accounts. Hold rate has two competing definitions, 15-second plays / 3-second plays and Meta's `video_p75_watched / video_3sec_watched`, so say which you mean and set an impression floor.

Meta and TikTok hook rates measure different things, since Meta's denominator is 3-second video plays while TikTok publishes 2-second and 6-second video views, where a 6-second view also counts a full play under six seconds or any engagement in the first six. TikTok's guidance puts the hook in the first six seconds and the proposition in the first three.
"85% of Facebook video is watched without sound" comes from a 2016 Digiday article citing three publishers' self-reported numbers, one of them a range of 50% to 80%, never Meta research and predating Reels. What holds up is weaker: Kantar with TikTok finds 88% of users say sound is essential, an attitudinal self-report, while Meta's feed guidance lists captions and sound as optional but recommended.
Two traps void a result before you read it. Unequal budgets make CPA comparisons invalid, because Meta allocates spend by estimated action rate, so the variant given more budget can post a worse raw CPA purely from reaching a larger, less efficient audience. And per Dmitriev and colleagues, across thousands of Microsoft experiments, a surprising result is usually an instrumentation problem.
How many ad creatives should you test per week?
More than most accounts do. Motion's Creative Benchmarks 2026 is the largest public creative dataset, $1.29bn of realised Meta spend, 578,750 creatives, 6,015 accounts from 1 September 2025 to 1 January 2026, where a winner takes ten times the account's median creative spend and at least $500, which measures spend concentration rather than business outcomes.
| Monthly spend | Creatives per week | Average hit rate | Winners per month (arithmetic) |
|---|---|---|---|
| Micro (under $10K) | 2.8 | 4.0% | 0.5 |
| Small ($10K to $50K) | 4.1 | 6.4% | 1.1 |
| Medium ($50K to $200K) | 6.6 | 8.1% | 2.3 |
| Large ($200K to $1M) | 11.2 | 8.6% | 4.1 |
| Enterprise ($1M+) | 18.8 | 8.8% | 7.1 |
Motion's own data, with no refresh published since. The last column is my multiplication of Motion's first two.
That last column is where you should get suspicious of me, because Motion separately counts winners directly and gets 0.2 a month for the average Small account and 0.5 for its top quartile at 8.0 creatives a week, against my multiplication's 1.1. The counted number is the one to plan on.
Creative throughput matters because creative does. Westwood One, summarising NCSolutions and Nielsen, puts creative at 49% of sales lift against targeting's 11% while the marketers surveyed guessed 20% and 24%, close to an earlier 2017 Nielsen study of around 500 campaigns. Read 49% as directional, since buyers killing losers suppress observational creative effects.
Spend that volume on concepts, since five variations of an unproven idea only teach you which shade of a bad idea is least bad. Rocketship HQ is the practitioner consensus, with no method published behind it: testing kept separate from scaling, ABO, 3 to 5 variants per ad set at 2 to 3 conversions a day, 30 to 50 conversions per variant across seven days. One vendor relaying another puts TikTok creative at half its effectiveness in 72 hours, down from 120 in 2024.
How much does an AI video ad cost?
Compute for a finished 30-second ad runs roughly $10 to $100 depending on model and rerolls, and a fully loaded 60-second ad is closer to $230 once editing, a 20% revision contingency and overhead go in, per AI Video Bootcamp, a vendor-adjacent stack that prices writing and directing at zero, the flaw in every cost comparison here including mine.
Routing alone can double a project. In a survey verified against provider pages on 22 August 2026, Seedance 2.5 at 720p ran $0.2312 per output second on ModelArk and on Replicate, $0.20 on WaveSpeed Turbo and $0.4730 on fal.
| Model | $/sec | Generations per keeper | One 30s ad |
|---|---|---|---|
| MiniMax H3 Max | $0.04 | 5 | $10.00 |
| Veo 3.1 Lite | $0.08 | 5 | $20.00 |
| Kling 3.0 Turbo 1080p | $0.14 | 4 | $28.00 |
| Seedance 2.0 | $0.151 | 4 | $30.20 |
| Wan 3.0 | $0.20 | 4 | $40.00 |
| Seedance 2.5 720p | $0.2312 | 4 | $46.24 |
| Veo 3.1 Standard | $0.40 | 5 | $100.00 |
I read these list prices off each provider's own rate card on 10 September 2026, so the table is my tally rather than anyone's published dataset, and only the Seedance 2.5 line is sourced above, to the cellcog survey. The last column costs ten usable five-second shots cut to a 30-second ad, so the generations-per-keeper column does all the work. An independent cross-check lands on the same ordering at roughly half the levels.
Yield decides your bill more than price does, and it tracks shot difficulty more than model choice, which is part of the case in my head-to-head for ad work:
| Shot type | Generations per keeper | Yield |
|---|---|---|
| Simple statics and product holds | 1 to 2 | 50% to 100% |
| Mixed shots with movement | 3 to 5 | 20% to 33% |
| Hands, lip-sync, walking, multi-subject | 6 to 10 | 10% to 17% |
| Ad-grade output generally, per invideo | 10 to 40 | 2.5% to 10% |
The first three bands are Poppify's, the last invideo's, and both sell into this market.
Published runs are self-selected, so read them as existence proofs. invideo's UGC batch burned 108 images and 103 videos for 49 used clips, with rejection near 85% on harder ads, and the Kalshi NBA Finals spot is reported at 300 to 400 generations for 15 usable clips, self-reported.
Cost per winner is the only unit that decides anything
Nobody joins these datasets, so take a Small-tier account shipping about 35 finished ads a month, its top quartile, where Motion counts 0.5 winners. invideo's batch figure of about $125 an ad makes that $4,375, the $145 it reports for localisation and character swaps $5,075, and Icon's human rate of $1,000 for six ads $5,833, so roughly $8,750 to $11,700 per counted winner either way.
Icon is no neutral benchmark either, since its page reads "Icon originally launched as The AI Admaker. Today, we make 6 Human UGC for $1000 (no AI / 100% real)", backed by 1,288+ vetted creators and 200+ in-house editors. The best-funded company in AI UGC walking back to human filming is revealed preference.
Where generation actually saves money
The savings sit in variant production at a fixed concept: localise a proven winner into twenty languages, swap the product inside a working structure, test five hooks against an identical body. The marginal clip approaches zero there, and every first-hand account I trust converges on that and on imagery you cannot shoot. Four things AI video still cannot do:
- Brand-consistent character work. Silverside's Svedka Super Bowl spot took about four months to reconstruct the Fembot character and train the models, per TechCrunch, longer than a conventional shoot.
- On-screen text, since signage, prices and small labels are unreliable across every model.
- High-AOV considered purchases and emotional dramatic work, where the craft gap Ipsos measured is the one you cannot brief past.
- Anything needing a specific real person, a legal problem before a technical one.

Dollar Shave Club is the useful model, producing about 90% of its advertising in-house with AI tooling while shooting its military campaign for real, because there authenticity is the message.
The prompt grammar that survives contact with a client
This comes from shipped pipelines rather than published documentation, so treat it as one operator's method. Instead of generating eight clips and fighting drift between them, build one wide 21:9 sheet of eight vertical 9:16 panels and feed it as the reference, and the model renders one continuous clip with eight internal hard cuts, panel K becoming cut K. All eight panels render in one pass against the same references, so wardrobe, face, lighting and product hold by construction, each reference gets one job (`@Image1 controls only the product geometry. Do not copy @Image1's studio background.`), and the literal string `Hard cut to.` goes between cuts, or they collapse into morphing motion.
Order the prompt as style and mood, a one-sentence summary, cut-by-cut description with timecodes, setting, audio, then the negative suffix, and budget 12 to 20 spoken words per ten seconds, my own rule of thumb rather than anything measured, which puts a 30-second ad at 36 to 60 words. Every writer I hand that to writes triple it. Cuts only snap when adjacent beats differ on point of view, framing distance and action at once, more than two simultaneous hand roles writes a phantom third hand, and her mouth moves only during her own lines fixes the most obvious tell in AI UGC.
Frozen-frame QA before anything ships
Freeze evenly spaced frames, every product close-up, and two or three mid-word frames:
- Exactly one hero product, no clones
- At most two hands per person, including at frame edges
- Absent features stay absent, and cap, button and prop states hold across cuts
- Labels are not gibberish, not mirrored, not a different real brand
- Product scale matches the holding hand
- No doubled lip edges, no face drift, no baked text
Item four is a legal check wearing a polish check's clothes.
What a finished one looks like
An ad that clears all six carries the de-slop grade, hard cuts on story beats, and captions burned from the transcript rather than generated in frame.
Do you have to disclose that an ad is AI-generated?
Reviewed 10 September 2026, and this area moves fast enough to re-check before relying on it. No platform imposes an advertiser disclosure duty on an ordinary US commercial ad, though state law reaches this creative and the FTC's endorsement rules never needed to mention AI to apply.
| Where | What is required |
|---|---|
| Meta, ordinary commercial ad | No advertiser duty. Meta labels ad images itself when created or significantly edited with AI, excluding resizing and colour correction |
| Meta, issue, election or political ads | Advertisers "are already required to disclose if the image, video or audio are created or edited with AI" |
| Google Ads, election ads | Required, and Google auto-generates it on mobile Feeds, Shorts and in-stream. Cropping and colour correction exempt |
| YouTube | Disclose realistic altered or synthetic content: real people saying things they did not, altered events, scenes that never occurred. Beauty filters exempt |
| TikTok | Creators must label realistic AI content, and TikTok auto-labels from C2PA whatever you declared |
| New York | GBL §396-b, in force 9 June 2026: conspicuous disclosure that a "synthetic performer" appears, where the advertiser has actual knowledge. $1,000, then $5,000 |
| California | SB 942 as amended by AB 853, operative 2 August 2026 at $5,000 per violation per day. No advertiser duty, though your output arrives carrying provenance |
| EU, any audience | Article 50(4) puts deepfake disclosure on the deployer, which is you, enforced from 2 August 2026, penalties in the €15M or 3% tier |
| Anywhere, synthetic testimonial | The FTC's fake reviews rule prohibits misrepresenting that a reviewer or testimonialist exists |
The Endorsement Guides are the piece most AI UGC pipelines miss. An endorsement is any message consumers are likely to believe reflects the opinions of a party other than the advertiser, and an endorser "could be or appear to be" an individual, so the test turns on consumer perception rather than on whether the endorser exists. The FTC's endorsement FAQ carries no example addressing AI endorsers, and it set aside its Rytr order in December 2025 as unduly burdening AI innovation, leaving liability with whoever publishes the claim.
Invent a person and you owe a label; evoke one and you owe consent. Tennessee's ELVIS Act protects a voice readily identifiable and attributable to an individual whether it is the actual voice or a simulation, and John R. Cash Revocable Trust v. The Coca-Cola Company, filed November 2025 over an alleged soundalike singer, sits at the pleading stage. California's Civil Code §3344 covers knowing use of a real likeness in advertising at the greater of $750 or damages plus profits, and under Labor Code §927 a digital-replica clause is unenforceable without "a reasonably specific description of the intended uses".
Trademark moves faster than copyright, since Cameo won a preliminary injunction against OpenAI over the naming of a generative video feature within about four months, Getty won in part on trade marks in the UK over generated images carrying its watermarks, and Google's indemnity carves out trademark claims from use in trade or commerce. Music is weaker still, since Suno's terms assign whatever rights Suno holds and then state it makes no representation that any copyright vests in it, and the US Copyright Office holds that purely AI-generated material is not protected.
Do people hate AI ads?
The IAB's AI Ad Gap survey of 505 US consumers and 104 ad executives, published by the trade body for the industry it surveys, found 82% of executives believing consumers feel positive about AI ads against 45% of consumers who do, and Gen Z negative sentiment nearly doubling from 21% to 39% between the 2024 and 2026 waves. Coca-Cola's AI Christmas work drew two years of criticism and System1 still scored it 5.9 stars, the top of its own scale, and no sales decline or brand-equity loss is documented against any AI ad backlash.
The study most readers arrive already believing is NIQ's, and it needs its date attached. That EEG and eye-tracking research, 2,000+ participants with about 150 on EEG, found AI ads rated more annoying, boring and confusing with weaker memory activation, in a press release from 12 December 2024 that carries no per-format numbers and no stimuli, and it predates every current video model.
Sizing the first test
Fix the offer first, write concepts rather than variations, clear the rights before a single reference is generated, then size the test: 30 conversions across twelve variants needs roughly 360 conversions, which is $18,000 at a $50 CPA, while producing those twelve at invideo's $125 an ad runs about $1,500, so the media cost is what decides whether you can run the test at all. Under roughly $20K a month, ship 3 to 4 concepts instead, since Motion counts 0.2 winners a month for the average account under $50K, and no public dataset on AI creative carries an incrementality test, a geo-lift, a holdout or a margin, so profit stays the target while your measurement stops at clicks and spend.
Frequently asked questions
Do I have to disclose that an ad is AI-generated?
No US platform requires it for an ordinary commercial ad, but New York's GBL §396-b has required a conspicuous synthetic-performer disclosure since 9 June 2026, and disclosure is required for political and election ads on Meta and Google, realistic synthetic depictions on YouTube and TikTok, deepfakes shown to EU audiences, and any synthetic testimonial.
Does Meta penalise AI-generated ads?
No delivery or ranking penalty is documented anywhere I can find, and Meta labels AI-created or significantly AI-edited ad images itself while promoting AI creative through Advantage+. The cost of a label is reader response, not distribution, and TikTok and YouTube both say labelling does not affect reach.
Can I get fined for using AI-generated ads?
Not in the US for using AI as such. New York's synthetic-performer rule carries $1,000 and $5,000 penalties for a missing label, and EU Article 50(4) deepfake disclosure has applied since 2 August 2026 with penalties in the €15M or 3% of turnover tier. Realistic US exposure is the FTC endorsement and testimonial rules and likeness claims.
Is AI UGC better than hiring real creators?
At low volume, no. Icon sells human-filmed UGC at about $167 an ad in a six-ad package, beating most AI UGC subscriptions on price, and human filming does not dodge the aesthetic penalty either, since roughly a quarter of genuinely human images in the field data still read as AI. AI wins where volume is the constraint.
Keep reading
How to Use Seedance 2.5 for Ads Without Burning Credits
How to use Seedance 2.5 for ads: the real resolution ceiling, what each of the four modes actually bills on, and the prompt grammar that stops cuts morphing.
The Best AI Video Model for Ads in 2026: Five Boards, Four Winners
Five leaderboards name four different winners. The best AI video model for ads in 2026, judged on what a finished 30-second spot really costs to make.
Why UGC Ads Beat Polished Video Ads
Phone-shot creator ads outperform studio video on nearly every paid metric. Here's what the numbers actually say, who published them, and where the advantage stops.
How Many Ad Creatives Should You Actually Test?
Only 4-8% of Meta ads become winners, and the median advertiser ships 6-7 a week. Run the arithmetic on what that means for a brand shipping one video a month.
