AI Search · Risk Analysis

AI Content Penalties: Or Is It Thin Content in a Hat?

Every guide on this answers the easy question, which is whether Google has an AI detector. A practitioner asked the harder one: has anyone ever isolated the variable? Here is the honest answer, and the part of the risk that is genuinely new.

By Linkeddit·Updated 25 August 2026·15 min read

Key takeaways

  • Google's policy targets scaled content abuse, not authorship. Its published guidance says appropriate use of AI is not against the guidelines.
  • Nobody has produced a case where authorship was the only variable. Every documented penalty also involved thin pages, no internal links and industrial volume.
  • The genuinely new risk is that the production pattern is visible: publishing velocity, repeated templates, URL structures and sudden growth are all detectable regardless of writing quality.
  • How much you can safely publish depends on the site. Twenty thousand pages on a fresh domain is a different act from gradual growth on an established brand.
  • Blocking a page in robots.txt is not the same as noindexing it, and getting that backwards leaves the page indexed with no way to signal otherwise.

01The question, asked properly

Most articles on this topic answer a question nobody serious is asking. The better version came from a practitioner running product content, posting in r/AI_SearchOptimization:

Has anyone actually tied a ranking drop to AI-written product copy, or is it always thin content wearing a hat? The mass-generated stuff that gets buried is always also duplicate, also thin, also zero internal links. Never seen an isolated variable.
via r/AI_SearchOptimization

They finished with the exact request that this article is organised around: anyone got a case where the only thing that changed was who wrote it.

That is the right question, and the honest answer is no, not that we could find. Which does not mean the risk is imaginary. It means the risk is not where most people are looking for it, and the rest of this piece is about where it actually sits.

900+
AI posts in the documented manual-action case
~1 year
How long that experiment ran before consequences
~96%
Less traffic from AI chatbot referrals
~1%
Click-through rate even when cited

02What Google actually says, in its own words

The policy is not ambiguous, and it has not changed much since it was published. Google’s guidance on AI-generated content states that appropriate use of AI or automation is not against its guidelines, and defines the problem as content used to generate output primarily to manipulate search rankings.

Read that carefully, because the operative phrase is not about how content was written. It is about why. The policy is aimed at intent expressed through production behaviour, which is why the enforcement mechanism is called scaled content abuse rather than anything mentioning AI.

This is also why the entire genre of articles asking whether Google can detect AI writing misses the point. Whether a classifier can identify generated prose is close to irrelevant, because the policy does not require it. Volume, templates and site history are far easier to measure and far harder to fake.

03The one documented case, in detail

There is a public case with Search Console screenshots, surfaced by SEO practitioners and widely discussed. The details matter more than the headline.

A practitioner published over 900 AI-generated blog posts on an idle domain, with no internal linking and, by their own description, little regard for quality. Google returned manual actions. The warnings covered reduced visibility, lower rankings and potential removal from search results, and the site was cited for aggressive spam practices including scaled content abuse, scraping and repeated violations of spam policies, affecting all pages.

Two details are easy to miss and important. The experiment ran for roughly a year before the consequences appeared, which means the absence of a penalty at month three tells you very little. And the action affected all pages, not just the generated ones, which is the part that should worry anyone running this on a domain that also hosts a real business.

The practitioner’s own summary was blunt: if you are mass-producing low-quality AI content in pursuit of bot traffic rather than real users, this is the outcome you should expect.

04Was the variable isolated? No.

Back to the original question honestly. In that case, authorship was not the only thing that differed. The pages were also:

Variable presentWould this alone risk a penalty?
Written by AINo, per Google's published policy
Published at scale, 900+ pagesYes, this is the named policy violation
Thin, low quality by the author's own accountYes, independently of authorship
No internal linkingContributes to a machine-built pattern
On an idle domain with no historyAmplifies every other factor

So the skeptic is substantially right. Every documented penalty we could find is over-determined: remove the AI and leave the volume, thinness and lack of links, and the same outcome is plausible. Thin content wearing a hat is a fair description of most of it.

Where the skeptical position needs adjusting is the assumption that quality is therefore a sufficient defence. It is not, and that is the genuinely new part.

05What is actually detectable, and why quality is not enough

The important claim in practitioner discussion is this: high-quality writing does not automatically make scaled content safe, because the production pattern is visible independently of the prose.

The signals practitioners list are all structural rather than linguistic:

Publishing velocity. Twenty thousand pages appearing overnight against gradual growth. This requires no content analysis at all.

Repeated templates and boilerplate structure. Pages that share a skeleton and vary only in the slots.

URL patterns. Structures that are obviously generated.

Internal linking patterns. Link graphs that look machine-built rather than editorial.

Sudden site growth inconsistent with the site’s history. A domain that published four posts a month for three years and then four thousand.

Notice that none of these is an AI detector, and none can be defeated by better writing. That is why a team can produce genuinely good pages at volume and still land in the same place as the practitioner with 900 thin ones. The hat is not the problem. Wearing ten thousand identical hats is.

This also explains why so many teams report no problem right up until they have one. Each of those signals is a threshold rather than a switch, and a site can sit comfortably under all of them for months while its publishing rate climbs. The absence of a penalty is not evidence that the pattern is safe. It is evidence that you have not crossed the line yet.

06AI-assisted is not automated production

The line that actually separates safe from risky is not how much AI touched the text. It is whether a human made a decision about each page.

AI-assisted content means a person decided this page should exist, what question it answers and why anyone would need it, then used a model to research, draft or edit. The judgement about whether the page deserves to exist was made once per page, by a human.

Automated content production means a pipeline decided. A template plus a list of variables produced two thousand pages, and no human evaluated any individual one. That is the behaviour the policy names, and it is visible in the output pattern whether or not the prose is good.

The test is uncomfortable but clarifying: could you explain, for any given page, why it exists and who it is for, without referring to the template or the keyword list? If the only available answer is that the pipeline generated it, you are on the wrong side of the line no matter how well it reads.

This is also why word count and quality scores are poor proxies for risk. A three thousand word page that no human decided to commission is still automated production. A six hundred word page answering one real question that somebody specifically wanted answered is not.

07Your trust budget determines what is safe

The most useful framing to come out of this discussion is that sites have different tolerances, and the same action carries different risk depending on who takes it.

Practitioners point to site authority, branded search volume, direct traffic and existing search history as factors affecting how much content a site can safely publish. A strong brand has a larger budget of trust to spend. A fresh domain has almost none.

Site profileAdding 500 pages reads as
Established brand, strong branded search, years of historyExpansion
Mid-size site, steady publishing, some brand demandAmbitious but plausible
New domain, no branded search, no historyThe pattern the policy exists to catch
Idle domain suddenly reactivatedThe documented case above

The practical implication is that advice on this topic is nearly useless without knowing which row you are in. Publishing volume that is routine for an established publisher is the single riskiest thing a new domain can do, and most of the AI content services being sold do not make that distinction.

08Programmatic SEO is not automatically abuse

An important nuance, because a lot of legitimate businesses are structurally indistinguishable from spam operations.

Ecommerce catalogues, marketplaces, service-area pages and template-driven sites all produce many similar pages from a template. So does scaled content abuse. The difference is not the mechanism, it is whether each page answers a distinct real query with distinct real data.

Two practical cautions. First, collateral damage is a real risk for these business models, and the defence is demonstrable per-page value rather than intent. Second, and more mundane, plugins can generate thousands of URLs nobody asked for. Audit what your CMS is actually publishing before assuming your page count is what you think it is.

If you are running programmatic pages, the honest test is whether you could defend a random sample of ten of them individually. Not the template. Ten specific pages, each answering a real question with data that is not on the other nine.

09Diagnosing a traffic drop before you panic

A Search Console decline can mean four very different things, and they require opposite responses. Getting this wrong is how teams delete content that was working.

What happenedHow to tellResponse
AI Overview click lossImpressions stable or up, clicks down, positions unchangedNothing is broken. Adjust expectations and measure differently
Ranking lossPositions moved for specific queriesOrdinary competitive work
Authority tighteningBroad decline across unrelated sectionsBrand and quality investment, slowly
Manual actionIt appears in the Search Console manual actions reportCleanup and reconsideration request

The first row deserves emphasis because it is now the most common and the most misdiagnosed. If impressions hold and clicks fall, you have not been penalised. You are being summarised. That is a measurement problem rather than a quality problem, and we covered how to measure it in measuring AI search visibility.

Two index states are also worth reading as feedback rather than noise. Crawled, currently not indexed and Discovered, currently not indexed are Google indicating it has seen the pages and does not currently consider them worth including. On a site that just published at volume, a rising count in either is an early warning that arrives well before any penalty.

10Cleanup, done correctly

If you have concluded that thin machine-generated pages need to go, one technical detail causes more failed cleanups than anything else.

Remove or noindex the pages. Do not block them in robots.txt. Blocking crawl is not the same as noindexing, because Google cannot see a noindex tag on a page it is not permitted to crawl. Block first and you have removed your only mechanism for signalling that the page should be dropped, while leaving it in the index.

The correct order is noindex, allow crawling so the directive is seen, wait for the pages to drop out, and only then consider blocking crawl if bandwidth is a concern. This applies equally to the AI crawler decisions we covered in the server logs analysis.

On what to remove: be honest about which pages exist because a real person would search for them and which exist because a template could produce them. The second category is what you are cleaning up, regardless of who or what wrote the words.

11The uncomfortable traffic math behind all of this

There is a reason teams reach for scaled AI content, and it is worth naming: the traffic economics of AI search are genuinely bad, which creates pressure to make up volume elsewhere.

The figures circulating in practitioner discussion are stark. AI chatbot referrals are reported generating around 96% less traffic than search, and even when content is cited, click-through sits near 1%. Separately, an Akamai report put AI bot activity as having surged 300% in 2025, with media and publishing among the most targeted sectors.

One practitioner summarised the resulting bind precisely: you can win the prompt, get a glowing recommendation, sit front and centre in the response, and still walk away with zero measurable traffic. Visibility has stopped translating into visits.

The wrong conclusion is to publish more pages to compensate, which is exactly the behaviour the policy targets. The right conclusion is that the metric has changed: if users are not clicking, what matters is whether you made enough of an impression that they search for you by name later. Branded demand becomes the outcome, not sessions.

12Where AI content genuinely helps, and what format wins

None of the above is an argument against using AI to produce content. It is an argument against using it to produce pages at industrial volume on a domain without the standing to absorb it. Two uses are straightforwardly sound.

Finding gaps rather than filling them. The strongest application is analytical: running your prompt set across engines, tallying which sources get cited, and identifying the specific questions your category answers and you do not. That produces a short list of pages worth writing by hand, which is the opposite of scaled publishing. The method is in our citation benchmarks guide.

Drafting from real evidence you already hold. AI turning your own research, customer complaints or measured data into a draft is a production speedup on content that was going to be substantive anyway. The evidence is the asset; the drafting is the commodity.

On format, the honest position is that we have not tested video against text for AI citation ourselves. What the published source data shows is that YouTube accounts for roughly 19% of Google AI Overviews top-source share, making it the one serious video citation surface, while text remains dominant across ChatGPT and Perplexity. Log evidence also indicates that the user-session agents fetching pages for live AI conversations extract text only and pull no images, CSS or JavaScript. Reasonable reading: text carries AI visibility, video carries a specific slice of Google surfaces, and neither replaces the other.

13The defence that actually works

The most encouraging finding in this research is that the recovery path is known, and it is not a content path.

HouseFresh, a site that publicly documented its own visibility collapse, is cited by practitioners as having recovered by investing in its brand, its audience and its branded searches rather than by publishing more. Branded demand and off-site presence appear to be genuinely protective.

That is consistent with everything else in the AI search picture. Since a large share of AI citations concentrate on a handful of third-party domains you do not own, and since chatbot referrals convert to sessions at around 1%, the durable asset is being the brand somebody types into a search box, not the site with the most pages.

Practically, that means treating community and review presence as search assets rather than marketing extras, and it means slow deliberate publishing compounds where aggressive automation produces a good-looking chart followed by a cliff. It is a less satisfying strategy than a content pipeline, and it is the one with evidence behind it.

Build the off-site half

Linkeddit Compete tracks the community and review conversations that answer engines lean on when recommending software, and returns a weekly graded brief on what changed. It is the off-site signal layer that no amount of on-site publishing substitutes for.

See how Compete works

14Frequently asked questions

Frequently asked questions

Does Google penalize AI-generated content?+

Not for being AI-generated. Google’s published guidance states that appropriate use of AI or automation is not against its guidelines, and that the problem is content generated primarily to manipulate search rankings. What gets penalised is scaled content abuse, which is a production pattern rather than an authorship question. The distinction matters because it means good writing at industrial volume can still trip it.

Is there a documented case of an AI content penalty?+

Yes, and the details are instructive. A practitioner published over 900 AI-generated blog posts on an idle domain with no internal linking and little quality control, and Search Console returned manual actions citing scaled content abuse, scraping and repeated spam policy violations, affecting all pages. The experiment ran roughly a year before the consequences became visible. Note that authorship was not the only variable: the pages were also thin, unlinked and mass-produced.

What can Google actually detect about how content was produced?+

More than most publishers assume. Practitioners tracking this list publishing velocity, repeated templates, programmatically generated URL patterns, machine-built internal linking patterns, and sudden site growth inconsistent with a site’s history. None of those detect AI directly. All of them detect industrial production, which is what the policy actually targets.

Is programmatic SEO the same as scaled content abuse?+

No, but they can look identical from the outside. Genuinely useful pages built from real data are legitimate, and ecommerce catalogues, marketplaces and service-area pages are structurally similar to spam operations. That structural similarity is the risk: the defence is that each page answers a distinct real query with distinct real data, not that you had good intentions.

My traffic dropped. How do I know if it was a penalty?+

Diagnose before reacting, because four very different things look similar in Search Console. AI Overview click loss means you still rank and people do not click. Ranking loss is ordinary competitive movement. Authority tightening means Google trusts the site less broadly. An actual manual action appears in the Search Console manual actions report. Only the last one is a penalty, and the four need completely different responses.

Should I delete AI content that is not performing?+

Remove or noindex it rather than blocking it in robots.txt, and the distinction is technical rather than pedantic. Google cannot see a noindex tag on a page it is not permitted to crawl, so blocking in robots.txt leaves the page in the index while removing your ability to signal that it should not be. Noindex first, and only consider blocking crawl afterwards.