Programmatic SEO means generating many pages from one template and a dataset. The canonical examples — travel listings, city pages, comparison directories — are cited endlessly, usually without anyone checking whether they still work.
A good number do not. Whole sets have been deindexed, and the pattern in which ones survived is clearer than the usual “it depends” suggests.
The distinction that decides it
Ask one question of any templated page: does this page contain information that exists nowhere else?
A page listing the actual opening hours, prices and availability for one specific thing contains data. A page that fills a sentence template with a place name contains a sentence.

Both look identical in the CMS. Only one has a reason to be in an index, and that is the line the last few years of updates have been drawing.
What survived, what did not
| Page type | Unique data per page | How it has held up |
|---|---|---|
| Listings with live availability and price | High — changes daily | Still ranking, still expanding |
| Comparison pages between two real products | High — specifications differ genuinely | Holding, where the specs are accurate |
| Statistics or reference pages from a real dataset | High | Holding, often the strongest performer |
| “Best X in {city}” with the same ten entries | Low — only the place name changes | Largely gone |
| Spun definitions and glossary variants | None | Gone, and taking sitewide trust with it |

The dividing line is not page count. Sites with hundreds of thousands of templated pages still rank; sites with four hundred have been wiped out. It is whether each page earns its place independently.
Why the failures took the whole site down
The unpleasant part of a programmatic failure is that it rarely stays local. A large set of near-identical pages appears to affect how the rest of the domain is assessed, which is why recoveries so often involve deleting rather than improving.

Backlinko’s overview of how programmatic SEO is built and where it breaks covers the implementation side in more detail. The strategic version is shorter: if you cannot describe what is unique on page 4,000, do not publish pages 1 through 4,000.
The honest test before you build
- Write out, in a sentence, what differs between page 12 and page 4,000.
- If the answer is only a noun, stop.
- Check that someone is actually searching for the long tail you are targeting — autocomplete is a free, real signal.
- Set a minimum data threshold per page, and refuse to publish below it.
- Ship a small set first. Wait a full index cycle. Expand only if impressions hold.
That fourth point is the one people skip. A row count floor — twenty-five real data rows per page, say — converts a judgement call into a rule, and rules survive deadlines in a way judgement does not.

The internal linking problem nobody mentions
A programmatic set creates thousands of pages at once, and they all need to be reachable. Left alone, most sit four or five clicks from the homepage with a single link pointing at them.
Those pages get crawled rarely and rank poorly, which is frequently misdiagnosed as a content quality problem when it is a structural one.
The sets that work build the linking deliberately: hub pages that group the long tail, related-item links between pages that genuinely relate, and a sitemap that reflects the actual hierarchy rather than a flat list of everything.
This is unglamorous and it is often the difference between a set that earns and one that merely exists.
Scaled content and the AI question
Generating the prose rather than the data does not change the analysis. If the underlying page has nothing unique, writing it more fluently does not give it something unique.
Where generation genuinely helps is in presenting real data well — turning a row of figures into a readable paragraph that a person would actually want. The data still has to exist first.

Where the data usually comes from
The practical constraint on programmatic SEO is rarely the templating. It is finding a dataset good enough to justify thousands of pages.
The sets that work tend to own their data or generate it: a marketplace with live listings, a tool recording its own results, a company publishing what its operations produce. Licensed or scraped data usually appears on several competing sites at once, which removes the uniqueness that made it worth publishing.
If the dataset behind a proposed set is one anybody could buy, expect to be one of several sites publishing the same pages — and expect the eventual re-evaluation to treat all of you accordingly.
Finding current examples rather than cited ones
The examples worth studying are the ones ranking now, which is a different set from the ones in the blog posts. Filter for sites in the programmatic pattern and check how they moved through recent updates — a set that survived three core updates is evidence; a set that was cited in 2021 is a memory.
Our update view shows before-and-after traffic for every confirmed update, which is the fastest way to tell the two apart. The fuller argument about what Google stopped rewarding is in programmatic SEO in 2026, and if you are choosing where to point this technique, finding high traffic, low competition niches is the prior step.

Approached with that discipline, programmatic SEO remains one of the few ways a small team can compete on coverage. Approached as a way to publish quickly, it remains one of the fastest ways to lose a domain’s standing.
Frequently asked questions
Does programmatic SEO still work in 2026?
Yes, where each generated page carries data that exists nowhere else. Listings with live prices, genuine product comparisons and pages built from real datasets continue to rank. Template-filled pages that vary only by a place name largely do not.
How many pages is too many?
Page count is not the constraint. Sites with hundreds of thousands of templated pages still rank while others with a few hundred have been deindexed. What matters is whether each page would justify its own existence to a reader.
Can a site recover from a programmatic SEO penalty?
Often, but usually by removing the thin set rather than improving it. Large groups of near-identical pages appear to affect how the whole domain is assessed, so deletion tends to work faster than rewriting.
Does using AI to write the pages change anything?
Not the underlying problem. If a page has no unique information, generating the prose more fluently does not give it any. Generation helps most when it presents real data readably, which requires the data to exist first.
The takeaway Ask what differs between page 12 and page 4,000. If the honest answer is a noun, the set will not survive. Set a data floor per page, publish a small batch, and expand only once it holds.


