New: 50k+ fresh sites added in the latest refresh  ·  Try the live explorer →

Programmatic SEO in 2026: what still works, and what Google stopped rewarding

The 2021 playbook was simple: take a dataset, cross it with a set of modifiers, generate ten thousand pages from one template, and wait. It worked because the supply of those pages was low relative to demand. That is no longer true anywhere worth being.

What actually changed

Not the technique — the threshold. Google did not decide that templated pages are bad. It got much better at asking whether a given page adds anything the rest of the index does not already contain — a judgement now written directly into its spam policy on scaled content abuse. Ten thousand pages that each recombine public data now fail that test almost by definition.

So the constraint moved from can you generate at scale to does each generated page contain something scarce. That is a much harder bar, and it is why most PSEO sites in the index today are flat or falling.

Generation cost collapsed at the same time. When producing a page required an engineer and a pipeline, volume itself was a moat. Now anyone can produce fifty thousand competent pages in a weekend, so volume is not a moat at all — it is the baseline condition of the competition.

An automated conveyor line running through a factory
When generation required an engineer, volume was a moat. Now fifty thousand competent pages take a weekend, so volume is the baseline.Photo by Homa Appliances on Unsplash

Three eras, one shifting bar

EraWhat was scarceWhat won
2018–2021The pipeline itselfAnyone who could generate at volume
2022–2024Page quality within the templateBetter templates, real internal linking
2025–2026The underlying dataSites nobody else can replicate

Each era ended when its scarce input became cheap. That is the pattern to hold in mind, because it also tells you what the next threshold will be: whatever is still expensive when generation and interpretation are both free.

Trait one: the data is proprietary or expensive

The surviving programmatic sites are almost all sitting on a dataset that is genuinely hard to get. They scraped something nobody else scraped, they aggregated something that required permission, they ran the tests themselves, or they accumulated user contributions over years.

If your dataset is a public API that anyone can query, your pages are commodities, and the sites that outrank their authority are almost never built on one. The question to ask before building is not “can I get this data” but “how many other people can get this data in an afternoon”.

The strongest version of this is data that accrues. A dataset you bought is a depreciating asset; a dataset that grows every month because users add to it, or because you keep measuring something nobody else measures, gets harder to copy the longer you run it.

Rows of labelled archive boxes on deep shelving
A dataset you bought depreciates. One that accrues every month gets harder to copy the longer you run it.Photo by Nana Smirnova on Unsplash

The constraint moved from “can you generate at scale” to “does each page contain something scarce”.

Trait two: the template does real work

There is a meaningful difference between a template that formats data and a template that interprets it. The first produces “Population of Springfield: 116,000”. The second produces a comparison against similar towns, a trend over five years, and a note about what changed.

The interpretive layer is where the scarcity lives, and it is the part most PSEO builds skip because it requires thinking about the subject rather than the pipeline. Sites that do it are effectively publishing ten thousand short analyses. Sites that do not are publishing ten thousand database rows with headings.

A practical test: could a competent person with the same raw data write a better version of your page in ten minutes? If yes, the template is formatting rather than interpreting, and the page is not defensible.

Trait three: they prune

This is the least intuitive and possibly the most important. Successful programmatic sites do not index everything they can generate. They generate broadly, watch which pages earn impressions, and remove or consolidate the ones that never do. Impressions, not third-party traffic estimates, are the signal here.

A site with 2,000 pages that all get traffic outperforms a site with 40,000 pages where 1,800 get traffic — even though the second has slightly more winners. Aggregate quality is evaluated at the site level, and dead pages are a liability you pay for on every other page.

Pruning is also the discipline most teams cannot bring themselves to apply, because deleting pages feels like destroying work. Treat the index as a portfolio that gets rebalanced quarterly and it becomes routine.

Secateurs cutting through a branch in a garden
Generate broadly, index narrowly, and rebalance quarterly. Dead pages are a liability you pay for on every other page.Photo by Priscilla Du Preez 🇨🇦 on Unsplash

The mechanical parts still matter

None of the above removes the ordinary technical obligations, and at this scale they bite harder than they do on a fifty-page site:

  • Internal linking has to be earned, not generated. A related-items block that links by genuine similarity beats one that links by adjacent database ID, and Google can tell the difference between a structure and a shuffle.
  • Crawl budget is finite. Forty thousand URLs behind slow responses means much of the site is discovered late or not at all.
  • Near-duplicates need consolidating before launch. Two modifiers that produce the same page for most rows should be one modifier.
  • Every page needs a reason to be indexed. If you cannot name the query it serves, it should be noindexed or rolled into a parent.

A build order that works

  1. Prove the data is scarce before writing any templates. This is the step that disqualifies most projects, so do it first.
  2. Hand-write ten pages. If they are not interesting when a human writes them, no template will save them.
  3. Generate a few hundred, not thousands, and index them.
  4. Wait a full quarter and read the impression data honestly.
  5. Scale only the page types that earned impressions, and prune the rest before adding more.

Step four is where most projects die, because a quarter of waiting feels intolerable next to a pipeline that could emit forty thousand URLs this afternoon. But launching everything at once destroys the only cheap experiment you get: with a few hundred pages you can still change the template, and with forty thousand indexed you are committed to whatever you shipped. The restraint is the strategy, not a delay before it.

A scale architectural model under construction
Hand-write ten pages before building any template. If a person cannot make them interesting, no pipeline will.Photo by Blackcurrant Great on Pexels

How to check any of this for yourself

You do not have to take this on faith. Filter the database to programmatic and automated sites in a niche you know, then sort by how they moved through the last few core updates — the same method as reading a core update by its winners rather than its losers. Open the winners and the losers side by side and look at a single generated page from each.

In our experience the difference is visible within about thirty seconds of reading. The winners have something on the page you could not have produced without their dataset. The losers have a heading, a table, and three paragraphs of filler around it.

So is it worth starting one?

Yes, if you have or can build a genuinely scarce dataset and you are willing to invest in the interpretive layer. No, if the plan is volume. The economics have inverted: the expensive part used to be generating pages, and now the expensive part is deserving them. Budget accordingly.

An antique brass balance scale with its weights
The economics inverted. Generating pages used to be the expensive part; deserving them is the expensive part now.Photo by Diana ✨ on Pexels

Frequently asked questions

Is programmatic SEO still effective in 2026?

Yes, but only where the underlying data is genuinely hard to obtain. Generating pages at volume is no longer a moat because anyone can do it cheaply; what survives is a dataset competitors cannot replicate plus a template that interprets it.

How many pages should a programmatic site publish?

Fewer than it can generate. Launch a few hundred, wait a full quarter, then scale only the page types that earned impressions. A site where every page gets traffic outperforms a far larger one where most pages get none.

Does Google penalise templated content?

Templates are not the problem — pages that add nothing beyond what the index already contains are. Google’s spam policies target scaled content produced without adding value, regardless of whether a human or a pipeline made it.

The takeaway Programmatic SEO still works when the data is hard to get, the template interprets rather than formats, and the index is pruned like a portfolio. Volume was never the strategy — it was just the part that used to be scarce.