Research Synthesis

Programmatic SEO Architecture: When Data Is the Product

Programmatic SEO generates pages from structured data at scale. Most of those pages get zero traffic. The template does not separate the pages that rank from the pages that Google suppresses. The data under the template separates them, together with the systems that Google uses to tell them apart.

Compiled by Aviel Fahl · Last updated September 2, 2026

Key Findings

96.55% of all indexed pages receive zero organic traffic from Google. Programmatic SEO operates against this base rate. Every implementation that works at scale shares one pattern. Wise draws 60M+ monthly visits and Zillow draws 33M+, and in both the data is the product, not a template wrapped around public information. Zyppy's 2026 study of 400+ sites found proprietary assets on 92.9% of the winning sites and on 57.1% of the losing sites.

Google's unhelpfulness classifiers target the opposite case. Copia and Firefly detect content velocity without quality. The site-level Panda classifier demotes a whole domain when the ratio of thin pages to quality pages is too high. The N-gram quality prediction system fingerprints phrase distributions and catches formulaic template output.

Against that, programmatic pages that cover long-tail query variations with unique data answer the sub-queries of an AI fan-out. Surfer SEO analyzed 36M AI Overviews. Fan-out coverage showed a 0.77 Spearman correlation with AI citation likelihood, the strongest single predictor in that dataset. Google's 2026 generative AI guidance narrows that advantage. Google policy treats a page per query variation with no distinct data as scaled content abuse.

Contents

96.55%

of all indexed pages get zero traffic

60M+

monthly visits, Wise (data-as-product)

0.77

fan-out ↔ AI citation correlation, pages with distinct data (Surfer, 36M AIOs)

22%

HCU recovery rate

Zero traffic is the default outcome for a new page


Ahrefs analyzed approximately 14 billion pages and found that 96.55% of all indexed pages receive zero organic traffic from Google. Zero traffic is the default outcome for any page published to the web. Programmatic SEO operates against that base rate. Each generated page either beats the base rate with unique data, or joins it with commodity content.

The timeline data makes the challenge sharper. An updated Ahrefs study of 1.3M keywords found that only 1.74% of newly published pages reach the top 10 within one year. The 2017 version of the same study put that figure at 5.7%. The average page at position 1 is five years old, and 72.9% of the pages in the top 10 are more than three years old.

Of the pages that do reach the top 10, 40.82% get there in the first month. Early momentum matters.

Publish 10,000 programmatic pages of commodity content and approximately 9,655 of them get zero traffic. Volume amplifies signal, or volume amplifies noise. Volume creates neither. The clinical diagnostic framework calls this a ceiling problem, not a weight problem: the constraint is eligibility, not competition. No amount of on-page optimization changes the outcome when the underlying data does not differentiate. That ceiling is why the investment screen runs before the first programmatic page ships. Page category allocation comes first: is a programmatic page the right page type for this business model? Query-level expected value comes second: do the specific queries justify the investment?

Granularity, not volume, is the competitive lever


Beating that base rate takes data at a granularity that the competitors do not serve. Owning data is not the advantage. G2 generates 92% of its traffic from programmatic pages built on user-generated review data. More granular data produces more specific pages, those pages target more specific queries, and those queries carry less competition. Granularity is the mechanism that lets a low-authority challenger compete against an incumbent with no editorial team and no link campaign.

The pattern is consistent across every successful implementation at scale. The aggregator archetype, product-led, inventory-driven, SEO as the primary growth channel, is the natural home for programmatic approaches. Product-Led SEO, as Eli Schwartz defines it, treats SEO as a product experience, not as a traffic channel. He states the rule directly: “Build your product in the way that Google's algorithms optimize for.” The rendered data is the product, optimized for search.

Source:Multiple sources: Daydream, Growth Memo, Foundation Inc, upGrowth, Practical Programmatic. Traffic figures are third-party estimates (Ahrefs or Similarweb) unless a filing is named. Tools disagree by 30% or more on the same site.
CompanyData AssetScaleOutcome
Wise (TransferWise)Real-time exchange rates, corridor-specific fees537K+ pages60M+ monthly visits
NerdWalletProprietary calculators, real-time rate data, editorial overlayFinancial product comparisonsS-1: 70%+ unpaid traffic
Zapier5,000+ app integration database, pairwise integration pages50K+ integration pages5.8M+ monthly visits (third-party estimate)
ZillowZestimates, price trends, school data, walkability scores5.2M+ indexed pages33M+ monthly visits
G2User-generated reviews, feature comparisons, pricing data140K+ product pages92% traffic from pSEO
PayscaleCrowdsourced salary data with statistical distributions212K+ pages530K–2.9M monthly visits
CanvaProprietary template library + usage analytics2.2M+ template listing pages1.3M+ monthly from pSEO

Eli Schwartz draws a critical distinction: “Programmatic SEO is not product-led SEO.” Product-led SEO creates pages from product data that users need. Google's detection systems catch the other kind: a template that reformats public data and adds no value. The line runs between data-as-product, such as Wise's exchange rates and Zillow's Zestimates, and data-as-decoration, which is a public dataset in a better template.

Outcome data confirms the lever. Cyrus Shepard at Zyppy compared more than 400 winning and losing websites (April 2026) against their 12-month Ahrefs traffic trend. Proprietary assets showed the widest gap of the five features he measured: 92.9% of the winning sites held one, against 57.1% of the losing sites. The Spearman correlation between proprietary assets and the traffic trend was 0.357. The study defines a proprietary asset as a thing other sites cannot replicate with ease: a unique product, a database, user-generated content, software, or reviews.

Letterboxd and TodayTix sat on the winning side, with user-generated film data and live ticket inventory. Lifewire and The Spruce sat on the losing side, with tutorial and lifestyle content and few first-party assets. The five features were additive. Sites with four or more features won at 68.1%, and sites with none won at 13.5%. The gap between 92.9% and 57.1% is the strongest outcome-level evidence for data-as-product from a traffic study of this size.

The same year gave the counter-case. Lily Ray at Amsive measured the March 2026 core update across 2,076 domains with the SISTRIX Visibility Index. Comparison aggregators and review platforms lost: Tripadvisor -44.8, Yelp -33.1, Indeed -18.1, Realtor.com -13.7, NerdWallet -4.3. The originators of the underlying inventory gained, and Zillow did not lose.

The two studies agree once the lever has a precise definition. A page that re-presents inventory the originator already publishes has no proprietary asset. A page that holds data the originator does not publish, a cross-entity comparison or a derived metric, has one. For a challenger, the comparison and derived-metric pages are the defensible layer. A single-entity page that mirrors the entity's own page has no defence.

Banksparency (a project I built and operate) shows the pattern at a smaller scale. A daily pipeline ingests data from 80+ financial institutions, and the site draws 10K+ monthly visits with no link building and no editorial hours. The advantage comes from bank-specific metrics at a level of detail that the general financial comparison sites do not surface. The template and the domain contribute nothing to it.

The data moat timeline

A data moat needs 2 to 3 years of consistent investment before it delivers a competitive advantage. A competitor can replicate the technical infrastructure in months. The data asset takes years. The compounding gap is the moat, not the code. Morningstar's moat taxonomy names three sources that a data asset can hold at once.

Intangible assets are the proprietary datasets. Network effects are the user-contributed data flywheel. Cost advantage is the near-zero marginal cost of each new page once the infrastructure exists.

Four systems detect a thin programmatic build


Granularity sets the upside. Google's detection systems set the downside. Four overlapping systems detect and suppress low-quality programmatic content, and each works at a different level: the site, the page, and the phrase. Their combined effect decides whether a programmatic build ranks. The 2024 API leak revealed the specific mechanisms.

Panda scores the whole site on its thin-to-quality ratio

The Panda patent (US9031929B1) computes a site-level quality score from the ratio of navigational queries that reach a site against the informational queries the site answers. Every generated page raises that score or drags it down. Take a build of 50,000 pages where 40,000 are thin. Those thin pages lower the ranking ceiling of all 50,000, including the 10,000 that carry genuine value. The API leak confirmed this as pandaDemotion in CompressedQualitySignals, a pre-computed site-wide demotion that gates pages before query-time ranking begins.

N-gram fingerprinting catches formulaic phrasing

The N-gram quality prediction patent (US9767157B2) builds a phrase model from sites of known quality, using n-gram frequency patterns. The system flags template output that carries unnatural phrase distributions: repeated boilerplate, identical sentence structures, and shallow variable substitution. Swap one city name for another, and if 85% or more of the content stays identical, the model catches it.

Copia and Firefly watch publication velocity

The 2024 API leak revealed Copia and Firefly (QualityCopiaFireflySiteSignal), the scaled-abuse detection pipeline. Copia monitors content velocity: the ratio of the URLs a site generates to the substantive articles it produces. Firefly aggregates inputs from Copia, page quality scores, and NavBoost (Google's user-interaction signal system that records click behavior over a 13-month window) to make site-wide demotion decisions. The March 2024 core update, which incorporated these signals, reduced low-quality content in search results by 45%.

Source:Google API leak (May 2024) and patents, see google-api-leak.md
SignalScoringWhat It Detects
OriginalContentScore0–512How much content is unique vs. existing corpus
contentEffortLLM-scoredEstimated effort: unique images, original data, linguistic complexity
CopycatScoreFlagDetects near-duplicate content across pages
Copia velocity ratioSite-levelURL generation rate vs. substantive content produced
N-gram fingerprintPer-pagePhrase distribution compared against known-quality sites

The HCU classifier reads the ratio, not the count

The helpful content system joined core ranking in March 2024. A machine learning classifier produces one site-wide signal, and that signal marks unhelpful content. A study of 400 affected sites measured a recovery rate near 22%. A build that generates a high ratio of low-value pages to high-value pages risks the classifier at the site level. The ratio matters more than the count. A site with 1,200 pages of which 1,000 are strong sits far safer than a site with 50,000 pages where 49,000 are thin.

The quality synthesis

Q* (pronounced “quality star”) is Google's aggregate quality metric. The metric synthesizes the site-level signals into one score. A site that scores below 0.4 cannot win a rich result. That bar excludes it from a Featured Snippet and from a People Also Ask entry. For a programmatic build, Q* is the ceiling, and page-level optimization cannot clear a site-level quality debt. The API leak states the relationship: “E-E-A-T is the goal, Q* is the system, Site_Quality is the score.”

A template must produce three kinds of uniqueness


Those four systems set the bar that the template must clear. A programmatic template that ranks on page 1 carries at least three types of uniqueness, and it holds a minimum of 30–40% content differentiation between the pages that share it. That threshold separates a template that produces pages at the top of the quality distribution from a template that feeds the 96.55%.

1. Content uniqueness. Intent-specific text that changes per page, not variable substitution in boilerplate sentences. Each page addresses the specific problem its target query represents. “USD to EUR” has different user needs than “PHP to KRW”, and the fee structures, corridors, and provider availability differ. A template that treats them identically fails the information gain test.

2. Structural uniqueness (differential templating). Different modules render on different pages based on data attributes. A real estate template for a high-competition metro area shows more data-rich modules than one for a rural area with sparse data. This prevents the empty-module problem, where template shells with missing data sections read as thin content to quality classifiers.

3. Linking uniqueness. Internal links connect semantically related pages, not random pages from the same template. Sibling links from the same cluster, contextual cross-references based on data relationships. The linking architecture should reflect topical authority signals, with pages clustered by semantic proximity, not by template type.

Four more mechanisms compound the advantage. Proprietary data layers add calculators, real-time feeds, and comparison tables. UGC overlays add reviews, ratings, and community contributions. Semantic variation changes the titles and the headings. Multiple template variants serve different data density levels, which beats one template with empty modules.

Practitioner threshold

Industry consensus puts the floor at 500 unique words per programmatic page, with 30 to 40% content differentiation between the pages that share a template. A page under 300 words risks a thin content classification. Treat these numbers as practitioner heuristics. Google confirms no threshold, but the numbers track the survival rates observed after the helpful content update.

Scheduled revalidation is the safe rendering default


A template that produces three kinds of uniqueness still has to reach the crawler. The rendering strategy decides whether Google and the AI crawlers can process programmatic pages at scale. Onely found that Google never indexes 42% of JavaScript-rendered content, and Google needs 9x more time to crawl a JavaScript page than a plain HTML page. The AI crawlers, GPTBot, ClaudeBot, and PerplexityBot, do not execute JavaScript at all. Server-side rendering is therefore a requirement for AI search visibility.

Source:Next.js docs, Vercel, Onely, 2024-2026
StrategyBest ForSEO Tradeoff
SSG (Static Site Generation)Stable data, location pages, historical comparisonsFastest crawl, best CWV, lowest cost. Rebuild required for changes.
ISR (Incremental Static Regeneration)Periodically changing data, exchange rates, pricingPre-rendered with background revalidation. Best balance for most builds.
SSR (Server-Side Rendering)Real-time data requirements on every requestFull HTML per request. Higher cost, guaranteed freshness.
CSR (Client-Side Rendering)Interactive dashboards, gated tools only67% lower rankings vs. server-rendered. Invisible to AI crawlers.

Incremental Static Regeneration (ISR) is the default for a programmatic build at scale. The build generates each page statically, then revalidates it in the background on a set interval. Volatile data such as pricing revalidates hourly. Semi-stable data such as reviews revalidates daily.

Use SSG for a small and stable dataset. Use SSR only for a genuine real-time requirement. CSR is categorically wrong for a page that needs organic visibility or AI visibility.

Fresh data earns more frequent recrawls

Google's freshness system (patent US8549014B2) creates a feedback loop between data update frequency and crawl allocation. The API leak revealed the specific mechanisms: lastSignificantUpdate tracks substantive revisions, freshByDocFp uses document fingerprinting to detect actual content changes vs. cosmetic date edits, and bylineDateConfidence scores the accuracy of displayed publication dates.

Some pages rest on data that updates on its own: an exchange rate, a price, an inventory count, a statistic. Those pages hold a structural freshness advantage that static content cannot match. The advantage compounds. A page with a record of meaningful updates earns a more frequent recrawl. The recrawl finds the new data sooner, and the earlier discovery reinforces the freshness signal. NomadList updates internet speeds, temperatures, and air quality several times each day, and those freshness signals compound with its programmatic coverage.

Three decisions define the pipeline

The generic pipeline pattern (ingest, normalize, enrich, generate, monitor) is well-documented elsewhere. What matters is the architectural decisions within each stage. Consider Banksparency, which ingests data from 80+ financial institutions daily. Three decisions drove the build:

  1. Rendering choice: ISR with 24-hour revalidation. Bank data changes daily, not hourly. SSG would require full rebuilds on every data update. SSR would add latency without benefit. ISR hits the sweet spot: pre-rendered pages with background refresh matching the actual data cadence.
  2. Normalization as differentiation. Raw institutional data uses inconsistent naming, date formats, and metric definitions across 80+ sources. The normalization layer does more than clean the data. That layer creates the cross-institution comparisons that exist nowhere else. The stage that most teams treat as plumbing is the stage that produces the unique data asset.
  3. Differential templating by data density. Not every institution provides the same data fields. Instead of one template with empty modules, Banksparency renders different module sets from the available data. An institution with richer data gets a richer page. The variant approach avoids the thin-page penalty that one universal template triggers across the lower-data institutions.

Every template variable should inject meaningfully different data. Embed data quality tests in CI/CD pipelines. The same discipline applied to code should apply to data.

The template carries the internal link graph


Rendering gets a page to the crawler. The internal link graph decides whether the crawler reaches it. On a programmatic site the template must carry that graph from the start. The hub-spoke model is the structural default. A category hub links to every page in its group. Each spoke links back to the hub with descriptive anchor text, and it carries 3 to 6 sibling links from the same semantic cluster.

The click depth ceiling is structural. No programmatic page should be more than 3 clicks from the homepage. Botify's analysis of 6.2 billion crawl requests found a 33% crawl ratio drop for sites with 1M+ pages at depth 3–4. Orphan pages, those with no internal links, consume 26% of crawl budget while generating only 5% of organic traffic. Flat architecture ensures crawlability and link equity distribution.

Cluster-aware linking connects semantically related pages, not random pages from the same template. “USD to EUR” links to “USD to GBP” and to “EUR to JPY”. The page does not link to “PHP to KRW”, which shares no data relationship with it. When a hub page earns an external backlink, the equity flows to the connected spokes. When a spoke earns a link, the equity flows to the hub and to the connected siblings. The hub-spoke model compounds, so one acquired link benefits the whole cluster.

The build must generate the schema markup from the same data source that fills the template. That keeps the visible content and the markup in step. Each template type needs its own schema definition, and content engineering principles apply directly. Programmatic generation stops the drift between content and markup that a manual process introduces at scale.

A new domain pays a waiting period before the data counts


The link graph and the template set the ceiling. The age of the domain sets the clock. The API leak confirmed what Google had publicly denied. hostAge is a PerDocData attribute, and the documentation describes its use “to sandbox fresh spam in serving time”. A new domain serves a trust-building period whatever the quality of its content. It competes against sites that hold 13 months of accumulated NavBoost click signals.

The disadvantages compound on a new domain. The hostAge sandbox holds down the ranking potential while the engagement signals accumulate. At a 1.74% first-year top-10 rate, most programmatic pages on a new domain will not rank inside the first year, at any content quality. Build on an established domain, or acquire one, and the time to value falls sharply.

A 16-month SE Ranking experiment (March 2026) confirmed the pattern empirically. The researchers published 2,000 AI-generated articles across 20 new domains. Google indexed 71% within 36 days, and 28% reached the top 100 in the first month. By month 3 the rankings had collapsed to 3% of pages in the top 100, and month 16 showed no recovery. Google indexes new-domain content quickly, and it does not sustain the rankings without accumulated trust signals. The fast indexing created an illusion of traction, and the sandbox period corrected it.

Publish in phases and gate each phase on engagement

A launch of thousands of pages at once strains the crawl budget and risks a quality flag from Copia and Firefly. Practitioners recommend a phased launch, though no controlled study tests it against a full launch:

  1. Seed phase (50–100 pages). Launch highest-quality pages first. Monitor indexing, ranking, and behavioral signals for 4–8 weeks.
  2. Validation phase (500–1,000 pages). Expand to the next tier. Compare engagement metrics against seed phase. Pause if behavioral signals degrade.
  3. Scale phase (full inventory). Deploy remaining pages in batches of 20–50 per day. Monitor site-level quality signals between batches.
  4. Maintenance phase. Ongoing data freshness, UGC accumulation, template iteration, pruning of underperforming pages.

The traffic cliff pattern

Programmatic launches repeat one pattern. The pages gain rankings within weeks, and then traffic falls sharply, often by 80 to 90%, during the next core update. The hypothesized mechanism runs in three steps. Google assigns a preliminary ranking from surface signals. Behavioral signals accumulate over 2 to 3 months: NavBoost click data, pogo-sticking, and dwell time. The next core update reads those signals and demotes the pages that carry negative engagement.

Launch small, validate the engagement, then scale in steps.


The sandbox punishes a new domain in classic search. AI search reverses part of that penalty, and it hands programmatic pages a structural advantage that editorial content rarely matches. AI search platforms use query fan-out, decomposing a single query into 8–15 sub-queries, each retrieving candidate pages independently. Surfer SEO's analysis of 36M AI Overviews found that fan-out coverage has a 0.77 Spearman correlation with AI citation likelihood, the strongest single predictor in that dataset. AirOps (March 2026, 15,000 queries, 82,108 citations) corroborates the finding: 32.9% of cited pages appeared only in fan-out SERPs, not in the original query's top 20. And 95% of fan-out queries had zero traditional search volume, invisible to conventional keyword research. A programmatic build that covers the query space of a topic provides more surfaces for AI citation than a handful of editorial pages. That advantage holds only when each page carries distinct data.

Google narrowed the advantage in May 2026. Its generative AI optimization guide (updated July 2026) names the pattern directly. A separate page for every query variation, including a fan-out query, falls under the scaled content abuse policy. The condition Google states is the purpose: the pages exist primarily to manipulate rankings or generative AI responses. Google's spam policies draw the same line for doorway abuse. Variant pages that differ only by the substituted variable are doorways, and pages that each carry different data are legitimate.

The guide also states that a high quantity of pages does not make a site more relevant. Google's systems now understand relevance without an exact match between the query and the page. Read the 0.77 correlation as a statement about coverage achieved with real data. The correlation is not a licence to expand the page count against the fan-out query space. Where the build cannot supply distinct data for a variation, do not create the page.

Domain authority is not the gating factor for AI citation. Profound (250M+ AI responses, 3B+ citations) found that traffic explains 5% of citation behavior (r²=0.05). Backlinks explain 3.8% (r²=0.038). AirOps measured citation rates by domain authority: sites in the DA 20-80 range achieve 21.5-23.6% citation rates, and DA 80-100 sites drop to 15.0%. Retrieval reaches the high-authority sites more often, and those sites convert retrieval into citation at a lower rate. Broad coverage probably dilutes their topical precision. Challenger programmatic builds with deep, specific data can compete on citation rate without needing established authority.

Information gain decides what an AI system cites

Google's Information Gain patent (US11354342B2) scores documents on a 0–1 scale based on how much novel content they contain relative to the existing result set. Content that adds nothing new scores near zero. For programmatic SEO, this creates a clear hierarchy:

Source:Google Patent US11354342B2 / US20200349181A1
Gain LevelDescriptionExamples
HighestProprietary data existing nowhere else on the webWise corridor fees, Zillow Zestimates, Banksparency bank-specific metrics
ModeratePublic data combined in novel ways: calculators, visualizations, cross-referencesPayscale salary distributions, NerdWallet comparison tables
Near-zeroPublic data reformatted into templates without analysis or unique data pointsGeneric directory pages, thin aggregation with variable substitution

AI-cited content covers 62% more facts than non-cited content and is 25.7% fresher on average (Wellows, Digital Bloom). Programmatic pages with proprietary data inherently score high on information gain because the data is absent from every other document. Information gain is the mechanism that turns data granularity into an AI citation advantage.

Content structure matters equally for AI extraction. Data-rich pages with tables achieve 2.5x citation rates vs. paragraph text, and FAQ-structured content shows 28-40% higher citation probability (Onely compiled research). Programmatic templates are naturally suited to this: structured data renders as structured HTML (tables, definition lists, comparison grids), which AI extraction systems prefer.

Key findings

Topical authority accelerates programmatic visibility by 57% (Graphite). Programmatic pages that cover the long-tail variations build topical coverage, which the API leak measures as siteFocusScore. That coverage compounds back into ranking eligibility for the harder queries. The pattern is the evidence-builder loop: win achievable long-tail queries first, use those wins to build authority for harder queries.

Three conditions must hold at the same time


Those advantages and constraints reduce to three conditions, and all three must hold at the same time. The business owns a data asset that a competitor cannot replicate cheaply. The data maps to a query space with structured search demand. The organization has the engineering capacity to build the pipeline and to maintain it. Remove one condition and the approach fails.

Programmatic SEO works when:

  • The data asset is proprietary, crowdsourced, or exclusively licensed, not publicly available in the same form
  • Query demand follows a modifier × entity pattern (city × service, product × comparison, metric × time period) that maps to programmatic pages
  • Each page delivers different information, not variable substitution in boilerplate
  • The organization treats the build as product infrastructure, not as a marketing project. SEO ROI in financial services averages 1,031% over 3 years. Real estate averages 1,389% (First Page Sage), and only a sustained investment reaches those numbers

Programmatic SEO fails when:

  • The data is public and the template adds no analytical layer, producing generic directory pages that any competitor can replicate in a weekend
  • Content differentiation between pages falls below 30%, triggering N-gram fingerprinting and Panda quality scoring
  • The pages differ only by the substituted variable. Google's spam policies name that pattern doorway abuse, and the 2026 generative AI guide names a page per fan-out variation scaled content abuse
  • Pages launch at scale without phased quality validation, and the Copia/Firefly velocity detection fires before behavioral signals can accumulate
  • The build runs on a new domain with no existing authority. The shrinking click window means the sandbox period is more punishing than ever
  • Content half-life is collapsing (3–6 months for competitive topics), and the build has no refresh mechanism. Static programmatic pages decay faster than editorial content because the data they surface goes stale

The strategic frame matters. Kevin Indig's aggregator vs. integrator distinction decides whether programmatic SEO is the right approach at all. An aggregator business, product-led and inventory-driven, is the natural home. An integrator business supports a programmatic play only when it holds a structured data asset to scale against. Even then the play stays limited in scope.

The decision framework

Four questions come before the build. Does the data asset exist, or can the team build it? Does query demand follow a modifier × entity pattern? Can engineering support the pipeline? Will the organization invest 2 to 3 years in the data moat? If any answer is no, programmatic SEO is the wrong approach, and the diagnostic framework can name the right one.

The 96.55% base rate stays. Google's detection systems keep getting sharper. AI search keeps fragmenting the click window.

The same forces make data-as-product programmatic content more valuable. The supply of pages that carry unique data at granular query coverage grows far slower than the supply of pages that do not. The gap between the programmatic SEO that works and the programmatic SEO that Google suppresses widens, and it will keep widening. Your data decides which side of that gap you build on.