Research Synthesis

Topical Authority: From SEO Folklore to Confirmed Signal

Google measures how deeply a domain covers a subject. The 2024 API leak named three of the attributes that do it: siteFocusScore, siteRadius, and site2vecEmbeddingEncoded. For years before the leak, practitioners argued about whether topical authority existed at all. That argument is over. What remains is an engineering question. Content depth and entity coverage move a site through topic space, and off-topic pages pull it back.

Compiled by Aviel Fahl · Last updated September 1, 2026

Key Findings

The 2024 API leak confirmed topical authority as a measured system. siteFocusScore, siteRadius, and site-level topic embeddings quantify how deeply a domain covers a subject. The effect is the combined output of topic embeddings, NsrChunks, ClusterUplift, and pairwise quality comparisons. No single score exists for anyone to toggle.

Text relevance is the strongest ranking factor at 0.47 correlation across 16,298 keywords. Fan-out query coverage correlates with AI citation likelihood at 0.77 Spearman across 36 million AI Overviews. Content depth and entity coverage shift the site vector. Volume alone does not.

On this page

0.47

Text relevance correlation (strongest factor)

0.77

Fan-out coverage ↔ AI citation (Spearman)

57%

Faster visibility for high-authority content

7

Confirmed topical authority signals in API

The API Leak Ended the Debate


In May 2024, an automated bot uploaded thousands of pages of internal Google Search API documentation to GitHub. The Content API Warehouse leak revealed 2,596 modules containing 14,014 attributes across 2,500+ pages, the most comprehensive view of Google's ranking infrastructure ever disclosed. Among those attributes: siteFocusScore, siteRadius, and site2vecEmbeddingEncoded.

iPullRank's analysis and Hobo Web's breakdown identified what these signals do. siteFocusScore quantifies how dedicated a site is to a single topic, specialist vs. generalist. siteRadius measures how much an individual page deviates from the site's central theme. And site2vecEmbeddingEncoded is a compressed vector embedding of the site's overall theme, with pageEmbeddings measuring each page against that site-level vector.

The leak is not inference. Google acknowledged the documentation as authentic, and cautioned that the data was “out-of-context, outdated, or incomplete”. The caution is fair, because no one outside Google can read active weights or deployment status out of a file. The attributes still exist. The architecture behind them still exists. The debate about whether Google measures topical coherence at the site level is over.

What changed

Before the leak, skeptics like Kevin Indig argued that topical authority was an “SEO ghost concept,” a narrative that practitioners imposed on correlation data. The leak confirmed siteFocusScore and siteRadius. Indig then reversed his position in public, and proposed Topic Share as the operational metric for tracking topical authority over time.

Four Overlapping Systems Produce the Effect


Those attributes do not add up to one score. The effect that practitioners observe comes from four overlapping systems, and each one works at a different level of granularity.

QualityAuthorityTopicEmbeddings place a site in a vector space next to every other site. The embeddings propagate to SuperRoot, the final ranking layer. Close embeddings mean two sites cover related territory. Distant embeddings mean they do not. The embedding is how Google groups a personal finance blog with a bank's advice section, even when the two share no backlinks.

NSR (Normalized Site Rank) is a 63+ field site-level quality scoring system. Inside it, NsrChunks splits the site into topical sections and scores each one on its own. A blog section can hold a different NSR chunk score than the product pages on the same domain. ClusterUplift groups a site with similar sites, then applies a collective boost or demotion to the whole cluster. A quality problem inside the cluster demotes every site in it, including the clean ones.

PairwiseQ comparisons favor the site with deeper topical coverage when Google matches two competitors head to head. NLP entity coverage, the breadth and depth of entity recognition inside the content, feeds Google's Entity-Based Ranking patent (US10235423). That patent assigns a composite score from knowledge graph metrics, weighted by entity type.

Source:Google API Leak (May 2024), iPullRank / Hobo Web analysis
SignalWhat It MeasuresLevel
siteFocusScoreHow dedicated a site is to a single topicSite
siteRadiusHow much a page deviates from the site's central themePage → Site
site2vecEmbeddingEncodedCompressed vector embedding of a site's overall themeSite
pageEmbeddingsPer-page vectors compared against site embeddingsPage
QualityAuthorityTopicEmbeddingsMulti-dimensional vector positioning site relative to all othersSite
NsrChunksIndependent quality evaluation per topical section of a siteSection
ClusterUpliftCollective quality boosts/demotions applied to similar-site clustersCluster

A practitioner who reports “I tested topical authority and it did not work” tested one dimension of a system with four. Fifty thin articles do not move the topic embedding vector. Twenty deep, entity-rich, interlinked articles do. Depth and entity coverage shift the vector. Volume on its own does not.

Signal versioning

Google versions all of these signals, and runs live experiments with different weightings at the same time. One site can rank differently under two experimental versions of the same signal. Signal versioning is a source of ranking movement with no visible external cause.

Focus and Radius Sort Sites Into Three Positions


Two of those signals, siteFocusScore and siteRadius, interact in a way that puts every domain in one of three positions. Hobo Web's analysis of the leak data names them:

Source:Hobo Web, API Leak Analysis (2024)
ArchetypeDescriptionImplication
Perfect TopicalityEvery page has low siteRadius, tight coherence around a single themeMaximum siteFocusScore. The specialist advantage.
High Focus with Topical DriftStrong core topic, but outlier pages raise siteRadiusPruning or improving off-topic pages strengthens calculated authority.
Generalist with Niche CoreExpertise diluted by tangential content coverageThe blog-around-everything pattern. Depth obscured by breadth.

A position is not permanent. When a site removes or improves its off-topic content, the calculated authority rises. A site that already has depth on its core topic does not always need more content about that topic. It may need to remove the pages that dilute the signal. The Panda patent (US9031929) works the same way at the site level. Thin or irrelevant pages pull down the score for the whole domain.

Text Relevance Dominates Rankings


A position in topic space matters because of what it buys. The first return is relevance. The Semrush 2024 Ranking Factors Study analyzed 16,298 English keywords across the top 20 positions, and evaluated 65 factors. Text relevance, how closely the content of a page matches the query, showed the strongest correlation with rankings at 0.47. The next-strongest factor reaches less than half of that.

Source:Semrush (16,298 keywords), Surfer SEO (260K SERPs), Ahrefs (5 tools)
FactorCorrelationSource
Text relevance0.47Semrush (16,298 keywords)
URL organic traffic0.33Semrush (16,298 keywords)
Domain authority0.21Semrush (16,298 keywords)
Content quality score0.17Semrush (16,298 keywords)
Content comprehensiveness0.17Surfer SEO (260K SERPs)
Content tool score → rankingWeakAhrefs (5 tools tested)
Ranking factor correlations: Text relevance 0.47, URL organic traffic 0.33, Domain authority 0.21, Content quality 0.17, Content comprehensiveness 0.17

Pages ranking for one keyword have significantly better odds of ranking for related keywords, supporting the topical cluster thesis. But there is an important nuance: content comprehensiveness alone shows only a 0.17 correlation with rankings (Surfer SEO, 260K SERPs). And when Ahrefs tested five content optimization tools (Surfer, Frase, NeuronWriter, Clearscope, AI Content Helper), they found weak correlations across the board. Content tool scores do not strongly predict rankings on their own.

The distinction matters: text relevance (are you writing about the thing the user searched for?) is strong. Content comprehensiveness (did you cover every subtopic?) is necessary but not sufficient. Topical authority appears to operate as a multiplier on relevance. Deep topical coverage increases the probability that any individual page achieves high text relevance for its target query.

A separate Surfer SEO / WLDM study (~260,000 SERPs) found that page-level topical authority was the largest on-page ranking factor, stronger than domain monthly traffic volume. This distinguishes page-level from domain-level signals: a highly relevant page on a lower-authority domain can outperform an irrelevant page on a stronger domain.

Authority Accelerates Visibility


Relevance decides whether a page can rank. Authority decides how fast it gets there. A Graphite study tracked 332 URLs published across 12 domains in June and July 2023. Content on domains with high topical authority gained visibility 57% faster. That content was 62% more likely to get traffic in the first week, and it reached impression milestones 30% faster.

Source:Multiple sources (2024–2025)
MetricResultSourceConfidence
Visibility speed (high TA vs. low)57% fasterGraphite (332 URLs)Moderate
First-week traffic likelihood62% more likelyGraphite (332 URLs)Moderate
Impression milestone speed30% fasterGraphite (332 URLs)Moderate
Niche Expertise algorithm weight~13%First Page Sage (2025)Moderate
Topic cluster traffic (case study)500 → 190K monthlyHubSpot (2024)Moderate (single case study)
Fan-out coverage ↔ AI citation0.77 SpearmanSurfer SEO (36M AIOs)High
Topical authority accelerates visibility: 57% faster visibility, 62% more likely first-week traffic, 30% faster impression milestones, 500 to 190K topic cluster traffic (HubSpot case study)

The sample is small at 332 URLs. The controlled method and the consistent direction across the metrics still give it moderate confidence. First Page Sage's 2025 ranking factor analysis points the same way.

That analysis weights Niche Expertise, defined as 10 or more authoritative pages around one hub keyword, at about 13% of the ranking algorithm. That weight makes it the fourth-highest factor in their model. They also introduce “Net DR”. A DR 40 domain that outranks a DR 70 domain nearly always holds higher Niche Expertise.

Those numbers are the quantitative evidence for what the clinical diagnostic framework calls the evidence-builder loop. Authority priors from earlier wins compound the return on later content. Sequence matters. A publisher that moves into topics where it already has depth sees a faster return than one that scatters coverage across new topics.

Practitioner note

NavBoost operates on a rolling 13-month window of click data, segmented per topic. Topical authority compounds behavioral signals within a topic cluster. Each new page benefits from the accumulated click history of existing pages in the same cluster. New domains face a structural disadvantage: no click history means no NavBoost signal, regardless of content quality.

Fan-Out Coverage Predicts AI Citation


The third return is citation. AI answer systems select sources by topical coverage, and the strongest evidence for that comes from Surfer SEO's AI Citation Report (2025), which covers 36 million AI Overviews and 46 million citations. Pages that rank for fan-out queries, the sub-queries an AI system generates when it decomposes a question, are 161% more likely to earn a citation. Fan-out coverage and citation likelihood correlate at 0.77 Spearman.

Topical depth produces that coverage without separate work. Gemini 3 decomposes one prompt into 10.7 sub-queries on average. Content that addresses many facets of a topic captures more of those sub-queries, and this is where topical authority and AI citation mechanics converge.

A practitioner cannot chase the sub-queries directly. Only 27% of fan-out sub-queries stay stable across repeated searches. You optimize for topical coverage, and the fan-out capture follows from it.

Source:Surfer SEO, Digital Bloom, AiModeBoost (2025)
MetricValueSource
Fan-out coverage → AI citation likelihood161% more likelySurfer SEO (36M AIOs, 46M citations)
Fan-out sub-query stabilityOnly 27% stableSurfer SEO (36M AIOs)
Brand search volume → AI citation0.334 correlationDigital Bloom (680M+ citations)
Entity-rich vs. keyword-optimized content267% more AI citationsAiModeBoost
Entity ID matching (Wikidata Q-IDs)8.9x citation increaseAiModeBoost
Multi-platform presence (4+ channels)2.8x more AI mentionsDigital Bloom (680M+ citations)
AI citation predictors: Knowledge graph alignment 89% correlation, Fan-out query coverage 0.77 Spearman, Brand search volume 0.334 correlation

Coverage decides which pages qualify. Entity recognition decides which brand the system reaches for first. The Digital Bloom 2025 AI Visibility Report covers more than 680 million citations. It found that brand search volume predicts AI citations better than backlinks do, at 0.334 correlation. A brand present on four or more platforms is 2.8x more likely to appear in a ChatGPT response.

Topical authority in AI retrieval covers more than what a site publishes. It covers whether the system recognizes the site as an entity worth citing on the topic.

AiModeBoost's entity research (67,394 content pieces) gives the mechanism. Entity-rich content earns 267% more AI citations than keyword-optimized content. Entity ID matching, with Wikidata Q-IDs and Knowledge Graph MIDs, produces an 8.9x citation increase. Knowledge graph alignment and AI visibility correlate at 89%.

Recognition concentrates hard. Goodie AI's September 2025 analysis of 5.7 million citations found that the top 50 domains take about 53% of all citations, and that more than 40,000 sites share the rest. Brands in the top quartile for web mentions receive 10x more AI Overview citations than the quartile below them. Incremental mentions below that top quartile return very little. The curve is a threshold, not a gradient.

Third-party platforms move a brand toward that threshold. A ConvertMate study (method undisclosed, treat it as directional) found that active profiles on Trustpilot, G2, and Capterra correlate with a 3x higher ChatGPT citation rate. Review platforms validate an entity for the comparison and evaluation queries, where an AI system needs credibility before it cites a source.

The mix of cited content types moves with the intent. Omniscient Digital analyzed 23,387 sources in January 2026. For branded queries, third-party validation dominates. Reviews and social proof account for 57% of citations, directories for 17%, product pages for 12%, and thought leadership for 5.4%. Brand-level authority in AI search rests on third-party signals as much as on the depth of owned content.

The platform moves it too. Yext's analysis of 6.8 million citations (via Surfer SEO, February 2026) controlled for intent. Brand-controlled sources account for 86% of citations. Reddit accounts for 2% in that context. On Perplexity, Reddit takes 46.7% of evidence citations. Reddit dominance is specific to a platform and an intent.

Owned content and structured data perform better across AI systems, except where the platform privileges forum evidence.

Information Gain Decides Who Gets Cited


Depth gets a site into the candidate set. Depth does not decide which candidate the system quotes. Google's Information Gain patent (US20200349181A1, filed 2018, granted June 2024) scores a document on how much new content it holds beyond what the user already saw. Google can demote or drop a document that scores near zero.

Information gain and topical authority measure different things. Authority is credibility from consistent expertise, a record of deep and accurate coverage inside one topic. Information gain is the novelty of one contribution. The two signals complement each other.

Clearscope defines information gain in operational terms. Survey data, case studies, interviews, personal stories, and novel perspectives create it. A site with high topical authority that publishes the same synthesis as everyone else scores low. A site with moderate authority that publishes original research scores high. Clearscope names the target: “concepts and entities on the fringe of Google's Knowledge Graph for the topic.”

The strongest position holds both: enough topical authority to be credible on the subject, plus novel data to be worth the citation over a competitor. Programmatic SEO architecture built on a proprietary data asset holds both at once. Every programmatic page carries information gain, because the data under it does not exist anywhere else.

Off-Topic Pages Subtract From the Same Score


Both signals work in the negative direction as well. Off-topic content dilutes the score that depth builds. NsrChunks evaluates topical sections on their own, so a blog section full of off-topic pages drags down the chunk score for the core topic. The quality of the individual off-topic page does not save it.

Keyword Insights documented a case where a large travel client narrowed its property type pages from 413 to 85, and cut about 15 million URLs. Organic traffic rose 110% almost immediately. The pruning removed the pages that split signals and diluted the topical coherence of the domain.

The reverse error costs as much. Ahrefs identified 9,700 cases of “keyword cannibalization” and found that almost none needed a fix. In a sample of 80 keywords with multiple ranking pages, only one needed consolidation. Multiple rankings produce cannibalization and diversification at the same time.

Diagnostic approach

Not all keyword overlap is cannibalization. Over-consolidation can be as harmful as over-clustering. Diagnose via GSC: are multiple pages splitting impressions with declining CTR? If impressions are growing and CTR is stable, the multiple rankings are diversification, not cannibalization. Distinguish between the two before acting.

A site with broad coverage that underperforms on its core queries has a dilution problem until the Clinical Retrieval and Ranking Framework rules one out. Test for topical dilution before you treat the core content as the fault. The binding constraint may be overclustering.

Topic Share Is the Number You Can Measure


None of these signals appear in a report that a practitioner can open. The leak names what Google measures, and it does not expose siteFocusScore or NsrChunks. Kevin Indig's Topic Share framework fills the gap. Topic Share is the percentage of organic traffic that a domain captures from all the keywords inside a defined topic, measured against its competitors.

Measurement, through Ahrefs: identify the head entity or term, which must hold a Knowledge Panel. Extract the matching keywords at 10 or more monthly searches. Upload the keyword set in Keyword Explorer, then pull “Traffic Share by Domains”. Topic Share is the aggregate traffic percentage of your domain inside that topic.

The metric composites rank position, search volume competitiveness, multi-keyword rankings, SERP feature capture, and snippet performance into one number. Topic Share does not map to any single Google signal. It reflects the outcome of all of them together.

Benchmarks from Indig's data: in ecommerce (29K+ keywords), Shopify holds 11% Topic Share and BigCommerce 10%. In spend analysis (142 keywords), Jaggaer holds 15% and Sievo 13%. A monthly measurement cadence is correct, because anything more frequent is noise.

Limitation: Topic Share is a proxy for competitive share. The metric is not a direct measurement of the internal authority signals at Google. A site could hold high Topic Share from brand searches alone. Pair the metric with an entity coverage assessment and a content depth analysis for the full picture.

Folklore Became an Engineering Problem


Six conclusions follow from the evidence above, and each one names work that a team can do.

Depth beats breadth. Google's topic embeddings reward sites that go deep on a subject, not sites that go wide across many. Twenty deeply comprehensive, entity-rich articles will move the site2vec embedding vector more than fifty thin ones. The N-gram Quality Prediction patent (US9767157) detects thin content via phrase-frequency fingerprinting, the mechanism behind why volume without depth fails.

Pruning moves the score. When a team removes or noindexes off-topic content, the siteFocusScore calculation improves. The Keyword Insights case study is not a hypothetical. It showed a 110% traffic increase from a cut of 413 pages to 85. IBM, Progressive, and DoorDash have all reported organic traffic gains after pruning.

Sequencing matters. The evidence-builder loop is real: high topical authority content gains visibility 57% faster. NavBoost's per-topic click signals compound within topic clusters. Start with queries where you already have depth, accumulate authority priors, then extend into adjacent topics. The clinical framework's ceiling vs. weight distinction applies: topical authority is a ceiling issue, not a weight issue. No amount of link building overcomes weak topical coherence.

Entity coverage is the mechanism. Topical authority requires entity recognition depth, not surface-level topic coverage. The content must give every relevant entity its detailed attributes, facts, and relationships. Koray Tugberk's semantic SEO methodology gives the operational rules. Match entity types between question and answer. Reduce dependency hops, disambiguate entities, boost salience through co-occurring terms, and avoid unclear antecedents. His case studies show results without link building or brand power. GetWordly.com went from zero to 128,000 organic traffic in 123 days through semantic methodology alone.

Source:Oncrawl / Koray Tugberk (2023–2024)
ProjectResultMethod
GetWordly.com0 → 128,000 organic traffic in 123 daysSemantic SEO methodology
Interingilizce.com10,000 → 200,000+ monthly in 5 monthsSemantic SEO methodology
Third client600% growth in 5 months (10,000 → 70,000 monthly)Semantic SEO methodology

AI citation follows topical authority. Fan-out coverage at 0.77 Spearman is the strongest known predictor of AI citation. Entity-rich content earns 267% more AI citations. Knowledge graph alignment correlates at 89% with AI visibility. A team that builds topical authority builds AI citation infrastructure at the same time. The depth that signals expertise to Google's ranking systems creates the coverage that AI systems need for extractive grounding.

Cluster quality drags affect innocent sites. ClusterUplift means the quality of the sites Google groups you with sets part of your competitive position. A site clustered with low-quality sites in the same niche takes the cluster-level demotion, whatever its individual signals say. Cluster demotion explains why one niche loses visibility in an update while another holds its position. It also makes differentiation from the cluster a competitive necessity, through novel data, better entity coverage, or a stronger technical implementation.

Topical authority stopped being folklore in May 2024. The signals have names, a level, and a place in the ranking stack. The concept lost its mystique and gained a specification.

A team can name the pages that pull the site vector off center. The same team counts the entities that the topic requires. It then measures its share of the topic against the competitors Google clusters it with. Nobody has to argue about whether the signal exists. The engineering work under it still waits.