Note: This is part 2 of 3 articles today on NSR. Part 1 is here.
When the 14,000-file Google Content Warehouse API leak exposed the internal plumbing of Google search, lots of the SEO industry suffered from a collective case of shiny-object syndrome. Commentators (even myself) fixated on clickstream metrics, sandboxing debates, and isolated feature flags, treating the repository like an unorganised garage sale.
Meanwhile, right in plain sight, the master switchboard for domain-wide trust and site quality sat largely undocumented by the mainstream news: QualityNsrNsrData.
For over a decade, Google’s public relations team maintained a firm party line: Google does not use a generalised, site-wide quality score; every page stands on its own merits. The API leak indicates that the picture is more complicated.
Through the New Site Rank (NSR) architecture, Google maintains a sophisticated telemetry engine that grades, chunks, and scores entire domains before individual pages ever compete in real-time retrieval.
This article, I think, is the first public compilation of this proto as the sitewide quality list. QualityNsrNsrData is the actual module that lists Google’s sitewide ranking signals (although not the only one).
QualityNsrNsrData is New Site Rank.
Part 1: The Evolution from a mental or isolated concept.
Fast forward to May 2024, and the GitHub API source code revealed that NSR had matured into a sprawling, production-grade telemetry pipeline.
2019 Slides to 2024 API Schema
To understand the weight of QualityNsrNsrData, we have to look backwards. In 2019, internal slide decks leaked to the public, introducing the concept of New Site Rank (NSR)—an evaluation layer designed to assess new and existing domains utilising human rater training patterns.
At the time, sceptics dismissed it as an experiment.
The QualityNsrNsrData module is the code-level manifestation of those early slides. It bridges historical intent with modern machine learning, translating abstract quality rater guidelines into concrete, automated code variables.
Part 2: Why the SEO Industry Missed the Big Picture

While automated scrapers indexed the module name immediately, very few analysts recognised QualityNsrNsrData as the holy grail of site-wide algorithmic control. Why did the broader industry drop the ball?
- The Confirmation Bias Trap: Because Google spokespeople repeatedly denied the existence of domain-wide quality metrics, most SEO professionals assumed site-wide penalties and scores were a myth.
- Information Overload: Sifting through thousands of complex Elixir and Python data schemas requires deep forensic engineering patience. Most commentators lacked the bandwidth to trace how individual modules execute upstream.
- Failure to Connect the Dots: Very few researchers bridged the historical gap between the 2019 NSR slide leaks and the 2024 API codebase.
We can now unlock this module for what it truly is: Google’s master control panel for systemic domain evaluation.
I missed it too until I started to investigate what NSR meant in the Google leak (I had put working out that until later; there was no obvious proof of what it signified in the leak).
Part 3: Anatomy of the Site-Wide Telemetry Engine

The architecture of QualityNsrNsrData is divided into distinct functional categories that dictate how a website is perceived by core ranking systems like Mustang.
1. Core NSR & Trust Metrics
nsr: The primary numerical score assigned to a site chunk, grading overall domain authority and reliability.priorAdjustedNsr: A comparative baseline tracking whether a domain sits above or below the expected average for its specific content niche.nsrOverrideBid: An emergency manual switch (go/nsr-override-bid) enabling engineers to unconditionally override NSR during severe ranking anomalies or safety crises.
2. Statistical Variance and Consistency Checks
Google doesn’t just look at a site’s masterpiece articles; it evaluates systemic predictability:
siteQualityStddev: Estimates the standard deviation (spread) of page-level quality ratings across the domain. High variance signals an erratic site that mixes great content with low-grade filler, dragging down baseline trust.chardVariance&nsrVariance: Statistical error models that measure how uniformly good or bad a site’s text content is across its entire URL structure.
3. Spam, AI Abuse, and Clutter Filters
spambrainLavcScore: Direct integration with SpamBrain, evaluating site-level abuse, low quality, and security vulnerabilities.racterScores: Tied to internal Project Racter initiatives, identifying scaled, automated, low-value AI text generation.clutterScore: A penalty vector penalising domains flooded with excessive, distracting ads or aggressive UI pop-ups (go/clutter-v0).
4. Brand Affinity & Graph Topologies
pnav&pnavClicks: Navigational query fractions measuring direct brand search volume and user loyalty.site2vecEmbeddingEncoded: Compact vector representations of the entire site’s semantic profile, optimised for fast retrieval in Google’s “superroot” architecture.sitePr: Modernised domain-level PageRank estimates tracking incoming link equity distribution.
Part 4: What QualityNsrNsrData Means for Modern SEO
The exposure of QualityNsrNsrData fundamentally shifts how we must approach technical and content strategy:
- The Death of Mixed-Quality Content: Because variables like
siteQualityStddevpunish variance, publishing low-effort AI fluff alongside high-end journalism actively poisons the domain’s baseline trust score. - Systemic Health Over Isolated Optimisation: You cannot out-optimise a bad site architecture with keyword tweaks. Google evaluates the collective “content-DNA” of your site chunks.
- The Necessity of Brand Building: High
pnavand direct traffic fractions act as a defensive moat against core algorithm updates, proving real-world demand outside of algorithmic manipulation.
Reference: All QualityNsrNsrData Variables Explained
1. Core NSR & Scoring Metrics
nsr: The primary numeric New Site Rank score assigned to the site chunk.- What it is: The main overall score Google gives a website section to grade its general quality and authority.
priorAdjustedNsr: A list of versioned float signals representing NSR minus its expected prior.- What it is: A comparison score showing whether a site performs better or worse than the average baseline site in its category.
nsrEpoch: The specific temporal epoch or version from which the NSR value was compiled.- What it is: A version or date tag tracking which update cycle generated this specific score.
nsrOverrideBid: An emergency override control variable (go/nsr-override-bid).- What it is: A manual emergency switch that lets Google engineers force an unconditional score override if something breaks.
clusterId: An integer identifier for grouping sites together during ecosystem experiments.- What it is: An ID tag used to group similar websites together for testing new algorithm changes.
clusterUplift: Model data representing cluster-level score enhancements or adjustments.- What it is: A bonus score applied to a group of sites based on their collective category performance.
versionedData: A versioned map of experimental NSR values.- What it is: A container holding test scores used when Google experiments with future ranking updates.
ketoVersionedData: Versioned telemetry data managed via Google’s Keto configuration infrastructure.- What it is: Configuration and score data tracked through Google’s internal Keto system for testing.
2. Content Quality & Classifier Models
articleScore: Evaluates whether the domain publishes professional journalistic or editorial article content.- What it is: A score checking if a website looks like a real news or publishing site rather than a spam page.
articleScoreV2: An upgraded version 2 score from the site-level article classification pipeline.- What it is: An improved, updated version of the article quality checker.
chardEncoded: An integer-encoded version of the Chard score—a core content-based site quality predictor.- What it is: A numerical summary score measuring the overall text quality of the site.
chardScoreEncoded: A list of versioned integer signals representing site-level Chard scores.- What it is: Historical or tested variations of the text quality score encoded as numbers.
tofu: Unknown score.- What it is: Potentially Top-Of-Funnel content identifier, potentially Trust on First Use. This needs more analysis to confirm validity.
ugcScore: Evaluates the prevalence and quality of User-Generated Content on the domain.- What it is: A metric measuring how much user-created content (like comments or forums) exists and how good it is.
smallPersonalSite: A promotional quality score targeting small personal blogs.- What it is: A flag or score used to identify and appropriately handle personal blogs and hobby sites.
ymylNewsV2Score: A specialised Your Money or Your Life (YMYL) scoring model targeting news and high-stakes informational integrity.- What it is: A specific score evaluating sites that cover high-risk topics like health, finance, or news safety.
healthScore: Categorical signals assessing the health or medical informational quality of the domain.- What it is: A classification tag tracking how reliable health-related information on the site is.
3. Statistical Variance & Consistency
siteQualityStddev: A critical metric estimating the standard deviation (spread) of page-level Page Quality ratings across the entire domain.- What it is: A measure of how inconsistent a site’s page qualities are—high variance means it mixes great content with terrible content.
siteQualityStddevs: A list of versioned float signals tracking site quality standard deviations across multiple evaluation windows.- What it is: Historical or experimental versions tracking consistency changes over time.
nsrVariance: Predicts the statistical error or variance of the NSR score itself.- What it is: A measurement of how confident Google is in the accuracy of its main NSR score for that site.
chardVariance: Measures the spread or variance of content-based Chard scores across all pages on the site.- What it is: A check to see if a site’s text quality is steady or if it bounces wildly between good and bad articles.
chardScoreVariance: A list of versioned float signals detailing site-level Chard score variance.- What it is: Versioned records tracking how text quality variance changes across different updates.
4. Spam, Abuse & Machine-Generated Content
spambrainLavcScore: A site-level SpamBrain Low-Quality/Abuse/Vulnerability (LAVC) score.- What it is: A score from Google’s AI spam system flagging abuse, low quality, or security vulnerabilities.
spambrainLavcScores: A list of versioned float signals tracking SpamBrain LAVC evaluations over time.- What it is: Historical or version-controlled records of SpamBrain abuse scores.
racterScores: Site-level Automated Generated Content (AGC) classification scores tied to Project Racter.- What it is: A detection score that spots mass-produced, low-value AI-generated text on the site.
clutterScore: A delta site-level signal penalising sites loaded with excessive, distracting, or annoying UI clutter and ads.- What it is: A penalty score for sites flooded with annoying pop-ups, excessive ads, or distracting elements.
clutterScores: A list of versioned float signals tracking site clutter over time.- What it is: Versioned data tracking how cluttered or ad-heavy a site is across updates.
5. Video & Shopping Specific Models
videoScore: Aggregate score assessing the presence and quality of video assets on the domain.- What it is: A metric measuring how much and how good the video content is on the site.
shoppingScore: Aggregate score evaluating e-commerce and shopping functionality across the site chunk.- What it is: A metric checking if the site functions as an online store or shopping resource.
vlq: Score generated by the Video Low Quality model.- What it is: A specific penalty score identifying low-quality video pages.
vlqNsr: NSR derived from a specialised headroom model targeting low-quality video sites.- What it is: A site-level score focusing specifically on filtering out junk video domains.
isVideoFocusedSite: A boolean bit determining if a site has mostly video content but is not hosted on known video-hosting domains.- What it is: A true/false flag identifying websites that exist mainly to host videos rather than text.
6. Site Architecture, Chunking & Identifiers
siteChunk: The primary NSR site chunk identifier.- What it is: The core URL segment or section label Google uses to group pages of a website together.
secondarySiteChunk: A secondary site chunk providing more granular sub-host structural chunking.- What it is: A more detailed subsection tag used when a website needs to be broken down into smaller chunks than just the main domain.
siteChunkSource: Annotated metadata tracked strictly within the Goldmine NSR annotator.- What it is: An internal label showing where the site chunk data was generated or tested.
nsrdataFromFallbackPatternKey: A boolean flag indicating that exact NSR data was missing for the chunk, and values are sourced from an averaged fallback pattern of peer host chunks.- What it is: A safety flag that means “we didn’t have direct data for this exact site section, so we used an average estimate from similar sites.”
host: The string identifier for the target web host.- What it is: The domain name or web address (like example.com) being evaluated.
url: The specific representative URL string linked to the record.- What it is: The exact web link being looked at.
language: The integer code representing the primary language of the domain slice.- What it is: A number code identifying what language the site is written in.
i18nBucket: Internationalisation bucket classification.- What it is: A categorisation tag grouping sites by global region or language market.
7. Traffic, Engagement & Graph Topologies
impressions: Aggregate site-level search impression volume.- What it is: A count of how often pages from this site appeared in Google search results.
chromeInTotal: Telemetry tracking site-level view frequencies inside the Google Chrome browser.- What it is: Usage data showing how frequently real humans visit the site using the Chrome browser.
directFrac: The fraction of direct traffic attracted by the domain.- What it is: The percentage of visitors who type the URL directly into their browser rather than clicking a search link.
pnav: Fractional navigational query metrics measuring direct brand search affinity.- What it is: A score tracking how often people search specifically for this brand name by typing it into Google.
pnavClicks: The click denominator used in computing thepnavnavigational score.- What it is: The total click count used as a baseline math denominator to calculate brand search loyalty.
siteLinkIn: The average value of incoming link equity for pages within the site chunk.- What it is: A measure of how many other websites link to this domain and how authoritative those links are.
siteLinkOut: Aggregated value of outgoing link scores originating from the site chunk.- What it is: A metric measuring where and how this site links out to other parts of the web.
titlematchScore: Measures how effectively page title tags match real-world user search queries.- What it is: A check to see if the titles written on the pages actually match what people are searching for.
site2vecEmbedding: Full vector embeddings representing the semantic profile of the site.- What it is: A complex AI mathematical map capturing the overall “meaning” and topic profile of the website.
site2vecEmbeddingEncoded: Compact, encoded versions ofsite2vecembeddings optimised for storage and fast retrieval.- What it is: A compressed, lightweight version of the site’s AI topic map designed for fast searching.
siteAutopilotScore: Aggregated value of URL-level autopilot metrics assigned to the site chunk.- What it is: An automated performance or quality score managed by Google’s internal systems.
8. Authority & Special Signals
localityScore: The locality component of the Local Authority signal.- What it is: A score measuring how relevant and authoritative a site is for local geographic searches.
isCovidLocalAuthority: Boolean bit identifying domains holding verified COVID-19 local authority status.- What it is: A special true/false tag given to trusted official public health sources during the pandemic.
isElectionAuthority: Boolean bit designating trusted election authority status.- What it is: A special trust flag assigned to verified official election and voting information sites.
sitePr: Estimated site-level PageRank score.- What it is: Google’s modernised estimate of the site’s overall link authority at the domain level.
metadata: General container model holding auxiliary metadata for NSR records.- What it is: A storage folder for extra background information and notes attached to the record.
Conclusion
The discovery of QualityNsrNsrData as Google’s list of core site-wide ranking signals is a watershed moment for search marketing. QualityNsrNsrData is the leaked Protocol Buffer in which Google stores its bundle site-level quality scores.
It is the closest thing the 2024 documents give us to a list of sitewide ranking signals in one place.
It proves that Google’s algorithm is not merely a decentralised collection of isolated page evaluations, but a unified ecosystem governed by systemic, site-wide trust metrics. For advanced SEOs, understanding this schema transforms the job from a guessing game of heuristics into precise engineering alignment.
Disclosure: I use generative AI when specifically writing about my own experiences, ideas, stories, concepts, tools, tool documentation or research. My tool of choice for this process is Google Gemini. All content was verified as correct. See Hobo Web AI policy.
Disclaimer: This is not official advice from Google. It is SEO theory based on patents and code leaks. The metrics equations are borrowed from known IR (information retrieval) models rather than being extracted from leaks. Any article (like this) dealing with the Google Content Data Warehouse leak requires a lot of logical inference when putting together the framework for SEOs, as I have done with this article. I urge you to double-check my work and use critical thinking when applying anything from the leaks or patents to your site. My aim with these articles is essentially to confirm that Google does, as it claims, try to identify trusted sites to rank in its index. The aim is to irrefutably confirm white hat SEO has purpose in 2026 – and that purpose is to build high-quality websites. Feedback and corrections welcome.