

Google’s Ranking Factors, Evidence-Graded
Every “200 ranking factors” list gets copied forward from a 2013 blog post. This one doesn’t. It sorts every claim about how Google ranks pages into four evidence tiers — Google-confirmed, sworn testimony, correlation study, or unverified folklore — and shows its work for each one.
Evidence-Graded Primary Sources Cited No Invented DataThe recycled-listicle problem
Search “Google ranking factors” and you’ll get roughly the same list, republished every January with a new year in the title: title tags, backlinks, content length, Core Web Vitals, sprinkle in “E-E-A-T,” done. Most of it traces back to correlation studies from the 2010s, restated as if they were confirmed mechanics. Some of it is outright folklore Google has explicitly denied for a decade.
Two events in 2024 changed how much of this we can actually verify. First, in March 2024, more than 2,500 pages of internal documentation from Google’s Content Warehouse API — describing 2,596 modules and roughly 14,000 internal attributes — were published to a public GitHub repository, apparently by mistake, and later analyzed publicly by SparkToro’s Rand Fishkin and iPullRank’s Mike King. Second, sworn testimony from Google’s own Vice President of Search, Pandu Nayak, in the Department of Justice’s antitrust case against Google (United States v. Google LLC) put some of Google’s ranking mechanics on the public record under oath — including several statements that directly contradicted years of public messaging from Google spokespeople.
Neither of those sources hands you a weighted formula. Google has been consistent, before and after both events, that raw attribute names in leaked documentation don’t tell you what’s actually used in live ranking, or how heavily. That caution is fair — and it’s exactly why this piece grades every claim by where it actually comes from, instead of treating all of it as equally solid.
The Evidence Tier framework
Before ranking any factor, ask one question: who said so, and how would they know? Every claim in this piece — and, frankly, every claim in every ranking-factors article you’ll ever read — falls into one of four tiers.
| Tier | What qualifies | What it can’t tell you |
|---|---|---|
| Tier 1 | Google’s own Search Central documentation, named algorithm updates, on-the-record statements from Google Search Liaison | Exact weighting versus other signals |
| Tier 2 | Sworn DOJ trial testimony, the 2024 Content Warehouse API leak documents | Whether a documented attribute is actively used, deprecated, or experimental |
| Tier 3 | Large-sample correlation studies from SEO analytics vendors (Backlinko, Ahrefs, SEMrush) | Causation — correlation studies cannot show that changing X causes rankings to move |
| Tier 4 | Community folklore, single-site anecdotes, claims Google has explicitly denied | Almost anything reliable — treat as a hypothesis to test, not a rule |
Tier 1: what Google has actually confirmed
This is the shortest list, and that’s the point — it’s short because it’s solid.
- Tier 1 Core Web Vitals, updated. In March 2024, Google officially replaced First Input Delay (FID) with Interaction to Next Paint (INP) as the third Core Web Vital, alongside Largest Contentful Paint and Cumulative Layout Shift. This was announced and documented directly by Google, not inferred.
- Tier 1 The Helpful Content System. Google’s people-first content guidance, originally a standalone classifier introduced in 2022, was folded into Google’s core ranking systems in 2023–2024. Google’s own documentation says it rewards content demonstrating first-hand experience and penalizes content produced primarily to manipulate rankings.
- Tier 1 Scaled content abuse is an explicit spam policy. Google’s spam policies name the practice of publishing large volumes of content — AI-generated or otherwise — primarily to manipulate search rankings, independent of whether any individual page is well written.
- Tier 1 Mobile-first indexing is complete. Google confirmed it now uses the mobile version of a page’s content for indexing and ranking for essentially the entire index.
Notice what’s missing: no confirmed weighting, no confirmed “you need X backlinks” threshold, no confirmed word count. Google has repeatedly said publicly that it does not use a fixed checklist — and on this narrow point, Tier 1 evidence and Google’s messaging actually agree.
Tier 2: sworn testimony and the leak — where the real news is
This is the tier that actually moved the conversation in 2024, because it’s the first time outside observers got anything resembling primary-source material instead of inference.
The leak, briefly
In March 2024, internal documentation for Google’s Content Warehouse API — reportedly pushed live by an automated internal tool — sat briefly on a public GitHub repository before being discovered. Erfan Azimi of EA Eagle Digital found it, corroborated it with former Google employees, and shared it with Rand Fishkin, co-founder of SparkToro, on May 5, 2024. Fishkin spent weeks verifying it before publishing jointly with Mike King of iPullRank on May 27, 2024. The documentation describes 2,596 modules and roughly 14,000 attributes — module and field names, not a ranked list of factors, and critically, not their weights.
Google’s response, delivered through Search Liaison Danny Sullivan, was that outside parties risked drawing incorrect conclusions from documentation lacking the context of how — or whether — each attribute is actually used in live ranking. That’s a legitimate methodological caveat, and it applies to every specific attribute name floating around SEO Twitter since: a field existing in an API schema is not proof it drives rankings today.
The DOJ testimony, briefly
Separately, in the 2023 proceedings of United States v. Google LLC, the federal antitrust case that a U.S. District Court ultimately ruled against Google in August 2024, Google’s own Vice President of Search, Pandu Nayak, testified under oath about ranking mechanics — testimony that, unlike the leak, Google cannot dismiss as decontextualized, because Google’s own executive said it.
The headline revelation was NavBoost, a click-based re-ranking system Nayak testified has existed since roughly 2005 — directly contradicting years of public statements from Google representatives (including former head of web spam Matt Cutts) that click-through data was too noisy and too easy to manipulate to use in ranking.
How a page actually gets ranked, per the testimony
The most useful thing to come out of the trial isn’t any single factor — it’s the shape of the pipeline. Ranking is staged, not a flat scorecard:
Two things are worth sitting with. First, NavBoost isn’t a fine-tuning tweak — testimony describes it acting early, culling tens of thousands of candidates down to a few hundred before the more expensive machine-learned systems ever see them. If your page doesn’t survive that click-based cull, later-stage quality signals may never get the chance to matter. Second, “twiddlers” — Google’s internal term for post-hoc adjustment algorithms — exist specifically to correct cases where click data alone would produce bad results, like an engaging-but-inappropriate page outranking a duller, more authoritative one. NavBoost is powerful, but testimony indicates it does not have unchecked authority over the final result.
Data point worth flagging
Tier 2 Testimony indicates NavBoost’s click-data window was 18 months prior to roughly 2017, and has run on a rolling 13-month window since. That’s a real, sourced detail — not a stat we generated to sound rigorous.
Tier 3: correlation studies — useful, but read the fine print
Correlation studies from SEO analytics vendors are the backbone of most “ranking factors” content, and they’re not worthless — they just answer a narrower question than most articles imply. A correlation study can tell you that pages ranking on page one tend to share a trait. It cannot tell you that adopting the trait causes the ranking, because top-ranking pages might simply share some third cause — bigger budgets, older domains, more resources for original research — that produces both the trait and the ranking.
| Correlation commonly cited | Tier | The honest caveat |
|---|---|---|
| Longer content correlates with page-one rankings; industry studies commonly cite an average around 1,400 words for first-page results | Tier 3 | Longer pages often cover more subtopics and attract more links — length itself is likely a proxy, not the mechanism |
| Pages in the #1 position tend to have substantially more referring domains than pages ranked #2–#10 | Tier 3 | Older, more resourced sites accumulate both links and rankings — reverse causation is plausible |
| Lower bounce rate correlates with higher rankings | Tier 3 | Bounce rate is also a symptom of good intent-matching, which is itself Tier-1 confirmed to matter — this may be double-counting one cause |
The honest way to use Tier 3 evidence: treat it as a signal about what good pages tend to look like, not a spec sheet for what to build. Write the length the topic actually requires. Earn links by being citation-worthy, not by hitting a target number.
Tier 4: folklore and denied claims still floating around
These persist in briefs and client calls years after Google addressed them directly:
- Tier 4 The meta keywords tag affects rankings. Google confirmed this was ignored well over a decade ago; it still shows up in “on-page checklists” today.
- Tier 4 A fixed “sandbox” penalty holds back all new domains for a set period. Google representatives have given inconsistent public answers here; treat it as unconfirmed rather than as a rule to plan around.
- Tier 4 Domain age, by itself, is a ranking factor. No Tier 1 or Tier 2 source confirms domain age as an independent signal; older domains simply tend to have more of the things that are associated with rankings — content, links, history.
None of this means these tactics never “worked” for anyone. It means the mechanism claimed for why they worked doesn’t hold up against the sources we now actually have.
A mistake I made, and an unpopular take
Where I got it wrong
For years I told clients backlink volume was close to the whole game — get enough links from decent domains and rankings would follow, more or less regardless of how satisfying the page actually was to read once someone landed on it. NavBoost’s role in the pipeline argues against that mental model. If a click-based cull happens before the expensive quality models even run, a page that earns clicks but disappoints readers on arrival is fighting the pipeline, not working with it. Backlinks still matter — Tier 3 evidence is consistent on that — but they’re not a substitute for the page holding up once someone’s actually reading it.
The unpopular take
Most “ranking factors” content isn’t wrong so much as it’s answering a question nobody should still be asking: “what are the factors?” The more useful question, now that Tier 1 and Tier 2 sources actually exist, is “what tier is this claim in, and does that tier justify the resources I’m about to spend on it?” A lot of SEO budgets go toward Tier 4 folklore because it’s actionable and cheap to explain in a slide deck, while Tier 1 fundamentals — genuinely useful, well-differentiated content that survives a click-based cull — get treated as a given rather than the actual lever.
A self-audit checklist
Before publishing or refreshing a page, run it against what’s actually verified:
- Does the title tag and opening section make the page’s value obvious enough to earn a click and survive the moment right after the click — the “lastLongestClicks” concept from the leaked schema, in plain terms?
- Have you checked Interaction to Next Paint alongside Largest Contentful Paint and Cumulative Layout Shift — the current, Tier 1 Core Web Vitals — rather than optimizing for the retired First Input Delay metric?
- Does the page demonstrate first-hand experience or original analysis, per Google’s own Helpful Content guidance, or could it have been assembled entirely from other people’s pages?
- Is the word count driven by what the topic needs, not a target pulled from a correlation study?
- Would this page still be worth publishing if no one ever linked to it — i.e., does it stand on its own value, independent of the Tier 3 signals it might eventually attract?
- Have you removed any tactic on this page whose only justification is Tier 4 folklore?
Frequently asked questions
What exactly was the Google Content Warehouse API leak?
Internal documentation describing roughly 2,596 modules and 14,000 attributes from Google’s Content Warehouse API appeared on a public GitHub repository in March 2024, reportedly through an internal tool that pushed it live by mistake. Erfan Azimi discovered and corroborated it, then shared it with Rand Fishkin (SparkToro) and Mike King (iPullRank), who published joint analyses on May 27, 2024.
Does a field name in the leaked documents prove Google uses that factor?
Not by itself. Google’s own response cautioned that documentation describing what a system can store or compute doesn’t confirm what’s actively used in live ranking, at what weight, or whether it has since been deprecated. Treat leak-derived claims as Tier 2 — informative about Google’s architecture, not a confirmed weighting.
Is NavBoost a ranking factor, or something else?
It’s more accurate to call it a re-ranking and filtering system. Per sworn testimony, it uses historical click data to narrow a large candidate pool down to a few hundred results early in the pipeline, before machine-learned rerankers do final scoring — rather than adding a small boost at the very end.
Does Google use Chrome browsing data for rankings?
Google representatives publicly denied this for years. Reporting on the leaked documentation and DOJ trial materials describes this denial as contradicted by internal evidence, though the precise scope of what browser-level data feeds into NavBoost versus other systems isn’t fully public. This is a Tier 2 claim: credible, but not exhaustively detailed.
What replaced First Input Delay (FID) in Core Web Vitals?
Interaction to Next Paint (INP) officially replaced FID as a Core Web Vital in March 2024. This is Tier 1 — announced directly by Google.
Is content length a real ranking factor?
Not on its own. Correlation studies (Tier 3) consistently find longer content on page one, but length is best understood as a byproduct of topical depth and link-worthiness rather than a lever to pull independently. Padding a thin topic to hit a word count doesn’t recreate the underlying cause.
What are “Twiddlers”?
Internal Google term, surfaced in DOJ trial materials, for a class of algorithms that adjust rankings after the main scoring pipeline — for example, to protect diversity of results, demote low-quality or unsafe content that clicks well, or apply freshness boosts. They exist as a check on systems like NavBoost.
Did the DOJ ruling against Google change how rankings work?
The August 2024 ruling found Google acted illegally to maintain a monopoly in general search; as of this writing the case is in the remedies phase, and any structural changes to Google’s ranking or distribution practices would follow from that phase, not from the ruling itself.
Does Google penalize AI-generated content specifically?
Google’s spam policies target “scaled content abuse” — mass-producing content, by any method, primarily to manipulate rankings — rather than AI authorship itself. Google has stated that well-produced, genuinely helpful content is not penalized purely for involving AI assistance in production.
What should I actually prioritize given all this?
Spend most of your effort on Tier 1 fundamentals you fully control — genuine expertise and first-hand detail, page experience, technical health — treat Tier 2 mechanics as context for why those fundamentals matter, use Tier 3 studies as loose signal rather than targets, and stop budgeting time for Tier 4 folklore.
Sources
- Rand Fishkin, “An Anonymous Source Shared Thousands of Leaked Google Search API Documents With Me,” SparkToro, May 2024 — sparktoro.com
- Barry Schwartz, “Google Search Document Leak Reveals Ranking Details,” Search Engine Land, May 2024 — searchengineland.com
- Hobo Web, “Key Strategic SEO Insights from the U.S. DOJ v. Google Antitrust Trial,” analysis of Pandu Nayak testimony — hobo-web.co.uk
- JC Chouinard, index of DOJ trial exhibits including the Testimony of Pandu Nayak — jcchouinard.com
- Google Search Central, Core Web Vitals documentation (INP replacing FID, March 2024) — developer.chrome.com
- Google Search Central, Helpful Content and Spam Policies documentation — developers.google.com
This piece synthesizes public reporting on the 2024 Content Warehouse API leak, publicly available DOJ trial materials, and named third-party correlation studies. Where a claim rests on secondary analysis rather than a primary document we could verify directly, it is graded Tier 2 or Tier 3 rather than presented as confirmed.
Last updated: 28 August 2026 · Next scheduled review: when Google publishes further remedies-phase disclosures or a comparable primary-source event occurs
