


Everyone tells you to “add semantic keywords.” Almost nobody explains what Google actually does with them. This is that article — including the parts that should make you uncomfortable about how you’ve been writing content.
Here’s the conversation I keep having: someone shows me a piece of content that isn’t ranking, I look at it, and the first thing I notice is that it contains the target keyword 47 times. The second thing I notice is that it contains nothing else. No surrounding context. No related concepts. No entities that tell Google what conceptual territory this page occupies. Just a keyword, repeated, like a panic attack in HTML.
These people have heard about “LSI keywords.” They’ve bought tools that generate lists of them. And they’re confidently chasing a technology that Google doesn’t use — and hasn’t used — for web search. At any point in its existence.
This article is about what Google actually uses instead, and why the difference matters more in 2026 than it ever has before.
If you’ve been buying “LSI keyword” tool subscriptions, you’ve been paying for a 1980s Bell Labs technology that Google has repeatedly and publicly said it does not use. The rest of this article explains what you should be doing instead — and how the same underlying idea (surrounding context matters) is right, even if the implementation has been completely wrong.
The LSI Myth That Won’t Die (And Why It Keeps Getting Sold To You)
Latent Semantic Indexing was invented in the late 1980s at Bell Labs. It used a mathematical operation called singular value decomposition to find patterns of word co-occurrence in small, static document sets. It was impressive for its time. It is not used by Google. Full stop.
In 2019, John Mueller — Google’s Senior Webmaster Trends Analyst — was direct: “There’s no such thing as LSI keywords — anyone who’s telling you otherwise is mistaken, sorry.” This wasn’t casual dismissal. It was a correction aimed at an entire industry that had built tools, courses, and methodologies around a fabricated concept.
The reason the myth persists is that the underlying intuition is correct. Surrounding context does matter. Related terms do help Google understand your page. But the mechanism is nothing like LSI — and the difference between “LSI thinking” and “semantic keyword thinking” is the difference between guessing and understanding how the machine actually works.
What Semantic Keywords Actually Are
Semantic keywords are contextually relevant terms, entities, and concepts that help a search engine understand the meaning of a page — not just its string content. They are the supporting vocabulary that defines what conceptual territory a page occupies in Google’s internal representation of the web.
The key distinction: a regular keyword tells Google what words are on your page. A semantic keyword tells Google what your page is about. Those are wildly different things, and Google’s entire NLP stack has been built to bridge that gap.
Think about the word “bank.” Type it into Google and you don’t get results for river banks and financial institutions and airplane banking maneuvers all mixed together. Google knows — with remarkable precision — what you mean by “bank” in context. It does that by building a semantic field from the surrounding query terms and from the content of candidate pages. That’s semantic understanding.
How Google’s NLP Pipeline Actually Processes Your Content
Most semantic SEO advice stops at “use related terms.” That’s not enough to actually understand what you’re optimizing for. Google’s natural language processing doesn’t happen in one step — it’s a sequential pipeline where each stage can strengthen or lose signal from your content.
Stage 1 — Tokenization and Parsing
Before any meaning is extracted, Google’s crawler fetches your page and the NLP system breaks the text into tokens — individual words and sub-word pieces. Then it parses the grammatical structure: what’s the subject, what’s the object, how do clauses relate to each other. This stage is where poorly structured sentences can actually lose semantic signal. Google isn’t human. A tortured, passive-voice sentence that a human reader might understand can produce ambiguous parsing signals.
Stage 2 — Entity Extraction and Resolution
This is where Google identifies the entities in your content. Entities are not keywords. They are specific, identifiable things: people, places, organizations, concepts, products, events. Google can extract an entity mention and then resolve it — linking “the search giant” to the entity Google Inc., for example, even without the company name appearing in that exact phrase.
Resolution connects mentions to entries in the Google Knowledge Graph, a database of hundreds of billions of entity relationships. If your content discusses an entity and Google can confidently resolve it to a Knowledge Graph entry, your page inherits semantic relationships from that graph. You gain adjacency to related entities you may never have mentioned.
Stage 3 — Entity Salience Scoring
Every entity extracted from your content gets a salience score — a measure of how prominent and important it is within that specific document. Salience is not about frequency alone. Position matters: entities mentioned in H1, H2s, and opening paragraphs score higher salience than the same entity buried in paragraph 14. The surrounding co-occurring terms matter. The specificity of attribute coverage matters.
Adding a 500-word section about a secondary topic can actually lower your primary entity’s salience score because you’ve introduced competing entities into the same document pool. This is why topically focused pages consistently outperform sprawling “ultimate guides” that try to cover too many entities at once.
Stage 4 — Sentiment and Context Classification
Google’s NLP models classify content sentiment — not in a crude positive/negative way, but in a nuanced way that understands whether a piece of content is authoritative explanation, promotional copy, user question, opinion, or factual answer. Approximately 87% of top-10 search results carry positive sentiment signals, according to analysis of NLP scoring patterns. That’s not about writing cheerleader content — it’s about writing with the confidence and clarity of an expert rather than hedging every sentence.
Stage 5 — Semantic Indexing by Concept, Not String
By the end of the pipeline, Google is not storing your page as a collection of keywords. It’s storing a conceptual representation — a vector in high-dimensional semantic space — that captures what your page means. When a query comes in, Google’s retrieval systems compute semantic similarity between the query vector and page vectors. This is why a page can rank for searches that contain none of its exact words.
The Models Behind the Curtain: BERT, MUM, and Gemini
Understanding the evolution of Google’s NLP models matters because each generation fundamentally changed what “optimizing for semantic understanding” requires.
BERT (Bidirectional Encoder Representations from Transformers, introduced October 2019) was the genuine inflection point. Before BERT, Google read text in one direction. BERT reads in both directions simultaneously — understanding each word in the context of every other word in the sentence. The famous example: “Can you buy a car without a license?” Before BERT, Google might rank generic car-buying pages. After BERT, it understood that “without a license” was the core semantic element and served legal and regulatory content.
For SEOs, the practical implication of BERT was that negation, prepositions, and small function words suddenly mattered enormously. Words like “not,” “for,” “without,” “before” — which previous systems largely ignored — became load-bearing semantic signals.
MUM added multimodal understanding — Google can now process images, video, and text together — plus cross-language comprehension. Critically, MUM remains primarily deployed for specific features (Google Lens, vaccine-related searches, Related Topics in video) rather than core ranking. For general web ranking, BERT, RankBrain, and Neural Matching remain the primary systems as of 2026.
Gemini 3, deployed in AI Mode since January 2026, is a different beast. It’s the synthesis engine behind AI Overviews — it doesn’t rank pages in the traditional sense, it selects and synthesizes content from a wide pool of sources determined by what Google calls “query fan-out” (breaking the original query into multiple sub-queries, then aggregating sources that appear most authoritatively across all of them). This is where semantic depth becomes survival-critical.
Why Semantic Keywords Are Now Existential, Not Optional
Here’s where it gets uncomfortable.
In January 2025, Google AI Overviews appeared for 6.49% of search queries. By July 2025, that figure peaked at 24.61%. After recalibration it settled at roughly 15–16% in November 2025 — but BrightEdge’s commercial vertical tracking put it at 48% by March 2026, and Xponent21’s measurement put U.S. prevalence at 60.32% by April 2026. The trajectory is one-way.
When an AI Overview appears, organic CTR plummets. Seer Interactive’s September 2025 study — tracking 3,119 informational queries across 42 organizations over 25.1 million organic impressions — found organic CTR dropped 61%, from 1.76% to 0.61%. Paid CTR fell 68%. Traditional news publishers have lost 26–55% of their organic traffic in 12 months as zero-click searches rose from 56% to 69% of all queries.
But here’s the inversion point: the brands cited in AI Overviews earn 35% more organic clicks and 91% more paid clicks than uncited competitors on the same queries. AI-referred visitors convert at 23× the rate of traditional organic visitors, according to Ahrefs data from June 2025.
“The question is no longer ‘can I rank on page one?’ It’s ‘will Gemini cite me when it synthesizes an answer?’ Those are different games. The currency in the second game is semantic completeness, not keyword frequency.”
What determines AI Overview citation? Analysis of 15,847 AIO results found that semantic completeness is the strongest predictor (r=0.87 correlation). Content scoring 8.5/10+ on semantic completeness is 4.2× more likely to be cited. Pages with 15 or more recognized entities show 4.8× higher selection probability. Meanwhile, only 38% of AIO citations now overlap with the traditional top-10 organic results, down from 76% in mid-2025 — meaning ranking first is no longer sufficient for AIO inclusion.
What Semantic Keywords Actually Look Like in Practice
I’m going to use a real example here because the abstract description is too easy to misapply. Let’s say you’re writing a page about electric vehicle charging infrastructure.
The keyword-era approach: repeat “EV charging” as many times as possible, maybe add “electric car charging stations” as a “synonym.” Done.
The semantic approach requires building a topical field. Your primary entity is EV charging infrastructure. The semantic keywords are not alternatives to that phrase — they are the conceptual neighbors that define what the topic means:
| Semantic Layer | Examples for “EV Charging Infrastructure” | Why It Matters to Google |
|---|---|---|
| Entity attributes | charging speed (Level 1/2/DC Fast), kilowatt output, connector types (CCS, CHAdeMO, NACS) | Defines the entity’s specific properties — increases salience depth |
| Related entities | Tesla Supercharger, ChargePoint, Electrify America, NEVI Formula Program | Connects your entity to known Knowledge Graph nodes |
| Contextual concepts | range anxiety, grid capacity, utility rate structures, charging etiquette | Shows topical authority beyond surface coverage |
| User intent phrases | “how long to charge,” “where to find fast chargers,” “cost per kWh” | Matches PAA boxes, featured snippets, conversational queries |
| Co-occurrence patterns | Terms that appear together in authoritative content on this topic across the web | Google’s NLP weights terms by how often they appear alongside your entity in high-authority documents |
Notice what’s not there: a list of synonyms. “Electric vehicle power stations” is not a semantic keyword — it’s a restatement. Semantic keywords are conceptually adjacent, not lexically similar. They expand Google’s understanding of your topical coverage, not its confidence that your page is about the same thing.
The Entity Salience Optimization Model (ESOM)
This is the original framework I’ve developed from working with content across dozens of sites. I’m calling it the Entity Salience Optimization Model because the industry doesn’t have a good shorthand for what I’m about to describe, and “add semantic keywords” has been so watered down that it means nothing useful.
The core insight: every piece of content has a semantic centroid — the primary entity that Google should associate with your page. Everything else in the content either increases or decreases the salience of that centroid. Your job is to engineer the content so that every section, every heading, every supporting concept serves the salience of your primary entity.
Six levers that determine entity salience in your content
about, mentions, and sameAs properties.A Semantic Coverage Scoring System You Can Actually Use
I want to give you something concrete. Here’s the scoring matrix I use when evaluating content for semantic completeness — the same quality that predicts AI Overview citation probability.
This is built on the pattern that semantic completeness correlates at r=0.87 with AI Overview citation, and the observation that content needs to score 8.5/10 or higher to achieve 4.2× citation lift. The five dimensions below contribute to an aggregate semantic completeness score.
A real calculation: a 2,000-word product review that scores Entity Density 6/10, Attribute Depth 8/10, Intent Coverage 5/10, Extractability 7/10, Schema 3/10 — after weighting, that article scores roughly 6.1/10. Borderline for featured snippets. Unlikely to be cited in AI Overviews. Adding schema alone pushes it to 6.9/10. Adding self-contained answer blocks and covering more intent variants gets it to 8.1/10. At that point, AIO citation becomes meaningfully probable.
How to Actually Find Semantic Keywords (Not the Listicle Version)
Most guides tell you to use tools. The tools are useful but secondary. Here’s the hierarchy that actually produces semantic keyword sets worth using:
Method 1 — Google’s Own NLP API (Free, Underused)
Google’s Natural Language API is publicly available at cloud.google.com/natural-language. You can paste the text of any competitor page and see exactly what entities Google extracts, their salience scores, and how they’re categorized. Run your own content through the same API. Compare. The gap between your entity profile and the top-ranked page’s entity profile is your semantic keyword opportunity map.
I’ve done this on 40+ content audits and the pattern is relentlessly consistent: the pages that rank are not the ones with more keywords. They’re the ones where Google extracts more entities with higher average salience.
Method 2 — Query Fan-Out Mapping
Google’s AI Mode uses “query fan-out” — splitting a primary query into multiple sub-queries to find the most authoritative sources across all variants. You can reverse-engineer this. Take your target topic. Brainstorm every sub-query that a more specific version of your audience might ask. The overlapping terms across those sub-queries are, by definition, the terms that appear most often in the fan-out SERP pool — which means they’re the terms most likely to trigger AI citation of your page.
Method 3 — People Also Ask and Related Searches Mining
PAA boxes are Google telling you, in plain language, what entities and concepts it associates with your primary query. Every PAA question is a semantically adjacent concept that Google has determined belongs in the same topical cluster. Covering the key PAA questions within your content doesn’t just get you PAA appearances — it tells Google’s entity resolution system that your page comprehensively covers the topical cluster.
Method 4 — Competitive Entity Differential
Take the top three ranking pages for your target query. Extract their entities using the Natural Language API or a tool like Semrush or Clearscope. Build a master list of all entities present across all three. Find which entities appear in 2 or 3 of the top pages but are absent from your content. Those are your priority semantic keywords — not because tools said they’re “related,” but because Google has already validated their association with this topic by ranking pages that contain them.
Where I’ve Gotten This Wrong (And What It Cost Me)
I spent about eighteen months optimizing for topical coverage by obsessively expanding content. The logic was simple: more semantic keywords covered = more semantic signal. My average page length went from 1,800 words to 3,400 words. Rankings did not improve proportionally. On several pages they declined.
The mistake was conflating topical breadth with entity salience. When you add 1,500 words of tangentially related content to a page, you’re not increasing the salience of your primary entity — you’re diluting it with competing entities. A 1,800-word page tightly focused on one entity with high attribute depth will consistently outperform a 3,400-word page that wanders into adjacent topics.
The fix was to apply what I now call semantic containment discipline: every section of a page must trace back to the primary entity. If a section is primarily about a different entity, that section becomes a separate page with its own internal link. Topic clusters built on this principle — where each page has one clear entity with maximum salience, and internal links handle the semantic adjacency — outperform the “comprehensive single page” approach on entity-salience metrics in every test I’ve run.
The Unpopular Take: Most “Semantic SEO” Content Is Still Keyword Stuffing in Disguise
I want to push back on something the industry consensus is getting wrong.
The migration from “keyword SEO” to “semantic SEO” has largely been a rebrand, not a rethink. Most practitioners have simply replaced “add synonyms” with “add semantic keywords” — a change in vocabulary that doesn’t change the underlying behavior. The content being produced is still organized around keyword frequency, still structured to maximize term coverage, still evaluated by tools that count word appearances.
Google’s entity salience scoring does not reward synonym variation. Swapping “semantic SEO” for “entity-based search optimization” throughout a page does not improve entity salience. What improves entity salience is contextual depth: explaining an entity’s properties, mapping its relationships to adjacent entities, covering the scenarios in which it matters, describing its failure modes. That requires genuine expertise, not keyword engineering.
The uncomfortable implication is that the content that will win AI Overview citations in 2026 is content written by people who actually understand the topic — supported by good structure and schema, yes, but fundamentally grounded in domain expertise. No amount of semantic keyword insertion rescues a shallow article.
“Google is not reading your keyword map. It’s running your text through a pipeline that scores what you know about the topic. Semantic keywords are evidence of knowledge, not a substitute for it.”
The Revenue Math: Why Semantic Completeness Has Become the Highest-ROI SEO Investment
Let me put some numbers to this so you can make a resource allocation decision rather than taking my word for it.
Here’s the arithmetic with real inputs. Take a page ranking #3 for a query with 10,000 monthly searches. Without AI Overviews, it earns roughly 10% CTR = 1,000 visits/month. With an AI Overview present but without citation: CTR drops 61% to about 3.9% = 390 visits. With AI Overview citation: you recover 35% of the lost ground = 527 visits, plus your paid CTR improves 91% if you’re also running ads on that query.
Now factor in the conversion premium: Ahrefs’ June 2025 data showed AI-referred visitors converting at 23× the rate of traditional organic. 0.5% of total traffic from AI channels generated 12.1% of signups in their 30-day study. The visitors you do get from semantic-rich, AI-cited content are dramatically higher value than the volume-traffic you were getting from keyword-stuffed content.
The summary: semantic completeness investment isn’t just defensive (maintaining traffic against AI Overview CTR erosion). It’s offensive — the visitors who arrive through AI-cited content convert at rates that make them individually more valuable than organic volume visitors by an order of magnitude.
A Practical Semantic Keyword Workflow for Content Teams
Here’s how I’d operationalize the above for a content team producing multiple pieces per week:
From brief to entity-optimized draft in five steps
The Knowledge Graph: The Entity Database Google Uses to Validate Your Semantic Claims
A concept that doesn’t get enough attention in semantic keyword discussions: the Google Knowledge Graph is not just a display feature (those info panels on the right side of search results). It’s the validation layer that Google’s NLP uses to assess whether your entity mentions are resolvable and credible.
When Google’s NLP extracts an entity from your content and attempts resolution — linking your mention to a Knowledge Graph node — it checks whether the surrounding content is consistent with what it knows about that entity. If you write about “BERT” in the context of NLP and your entity profile matches the Knowledge Graph’s representation of BERT as a Google language model, resolution succeeds and your page inherits Knowledge Graph associations. If your content is thin or contradictory, resolution may fail or score low confidence.
The practical implication: adding sameAs links to authoritative external sources (Wikipedia entries for entities, official product pages, academic papers) in your JSON-LD schema helps Google resolve your entities with higher confidence. You’re essentially saying “when I mention BERT, I mean this specific, known entity” rather than leaving resolution to probabilistic inference.
What Changes With Gemini 3 and AI Mode (The 2026 Update)
Since January 2026, Google has been rolling out AI Mode globally — a search experience powered by Gemini 3 Flash that generates dynamic answers with interactive elements directly in the SERP. This isn’t AI Overviews plus. It’s a different architecture.
The citation dynamics have shifted. Ahrefs’ March 2026 data shows that only 38% of AI Overview citations now come from the visible top-10 organic results, down from 76% in mid-2025. Google confirmed the mechanism: AI Mode uses “query fan-out” — splitting queries into multiple sub-queries and aggregating sources from across all the resulting SERPs. A page doesn’t need to rank #1 for the primary query. It needs to rank authoritatively for multiple related sub-queries — which is exactly what strong semantic coverage achieves.
In practical terms: a page with high entity salience and strong semantic completeness that ranks #8 for the primary query might rank #2 for three related sub-queries. Under the query fan-out system, it accumulates citation probability across all four signals and may end up cited more reliably than the page that ranks #1 for the primary query but doesn’t cover the semantic field broadly enough to appear in sub-query results.
This changes the ROI calculation for semantic depth investment in a significant way: it’s no longer just about ranking for your target keyword. It’s about being the most authoritative document in the conceptual neighborhood — because Gemini is scouting the whole neighborhood, not just the front door.
Frequently Asked Questions
Semantic keywords are contextually relevant terms, entities, and concepts that help search engines understand the full topical meaning of a piece of content — not just its exact string content. They include related entities, attribute descriptors, co-occurring concepts, and intent-matching phrases that define what a page is about in semantic space.
No. Google’s John Mueller confirmed in 2019 that there is no such thing as LSI keywords in Google’s ranking system. LSI (Latent Semantic Indexing) is a 1980s Bell Labs technology built for static document databases. Google uses neural NLP models — including BERT, MUM, and Gemini — to understand semantic meaning through entity extraction, salience scoring, and vector-based conceptual indexing.
The framing is wrong: semantic keywords are not a quantity target. The goal is entity completeness — covering the primary entity’s attributes, relationships, and contextual scenarios in sufficient depth. Analysis of AI Overview citation patterns suggests pages with 15 or more recognized entities have 4.8× higher citation rates. But those entities need to be genuinely relevant and covered with depth, not mentioned as keywords.
Regular keywords are exact query strings. Semantic keywords are conceptually adjacent entities and terms that establish what a page means, not just what words it contains. Google can now rank a page for queries that never appear in the text if the semantic field is sufficiently strong — a capability that was not possible in the keyword-matching era and that makes semantic completeness far more valuable than keyword frequency.
Semantic completeness is the strongest predictor of AI Overview citation (r=0.87 correlation, per analysis of 15,847 AIO results). High semantic completeness signals that a page provides a full, self-contained answer for the topic — which is exactly what Google’s Gemini system prioritizes when selecting sources for AI Overviews. Pages with strong entity coverage also benefit from query fan-out: appearing in multiple related sub-query results increases cumulative citation probability even for pages that don’t rank #1 for the primary query.
The most direct method is Google’s own Natural Language API (free tier available), which shows you exactly what entities Google extracts from any page and their salience scores. For competitive analysis, Semrush’s Keyword Magic Tool, Clearscope, and MarketMuse surface co-occurrence patterns. Surfer SEO’s SERP analyzer compares your entity coverage against top-ranking pages. The critical point: use these tools to understand entity gaps, not to generate synonym lists.
The Bottom Line
Semantic keywords are not a trick. They’re not a synonym list. They’re not a box you check by running a piece of content through a keyword tool that generates “related terms.”
They are the evidence that you actually know the subject you’re writing about — presented in a structure that Google’s NLP pipeline can parse, score, and index at the conceptual level. In an era where AI Overviews are intercepting 15–60% of search queries and converting clicks into zero-click answers, the content that survives and thrives is the content that Google’s Gemini system trusts enough to cite.
That trust is earned through entity salience, attribute depth, Knowledge Graph adjacency, and structured data — not through keyword frequency or synonym variation.
The hard truth is that most content on the web will not earn that citation. The CTR erosion data from Seer Interactive — 61% drop in organic CTR when AI Overviews appear — is a permanent structural shift in the search economics. The way through it is not louder SEO. It’s deeper, more genuinely expert content that uses semantic signals to tell Google’s pipeline exactly what it needs to know.
Start with the Natural Language API. Audit your primary entity’s salience score. Find the entity gaps between your page and the top-ranking competitors. Fill them with genuine depth. Run it through the ESOM matrix. That’s the game now.
