Most guides on keyword research for SEO are built on practitioner intuition, tool vendor documentation, or advice recycled from articles a decade old. This one is different. Over the past ten years, more than a hundred peer-reviewed studies have examined how people form search queries, which terms actually drive traffic and conversions, and how language models can extend or distort keyword lists. The research spans top-tier computer science conferences — ACL, EMNLP, SIGIR, WWW — information science journals, and marketing publications. Taken together, it converges on conclusions that contradict several things the SEO industry still takes for granted.
That disconnect matters especially for businesses navigating an accelerating transition. Every best SEO agency in Canada is fielding more questions than ever about AI-powered search, semantic keyword strategy, and what the shift from blue-link rankings to generated answers means for organic visibility. The evidence base to answer those questions has been building in academic literature for a decade — but it rarely surfaces in practitioner guides. This article closes that gap.
What follows is a structured synthesis of the highest-priority findings from a curated database of 100 academic papers published between 2016 and 2026, covering four research areas: SEO keyword selection and search marketing; query expansion and semantic retrieval; keyphrase extraction and topic discovery; and search intent modelling through autocomplete and query suggestion. Each section moves from what the research shows to what it means for your day-to-day workflow — with specific methods, examples, and caveats drawn directly from the studies, not extrapolated from them.
The guide is particularly relevant for those evaluating or delivering AI SEO services for Canadian businesses in 2026 — a market where LLM-powered keyword tools are being adopted rapidly but where the academic evidence about their limitations and correct usage is still not widely understood. The aim is to give practitioners a research-grounded framework they can use immediately, and to flag specifically where commonly accepted SEO practices conflict with what the evidence actually shows.
The most foundational and highest-scoring paper in the entire database (Priority Score: 100) is Nagpal and Petersen’s 2021 study published in the Journal of Retailing: “Keyword Selection Strategies in Search Engine Optimization: How Relevant is Relevance?” Its central finding is both simple and persistently ignored: sorting keywords by search volume alone produces a suboptimal shortlist.
The reason is that keyword popularity interacts with at least four other variables before it translates into actual traffic value: competition intensity, keyword specificity, search intent alignment, and the existing page or domain authority of the site competing for the query. A high-volume keyword on a low-authority domain competing against entrenched SERP incumbents is not a better opportunity than a lower-volume keyword where specificity is high, the competition set is thin, and the search intent matches precisely what the page already delivers. The interaction effects between these variables are substantial enough to flip the ranking of opportunities entirely when you move from a single-factor to a multi-factor model.
The implication is structural. Keyword prioritisation needs a composite score, not a single sorted column. A practical opportunity score following Nagpal and Petersen’s variable framework looks like this:
Opportunity Score = Demand × (1 / Competition Difficulty) × Specificity Signal × Intent Match × Page/Domain Strength
None of these variables is difficult to estimate in practice. Demand comes from volume tools. Competition difficulty is available from Ahrefs, Semrush, or Moz keyword difficulty scores, or from manual SERP inspection — counting the number of high-authority domains in the top five positions gives a reasonable proxy. Specificity signal is approximated by word count and modifier density. Intent match requires reading the top-ranking SERP results and asking whether your page’s purpose aligns with what those results deliver. Page and domain strength is your existing authority in the topic area, estimated by comparing your domain rating to those of the incumbents.
Scoring on even a rough 1–5 scale for each dimension, then multiplying the components, produces a ranking that reflects actual opportunity far better than volume alone — and it surfaces low-volume, high-specificity terms that volume-only sorting buries.
Ramaboa and Fish’s 2018 work in Information Processing & Management added an important mechanism to the specificity variable: keyword length itself functions as an intent proxy. Longer, more specific queries — what the industry conventionally calls long-tail keywords — are statistically associated with more focused search intent and more predictable user behaviour. This is not practitioner intuition. The study showed that word count and the breadth of the match type together classify queries by funnel depth, with shorter head terms carrying more navigational and informational ambiguity, and longer specific phrases concentrating commercial and transactional intent.
The practical application is direct. A query like “hosting” is a three-character head term with maximum ambiguity — the searcher could be looking for a definition, a price comparison, a provider review, or a tutorial. A query like “best managed WordPress hosting for WooCommerce Canada” is an eight-word long-tail phrase with a heavily concentrated transactional and commercial intent. The content, the depth, and the call to action that serves one of those queries well will often actively fail the other.
Yang, Pancras and Song’s 2021 Decision Support Systems study on broad versus exact match in paid search reinforced the same dynamic: the performance difference between match types depends fundamentally on keyword specificity, not just on budget or bid level. The insight transfers directly to organic SEO content strategy. When you write content targeting a head term, you are competing across a query space that includes multiple intent types, some of which your page may not serve. When you target a specific long-tail variant, the intent signal is stronger, the competition set is usually narrower, and a well-matched page has a meaningfully higher probability of both ranking and converting.
Yang, Jansen, Yang, Guo and Zeng’s 2019 IEEE Intelligent Systems paper proposed a multi-level computational framework for keyword optimisation that treats targeting, assignment, and grouping as a joint problem rather than three separate sequential decisions. The finding — that the value of a keyword depends on which targeting rule and assignment structure is applied to it — maps directly to a common SEO mistake: selecting keywords and then deciding later which pages they go on, rather than designing keyword-page assignments as a unified system.
In practice, this means the question “which page should own this keyword?” should be answered during the selection stage, not after. A keyword that belongs on a blog post and a keyword that belongs on a product landing page have different opportunity metrics even if their volume and difficulty scores are identical — because the page type constrains conversion path, CTA structure, and the authority signals the page can accumulate.
One of the most underabsorbed findings in the database comes from Symitsi, Markellos and Mantrala’s 2022 study in the European Journal of Operational Research: keyword sets should be managed as portfolios, not collections of independent targets. The study applies financial portfolio theory — specifically the principle of diversification — to keyword selection. It shows that combining keywords with different trend trajectories, competition profiles, and return characteristics improves risk-adjusted performance relative to concentrating on any single keyword class.
The analogy to financial diversification holds in a practical sense. Head terms (high demand, high competition, slow ROI) behave like growth stocks: high potential but long lead times, volatile to algorithm changes, and concentrated in a few high-value pages. Long-tail terms (lower individual demand, faster ranking, stronger intent, often better conversion) behave more like fixed income: lower individual return but faster realisation, more stable, and diversifiable across many pages. Mid-tail modifier terms occupy the middle: moderate competition, meaningful specificity, useful for connecting the head and long-tail layers into a coherent topic cluster.
A keyword strategy built entirely on high-volume head terms is unnecessarily concentrated. A strategy built entirely on long-tail specifics misses the brand visibility and topical authority signals that head terms provide. The evidence supports holding all three layers — not for intuitive balance, but because they respond differently to algorithm changes, seasonal demand shifts, and competitor activity, which is the precise definition of a diversification benefit.
Erdmann, Arilla and Ponzoa’s 2022 Journal of Business Research study extends the portfolio model into a second dimension that most SEO workflows ignore entirely: time horizon. Keyword choice is not a one-period decision with a single expected payoff. The study demonstrates empirically that branded and generic keywords have different long-run economic trajectories and should be managed as separate sub-portfolios with distinct performance expectations and distinct measurement timelines.
Branded terms typically convert better in the short term, build faster, and protect existing customers — but they require ongoing authority maintenance and do not compound indefinitely because their demand ceiling is bounded by brand awareness. Generic terms are slower to rank and slower to convert, but they compound over time: a well-ranked generic page continues to accumulate traffic and backlinks for years with minimal ongoing investment. Evaluating both types by the same three-month ranking timeline systematically undervalues generic keyword investment and causes teams to abandon generic content too early.
Two papers in the SEO selection cluster address the gap between what external keyword tools surface and the language your actual customers use. They suggest the gap is larger than most practitioners suspect — and that the solution is already available in data most teams already have.
Qiao, Zhang, Wei and Chen’s 2017 paper in Information & Management built a topic-based competitive keyword discovery method from query logs. The core finding: relevance matching alone — the model underlying most keyword tools — misses keywords that compete for the same search demand through different topic pathways. Query logs reveal semantic neighbourhoods that frequency-based tools ignore, because those tools look at individual term co-occurrence patterns rather than at the topical structure of how users navigate a query space.
For practitioners without access to raw query logs, Google Search Console is the closest practical equivalent. The key filter is not the keywords with the most clicks. It is the keywords with meaningful click counts (three or more clicks) but a below-average click-through rate. These are terms where you are already competing — Google has ranked your page for them — but where your title and meta description are underperforming the user’s intent. They represent demand you have already partly won and have a high probability of winning more completely with targeted content improvement. That signal is more reliable than any third-party volume estimate because it reflects actual search engine behaviour on your specific pages.
Scholz, Brenner and Hinz’s 2019 Decision Support Systems study demonstrated a second frequently overlooked source: internal site-search data. Terms typed into a site’s own search box represent what real, already-engaged visitors wanted but could not immediately find on the pages they had visited. This is precisely the vocabulary that would have converted if it were in the page copy, or that would drive organic traffic if it were targeted by a content page. Mining internal search logs, support ticket language, live-chat transcripts, and email enquiries for query patterns consistently uncovers commercially useful terms that third-party keyword databases never carry — because those databases are built from aggregate search behaviour, not from the specific intent of your actual customer base.
A practical workflow combining these two sources: pull your Search Console queries filtered by impression count above a threshold and click-through rate below average; pull your internal site-search terms by query frequency; merge and deduplicate the two lists; and run the merged list through a standard opportunity-scoring step before adding terms to your content brief. This two-source process consistently identifies ten to thirty percent more actionable keywords than starting from a seed-and-expand model alone.
The largest cluster in the database — 45 papers — covers query expansion and semantic retrieval. This is where research activity has been most concentrated since 2020, and it is where the underlying technology has changed most dramatically. Understanding that shift is essential for any practitioner deciding when and how to use AI tools in a keyword research workflow.
Before large language models, query expansion meant statistical co-occurrence methods: find terms that appear frequently alongside the seed query in a large text corpus, score them by relevance using weighted term frequency or pointwise mutual information, and add the strongest to the keyword list. This approach is not fundamentally wrong — the survey by Azad and Deepak (2019, Information Processing & Management, Priority Score: 92) catalogued the family of statistical methods and showed they provide consistent retrieval improvements on standard benchmarks. But these methods are bounded by the corpus they are trained on. If a term is not well-represented in the co-occurrence data, it will not be suggested regardless of whether it reflects real search demand.
LLMs changed that ceiling. They can generate semantically plausible vocabulary for a topic from parametric knowledge — information encoded in model weights from training — without any corpus retrieval step at all. This produces substantially broader candidate sets, but it introduces a different failure mode: the model’s vocabulary reflects what it was trained on, not what searchers are currently typing. Managing that trade-off is what the 2023–2026 research is primarily about.
This shift is directly relevant to anyone delivering or evaluating AI SEO services for Canadian businesses in 2026. Canadian search behaviour carries specific characteristics — bilingual markets, provincial regulatory vocabulary, distinct ecommerce patterns, and a concentration of search traffic on Google.ca — that generic LLM training corpora represent unevenly. Understanding where LLM-generated keyword vocabulary is reliable and where it requires domain-specific grounding is therefore not an abstract academic concern; it has direct consequences for keyword plans built for Canadian audiences.
Wang, Yang and Wei’s 2023 EMNLP paper “Query2doc: Query Expansion with Large Language Models” (Priority Score: 99) is the single most practically important LLM paper in the database for SEO keyword work, even though it was designed for information retrieval systems rather than for SEO directly. The core idea is straightforward: instead of searching for co-occurring terms, ask an LLM to write a short pseudo-document — a plausible answer or brief article — for each seed keyword, then mine that generated document for the entities, attributes, questions, and subtopics it naturally uses.
The mechanism is important to understand. When a language model generates a pseudo-answer to “how to choose managed WordPress hosting for an ecommerce site,” the generated text does not simply list synonyms of the seed keyword. It produces the vocabulary a knowledgeable person would naturally use to discuss the topic: uptime guarantees, server-side caching, automatic core updates, PHP version management, staging environments, WooCommerce compatibility, developer tool access, and so on. Each of those concepts corresponds to a cluster of real search queries — “WordPress hosting with staging environment,” “managed hosting PHP version control,” “WooCommerce hosting uptime SLA.” That vocabulary is your semantic keyword map.
Query2doc experiments showed statistically significant BM25 retrieval gains across multiple benchmark datasets. For SEO, the workflow is: generate the pseudo-document; extract the noun phrases, entities, and technical terms; cross-check each extracted term against real search-volume data; promote confirmed terms to your keyword plan; and park unconfirmed terms in a demand-monitoring folder rather than discarding them — some of them will develop search volume as the topic matures.
Lei, Shen and Yates’ 2025 EMNLP paper “ThinkQE: Query Expansion via an Evolving Thinking Process” (Priority Score: 96) formalised a refinement on the single-pass LLM expansion model. Rather than generating all expansion terms in a single prompt, ThinkQE iterates: generate an initial set, retrieve evidence using those terms, identify facets the initial expansion failed to cover, and then expand again with the newly identified gaps as the prompt context.
For practical keyword research, this iterative model is more reliable than single-pass generation because it incorporates real retrieval feedback — in SEO terms, actual SERP results — at each iteration. A practical implementation: run the initial LLM expansion, search for each generated cluster on Google, review the top results for topic facets not yet in your keyword plan, and then run a second expansion pass specifically targeting those gaps. This process typically surfaces two to four additional topic clusters that the first pass missed, and it grounds the second expansion on real SERP evidence rather than model assumptions.
Jagerman, Zhuang, Qin, Wang and Bendersky’s 2023 Gen-IR paper “Query Expansion by Prompting Large Language Models” (Priority Score: 98) found that the quality of LLM-generated expansion terms is highly sensitive to prompt structure. Reasoning-oriented prompting — where the model is asked to think through the question before generating terms — consistently outperformed direct expansion prompting across retrieval benchmarks. The performance difference between prompt styles was large enough to dominate the choice of model architecture.
For SEO keyword research, prompting for specific facets outperforms asking broadly for “related keywords.” Effective prompt facets, drawn from the paper’s analysis, include:
Each of these facets produces a different slice of the keyword space, and combining them across a structured prompt sequence gives broader, more organised coverage than any single broad request. The output of a well-structured facet-based expansion is a candidate set naturally organised by intent type — which feeds directly into the clustering stage of the keyword workflow.
The research literature that emerged in 2025 introduced two documented failure modes that every practitioner using LLMs for keyword work needs to understand.
Yoon, Jung, Yoon and Park’s 2025 paper “Hypothetical Documents or Knowledge Leakage? Rethinking LLM-based Query Expansion” (Priority Score: 95) demonstrated that apparent retrieval gains from LLM expansion can partly reflect memorised benchmark or domain data rather than genuine understanding of the query context. A language model trained on large web corpora has seen most popular web topics extensively. When asked to expand a query about a well-documented topic, it produces fluent, confident-sounding terms — but that fluency is evidence of training exposure, not evidence of search demand. The model’s output looks useful because it is topically coherent, not because the terms it suggests are actually searched.
The parallel failure mode is documented in Abe, Takeoka, Kato and Oyamada’s 2025 paper “LLM-Based Query Expansion Fails for Unfamiliar and Ambiguous Queries” (Priority Score: 95): when the seed keyword is outside the model’s training distribution — niche industry jargon, newly emerged product categories, ambiguous acronyms, regional terminology — LLM expansion actively degrades retrieval performance. Rather than generating useful semantic neighbours, the model over-commits to one plausible interpretation of an ambiguous term and narrows the expansion space in ways that exclude genuinely relevant queries.
These two failure modes define the boundary conditions for LLM-based keyword expansion. The method works well for well-documented topics where the model has genuine parametric knowledge and the seed terms are unambiguous. It fails predictably for emerging topics, niche vocabularies, and ambiguous terms. The practical rule is non-negotiable: treat every AI-generated keyword as a hypothesis, not a recommendation. Validate demand before acting. Do not infer search popularity from model confidence or output length.
Lei, Cao, Zhou, Shen and Yates’ 2024 “Corpus-Steered Query Expansion with Large Language Models” (Priority Score: 97) identified the root cause of the generic-suggestions problem and proposed a principled solution. The problem is that unprompted LLMs draw on the full breadth of their training distribution. Asked to expand “managed WordPress hosting,” a model without corpus context generates terms common across all hosting topics — performance, security, uptime — rather than terms specific to managed WordPress as a product category and the vocabulary that distinguishes it from shared or VPS hosting.
The solution is a retrieval step before the expansion prompt: identify the most relevant and topically pivotal sentences from a domain-specific corpus (competitor pages, your own content, Search Console queries, support documentation), and include those sentences as context in the expansion prompt. This grounds the model in niche vocabulary and prevents it from retreating to generic topic associations. Corpus-steered expansion consistently outperformed unprompted LLM generation in the study’s retrieval experiments, with the largest gains precisely on the niche topics where unprompted generation is weakest.
The SEO implementation is straightforward. Before running an LLM keyword expansion, paste in: the three highest-ranking competitor pages for your target keyword (as text); your own best-performing existing content on the topic; and the twenty to thirty most common queries from Search Console related to the seed term. Then prompt for expansion within that context. The resulting suggestions will reflect the specific vocabulary your market uses rather than the generic vocabulary of the broader topic area.
Chen, Chen, He, Wen and Sun’s 2024 Findings of ACL paper “Analyze, Generate and Refine: Query Expansion with LLMs for Zero-Shot Open-Domain QA” (Priority Score: 97) formalised the most complete practitioner-ready LLM expansion workflow in the database. The three-stage process consistently outperformed single-pass LLM expansion in the paper’s zero-shot retrieval experiments, and it maps directly into an SEO keyword research context:
Stage 1 — Analyze: Review the seed keyword and available SERP data. Identify which topic facets are currently underserved in the top-ranking results. Make an explicit list: what questions do searchers for this term have that the current top results do not answer? What related subtopics appear in related searches but are absent from the main content? This analysis step prevents the next stage from generating more of what already exists.
Stage 2 — Generate: Run a corpus-steered LLM expansion (as described above) with the gap analysis from Stage 1 as additional prompt context. At this stage, optimise for breadth: generate a large candidate set covering all identified facets, including terms you are uncertain about. Quantity over precision here — you will filter in the next stage.
Stage 3 — Refine: Filter the Stage 2 candidate set against real demand data. Check each candidate in Search Console, a volume tool, or autocomplete. Keep terms that show confirmed demand above a minimum threshold. Keep terms with no current demand but clear content logic in a monitoring list — if demand develops, you have already mapped the content opportunity. Discard terms where neither demand nor content logic justifies inclusion.
This pipeline eliminates the two main failure modes of single-pass LLM expansion: it prevents the analysis gap (generating more of what already exists by not examining the SERP first) and prevents the demand gap (keeping terms with no search evidence by validating before acting).
Jia, Liu, Zhao, Li, Hao, Wang and Yin’s 2024 NAACL paper “MILL: Mutual Verification with Large Language Models for Zero-Shot Query Expansion” (Priority Score: 94) added a validation layer that is particularly useful for practitioners who do not have access to high-quality volume tools. The mutual-verification approach cross-checks LLM-generated expansion terms against retrieval evidence — essentially asking: when you search for this generated term, do the results look like content relevant to the original seed query? Terms that generate irrelevant retrieval results are discarded; terms that generate relevant results are retained.
For SEO practitioners, this translates to a quick manual step: search each generated candidate on Google and review the top three results. If the results are clearly relevant to the topic area — they discuss the same customer problem, product category, or use case as the seed keyword — the candidate is confirmed. If the results show an unrelated topic or a completely different intent type, the candidate is discarded regardless of how confident the LLM sounded when generating it.
While the query expansion literature focuses on generating new keyword candidates from a seed, the keyphrase extraction cluster — 20 papers published between 2017 and 2025 — addresses a different but equally valuable problem: automatically identifying the important terms and phrases already present in competitor pages, SERP snippets, and your own content corpus.
The fundamental insight from this body of work is that your competitors have already done a significant part of your keyword research. The pages ranking in positions one through ten for any target keyword collectively contain the vocabulary that search engines have rewarded as relevant, comprehensive, and authoritative on that topic. Systematic extraction of that vocabulary, organised by frequency and prominence, gives you a research-grounded vocabulary inventory with substantially less work than building a keyword list from scratch.
The most practically accessible method in this cluster is YAKE (Yet Another Keyword Extractor), introduced by Campos, Mangaravite, Pasquali, Jorge, Nunes and Jatowt (Priority Score: 97, Information Sciences). YAKE ranks keywords from a single document using five local statistical features: term frequency normalised by document position, casing (terms that appear in mixed or uppercase often signal named entities or technical vocabulary), position within the document (early-appearing terms receive higher weight), co-occurrence with other candidate terms in a local context window, and the frequency of sentences that contain the candidate term.
Critically, YAKE requires no training data, no external model, and no API calls. It runs on a single web page in milliseconds. This makes it practical for competitive analysis at scale: download the HTML of the top ten ranking pages for your target keyword, strip to body text, run YAKE on each document, and collect the resulting term lists.
The analytical step is the intersection and union of those lists. Terms that appear in YAKE’s top twenty for every competitor document are the vocabulary of comprehensive coverage — the minimum terms your own content needs to address. Terms that appear in only one or two competitor documents represent differentiation opportunities: topics adjacent competitors have not yet converged on. Terms that appear in no competitor document but are mentioned in related searches or autocomplete represent content gaps that no one in the SERP is yet filling — higher risk but potentially high reward.
The 2025 COLING paper by Galletti, Prevedello, Brugnoli, Lo Sardo and Gravino — “Are Your Keywords Like My Queries? A Corpus-Wide Evaluation of Keyword Extractors with Real Searches” (Priority Score: 97) — raised a finding with direct implications for every practitioner who relies on keyphrase extraction for SEO purposes: keyphrases extracted from documents do not reliably match the phrases people actually type into search engines.
The paper’s methodological innovation was to evaluate keyphrase extractors not against author-provided keyword labels (the standard NLP benchmark) but against real Google Trends searches on the same topics. This is precisely the evaluation question that matters for SEO — not “did the extractor identify what the author intended?” but “did the extractor identify what searchers actually type?” KeyBERT, which uses BERT embeddings to rank candidate phrases by semantic similarity to the document, performed strongest on this real-search benchmark. The simpler RAKE (Rapid Automatic Keyword Extraction) algorithm was a surprisingly competitive baseline despite its simplicity.
The implication for SEO is a required validation step that most practitioners skip. Running a keyphrase extractor on competitor pages produces a vocabulary inventory of terms that appear in the documents and are topically salient — but those extracted phrases need to be checked against actual search demand before being promoted to SEO targets. An extracted phrase like “robust infrastructure supporting high-demand concurrent sessions” may appear prominently in a competitor’s technical documentation but generate essentially no organic search volume; the corresponding search-phrased equivalent “hosting for high-traffic websites” may have substantial volume and commercial intent. The extraction gives you the topic space; the demand validation gives you the search-ready phrase.
For teams with access to pre-trained language models — either locally or via API — SIFRank (Sun, Qiu, Zheng, Wang and Zhang, 2020, IEEE Access, Priority Score: 96) and EmbedRank (Bennani-Smires, Musat, Hossmann, Baeriswyl and Jaggi, 2018, CoNLL, Priority Score: 95) offer semantically richer extraction than frequency-based methods like YAKE or RAKE.
SIFRank computes sentence and phrase embeddings using a pre-trained language model, then ranks candidate keyphrases by their semantic similarity to the document’s overall embedding — essentially asking: which phrases are most representative of what this document is collectively about, rather than merely what it mentions most often? This method surfaces conceptually central phrases that frequency methods miss. A page about WordPress performance optimisation may mention “caching” only twice but discuss it in a way that is semantically central to the document’s purpose; SIFRank would surface it; YAKE might not.
EmbedRank extends this by adding a diversity mechanism through maximal marginal relevance. After selecting the highest-scoring candidate, the algorithm penalises subsequent candidates that are semantically similar to already-selected ones, producing a set of keyphrases that covers multiple distinct facets of the document rather than clustering around a single dominant concept. For SEO topic clustering specifically, this diversity property is valuable: you want cluster keywords that cover different subtopics and user intents within a theme, not five near-synonyms of the same phrase.
A practical extraction workflow for a competitive analysis:
Florescu and Caragea’s 2017 ACL paper PositionRank (Priority Score: 94) established empirically that phrase position in a document is a significant signal for keyphrase importance, beyond frequency. Terms appearing in the title, opening paragraph, section headings, and the first sentence of each section are consistently more likely to be genuine keyphrases than terms appearing only in the body. The mechanism is intuitive: well-structured content introduces its important topics early and in structurally prominent positions, which is the same signal that makes positional weight a reasonable proxy for topic salience.
Boudin’s 2018 NAACL-HLT paper on multipartite graph keyphrase extraction (Priority Score: 92) extended this with an explicit diversity-through-topics model. By representing a document as a graph where topics are distinct partitions and candidate keyphrases are nodes within those partitions, the method explicitly promotes diversity across topic clusters rather than concentrating all selected keyphrases on the single most salient topic. The result is an extraction set that better maps to the full topic space of a document, which for SEO purposes translates to a more complete vocabulary of subtopics the content covers.
The combined SEO implication from both papers: when extracting keyphrases from competitor long-form content, weight terms appearing in titles, H2s, H3s, and first sentences of sections more heavily than terms appearing only in paragraph bodies. Conversely, when writing your own content, ensure the terms you intend to rank for appear in structurally prominent positions — not because of arbitrary SEO rules, but because positional prominence is a genuine signal of topic importance that both readers and automated extraction systems use to identify what a document is fundamentally about.
| Method | Setup required | Best for | Main limitation |
|---|---|---|---|
| YAKE | None | Fast corpus-wide extraction; competitive analysis across many pages | May surface technical jargon over search-phrased variants |
| RAKE | Minimal | Speed and surprising search-query alignment (per Galletti 2025) | Limited semantic understanding; favours noun phrase frequency |
| KeyBERT | Pre-trained BERT model | Best alignment with actual search phrasing; semantic relevance | Requires model access; slower at scale |
| SIFRank | Pre-trained LM | Conceptually central phrases; finds important terms that are infrequent | Higher computational cost; over-weights document-level themes |
| EmbedRank | Pre-trained LM | Topic-diverse extraction for cluster planning | Can lose precision by over-penalising semantically similar terms |
The 20 papers covering search intent, autocomplete, and query suggestion form the research area most commonly oversimplified in SEO content. The standard intent taxonomy — navigational, informational, commercial, transactional — is a useful starting framework. But the research shows it substantially misrepresents how real search intent operates, and that misrepresentation leads to concrete content failures.
Cai and de Rijke’s 2016 foundational survey of query autocomplete (Foundations and Trends in Information Retrieval, Priority Score: 95) documented that search intent is inherently contextual, temporal, and personal. The same query typed by two different users in two different sessions can reflect genuinely different underlying needs. A search for “WordPress hosting” by a developer comparing technical specifications for a new client project is not the same intent as the same search by a small business owner who does not know what WordPress is. Intent is not a fixed property of a keyword — it is a probability distribution over user goals that shifts based on the searcher’s session history, device, location, time of day, and previous queries in the session.
The four-box taxonomy pretends intent is static. The research shows it is dynamic. This matters practically because content designed for a single static intent label will underserve the portion of the query’s user population that carries a different intent — even if the page ranks well.
The autocomplete literature makes a strong, consistent case for treating query suggestions as primary research data — not as a supplementary brainstorming source to be used after “real” keyword research is complete.
Cai, Liang and de Rijke’s 2016 IEEE Transactions on Knowledge and Data Engineering paper (Priority Score: 95) demonstrated that autocomplete quality improves substantially when personalization and time-sensitivity signals are incorporated. The practical implication for practitioners is significant: a one-time scrape of Google autocomplete suggestions in a single browser, one location, one time of day, gives you a narrow, unrepresentative slice of the actual suggestion distribution. Personalization filters mean your autocomplete reflects your own search history. Time sensitivity means today’s suggestions differ from next month’s for seasonal or trending topics. Location sensitivity means suggestions differ between markets.
The research recommendation translates to a collection protocol: gather autocomplete data across at least three different time windows (including times of seasonal interest), across incognito and standard browser sessions, and across at least the major geographic markets you target. The union of those collection passes gives a substantially more complete picture of the actual query space than any single session.
Fiorini and Lu’s 2018 NAACL paper on personalised neural language models for autocomplete (Priority Score: 95) showed that modern autocomplete can generate suggestions for queries that have never been searched before — using subword and semantic modelling to extrapolate into low-volume and previously-unseen formulations. For keyword research, the practical consequence is that autocomplete surfaces not only confirmed high-volume terms but also plausible emerging variants that have not yet accumulated the search history required to appear in volume-based keyword tools. This is an early-demand signal, not a confirmed demand signal — but treated appropriately, it is genuinely useful for identifying where demand is developing before competitors do.
Jiang, Chen and Cai’s 2017 paper in Mathematical Problems in Engineering (Priority Score: 93) demonstrated that query autocomplete exhibits both periodicity and burstiness. Periodicity refers to seasonal patterns — the same modifier terms appearing in autocomplete in predictable annual cycles. Burstiness refers to sudden interest spikes that emerge rapidly and then may or may not sustain. Critically, both patterns are detectable in autocomplete data before they appear as significant volume in historical keyword tools.
This converts autocomplete from a keyword discovery tool into an early-warning system for emerging demand. Monitoring autocomplete for a set of core seed terms on a regular basis — every four to six weeks — surfaces new modifiers, emerging questions, and developing subtopics while there is still time to create content ahead of the interest peak rather than in response to it. The lead time advantage can be meaningful: autocomplete reflects query behaviour in near-real time, while volume tools typically require three to six months of search history before a term shows sufficient data to appear in their databases.
Wang, Ouyang, Deng and Chang’s 2017 IEEE Transactions on Knowledge and Data Engineering paper (Priority Score: 94) on online learning for adaptive autocomplete formalised the mechanism: the most current autocomplete systems update suggestions based on recent user feedback in near-real time, meaning they reflect current and emerging query behaviour more accurately than any historical database. For practitioners, this means fresh autocomplete collection is not redundant with monthly volume checks — it is complementary, and it captures different information.
Chen, Cai, Chen and de Rijke’s 2020 Information Processing & Management paper on hierarchical query suggestion (Priority Score: 95) showed that individual search sessions have predictable structure. A user’s queries within a single session follow patterns of refinement, elaboration, and intent shift as their understanding develops and their information need becomes clearer. They begin with a broad term, encounter results that refine their understanding, and progressively move toward more specific formulations that reflect a sharper, more informed intent.
This session structure has a direct and underappreciated content implication: the keyword cluster for a topic should map to a user’s natural query progression through that topic, not merely to lexical variations of a single phrase. A cluster built around “managed WordPress hosting” needs to include not only synonyms and modifiers of that exact phrase, but the upstream awareness queries (“what is managed WordPress hosting,” “WordPress hosting types explained”) and the downstream decision queries (“managed WordPress hosting comparison,” “4GoodHosting managed WordPress review”) that bracket the session in which the target query appears.
The classical funnel model — awareness at the top, consideration in the middle, decision at the bottom — is a coarse approximation of this session-progression principle. The research adds two important refinements: first, the transition points between intent stages are where content gaps are most costly (a user moving from awareness to consideration who finds no comparison content from your site will find it from a competitor); and second, the progression is not linear in reality — users loop back, skip stages, and change intent direction based on what they find.
Xu, Ma and Lin’s 2019 Journal of Intelligent & Fuzzy Systems paper on hybrid deep neural intent classification (Priority Score: 87) added that even within a single query, intent classification at the level of individual tokens and phrases can improve on document-level classification — essentially, a seven-word query may carry two distinct intent signals that point to different content needs simultaneously. For practical keyword research, this means that ambiguous long-tail queries sometimes belong on two different pages with different primary purposes, linked to each other, rather than on one page trying to satisfy two distinct intent profiles at once.
QUIDS (Wang, Chen and Verberne, EMNLP 2025, Priority Score: 96) represents the most recent advance in the search intent modelling literature and introduces a concept with direct practical value for keyword research: intent description rather than intent classification.
Instead of assigning a keyword to a category label, QUIDS generates a plain-language description of what a user with that query probably wants — their specific goal, their information gap, the decision they are trying to make. This output is far more useful for content planning than a single label. “Commercial investigation” tells a content strategist that the user is comparing options. “The user wants to understand whether managed WordPress hosting is worth the monthly cost premium over shared cPanel hosting for a WooCommerce store with approximately one hundred products and under a thousand daily visitors” tells them what the page needs to say, how detailed the cost comparison should be, and what the conclusion should help the user do.
The practical implementation for most teams does not require deploying the QUIDS system directly. A simple prompt to any capable language model — “Describe in one paragraph what a user who searches [keyword] is probably trying to do, what they already know, what gap they’re trying to fill, and what outcome would satisfy them” — produces intent descriptions useful enough for content brief writing. Running this process on your top twenty to thirty target keywords before assigning them to pages typically changes at least a third of the assignments and restructures the depth and CTA of another third.
Huang, Cautis, Cheng, Zheng, Mamoulis and Yan’s 2018 ACM Transactions on Knowledge Discovery from Data paper on entity-based query recommendation for long-tail terms (Priority Score: 95) addressed a specific problem that affects keyword research on niche topics: when search volume for a seed term is very low, co-occurrence-based methods and historical log methods have insufficient data to suggest meaningful expansions. The solution demonstrated in the paper is to use entity relationships from a knowledge graph — essentially, known semantic connections between named entities and their attributes, categories, and related concepts — to generate expansion terms even where search log data is sparse.
For SEO practitioners, the accessible equivalent of this method is to identify the key entities in your topic area (the named products, services, organisations, people, and technical concepts) and use structured knowledge sources — Wikipedia categories, Google’s Knowledge Panel, industry taxonomy databases — to map their semantic relationships. The resulting entity map suggests query variations that users might plausibly search even when existing search volume data does not yet confirm demand. This is particularly valuable for emerging technologies, newly-launched products, or niche industry topics where the search behaviour is still developing.
Yang and Li’s 2023 Information Processing & Management literature review (Priority Score: 95) provides the most operationally useful taxonomy of the keyword management process in the database. The review identifies four recurring problem types across the sponsored search literature, each of which maps cleanly to an organic SEO workflow stage:
The diagnosis from the review is that most practitioner workflows treat this as a two-stage process: generate and assign. They collapse stages two and three — often by assigning terms to pages based on intuition rather than explicit opportunity scoring and intent mapping — and they omit stage four almost entirely, treating keyword research for SEO as a deliverable rather than as an ongoing process. Understanding which stage is failing is the first step to improving outcomes.
Research-supported candidate generation draws on five distinct evidence sources, each of which surfaces vocabulary the others miss:
Source 1 — Seed terms from the brief. The starting vocabulary provided by the client, stakeholder, or topic brief. These are typically obvious and well-known terms — necessary but never sufficient on their own.
Source 2 — Internal data: Search Console and site-search. Search Console’s query report filtered for the terms that already bring impressions, especially those with low click-through rate; and site-search logs filtered by query frequency. This is the vocabulary of users already in your market. It requires no external tool and is based on real behaviour, not modelled demand.
Source 3 — Competitor page extraction. YAKE and/or KeyBERT run on the full text of the top ten ranking pages for each primary seed term. This gives the vocabulary that search engines have already determined to be relevant and authoritative on the topic — a proxy for what comprehensive coverage requires.
Source 4 — Autocomplete collection. Gathered across at least three time windows, two session types (incognito and standard), and the relevant geographic markets. Particularly valuable for discovering modifier patterns, question formulations, and emerging terms.
Source 5 — LLM pseudo-document expansion. Using a corpus-steered LLM (seeded with competitor content and Search Console queries) to generate pseudo-documents for each primary seed term, then extracting their vocabulary as described in the Query2doc method. Particularly valuable for discovering expert vocabulary that users may not yet be searching for but that distinguishes comprehensive coverage from shallow coverage.
The union of terms from all five sources, before any scoring or filtering, will typically be three to five times larger than the initial seed list. That breadth is the point — you filter down in Stage 2, but you need the breadth first to avoid the systematic blind spots in any single source.
Apply the composite opportunity score to the full candidate set before committing any terms to a content brief. The five components, drawn directly from Nagpal and Petersen’s framework, with practical estimation guidance:
| Dimension | What it measures | How to estimate it |
|---|---|---|
| Demand | Monthly search volume and trend direction | Volume tool; Google Trends for direction |
| Competition difficulty | How hard it will be to rank | Keyword difficulty score or manual SERP analysis |
| Specificity signal | How focused the query intent is | Word count, modifier presence, question vs. head term |
| Intent match | How well your page serves this query’s purpose | SERP review — do the top results match your page’s purpose? |
| Page/domain strength | Your realistic ranking probability | Compare your DR/DA to incumbents for this specific topic |
Score each dimension on a 1–5 scale. Multiply the scores together (or weight and sum, depending on your model). Sort by composite score descending. The terms at the top of this list represent your highest-probability opportunities — not necessarily your highest-volume terms, and often not your most exciting terms, but the ones where the evidence most supports investment.
Group the scored and filtered candidates by semantic similarity and intent type. A workable clustering approach for teams without specialised tools:
First, group by head concept — all terms that are clearly about the same core topic. Second, within each head-concept group, sort by the intent description you generated using the QUIDS-inspired prompting approach. Third, check whether the intents within a group are compatible — can one page realistically serve all of them without splitting its focus, or do some intent types within the group belong on separate pages?
Peng, Li, Jiang, Wang, Ou, Zeng, Xu, Xu and Chen’s 2024 WWW paper on long-tail query rewriting (Priority Score: 98) contributed a nuance that standard clustering misses: long-tail variants of the same topic often require semantic rewriting rather than just grouping, because the vocabulary a searcher uses when unfamiliar with a product or service differs substantially from the vocabulary an expert would use to describe the same thing. A page about “managed WordPress hosting” needs to cover both the expert vocabulary (server-level caching, PHP worker processes, auto-scaling) and the searcher vocabulary (“does managed hosting update WordPress automatically,” “what’s included in managed WordPress hosting”) simultaneously. The keyword cluster defines which vocabulary the content must bridge.
This is the stage the research literature adds that most practitioner guides omit entirely. Before any keyword is finalised in a content brief or assigned to a writing project, validate it against real demand evidence using a specific question: does this term appear in Search Console, autocomplete, or volume tools with measurable demand at or above a minimum threshold? If the answer is no, the term does not enter the active content plan. It enters a monitoring folder where it is checked quarterly — and if demand develops, you have already mapped the content opportunity and can act quickly.
The validation gate is not a formality. The 2025 COLING paper by Galletti et al. showed that extracted keywords frequently fail to match actual search phrasing. The two 2025 LLM caution papers showed that AI-generated terms can be fluent and topically coherent but demand-free. Without a validation gate, both of these failure modes produce content investment that generates impressions for uncompetitive or low-demand terms while the actual high-opportunity terms remain unaddressed.
The five-stage workflow above is market-agnostic at the method level, but implementation details vary for Canadian businesses and the agencies that serve them. Several considerations are specific enough to warrant explicit attention.
Bilingual keyword architecture. Canadian businesses serving both English and French markets face a keyword research problem that does not appear in most published guides: the two languages are not simply translations of each other at the keyword level. Quebec French search behaviour uses distinct vocabulary — hébergement web for web hosting, référencement naturel for SEO, infonuagique for cloud computing — that does not always map directly to English equivalents. Running the five-source generation process separately for each language, rather than translating the English keyword plan, produces substantially better coverage. This applies both to keyword selection and to LLM pseudo-document generation: corpus-steered expansion for French-language targets should be seeded with French-language competitor pages and Search Console queries, not translations of English prompts.
Provincial regulatory vocabulary. Canadian search queries in regulated industries — financial services, healthcare, legal, real estate, and construction — carry provincial regulatory terms that differ across provinces and are often absent from LLM training corpora because they appear primarily in government and regulatory documents rather than in the high-traffic web content that dominates model training. PIPEDA-related privacy queries, provincial securities terminology, and construction code references are examples. For businesses in these sectors, the internal-data sources (Search Console, site-search logs, support queries) are more reliable than LLM expansion for surfacing the regulatory vocabulary their actual customers use, because that vocabulary comes from real Canadian search behaviour rather than from a model trained predominantly on US-origin content.
Canadian ecommerce search patterns. A best SEO agency in Canada working with ecommerce clients needs to account for the specific search modifiers Canadian shoppers use: “free shipping to Canada,” “ships to Ontario,” “Canadian retailer,” and “HST included” are examples that carry meaningful search volume and commercial intent in the Canadian market but are underrepresented in generic keyword tools calibrated primarily on US search data. The Taobao long-tail query rewriting research (Peng et al., 2024) showed that long-tail ecommerce queries are often formulated in vocabulary that differs from how products are described internally — and this gap is larger in Canadian ecommerce because the additional geographic and regulatory modifiers multiply the long-tail possibilities.
Search Console Canada-specific filtering. The internal-data source (Stage 1 Source 2) is particularly powerful for Canadian businesses because Google Search Console allows performance filtering by country. Filtering to Canada-only queries eliminates the noise from US traffic that dominates many tools’ volume estimates and reveals the specific phrasing Canadian users actually use — often meaningfully different from the US-phrased equivalents that dominate global keyword databases. Running the opportunity-scoring step on Canada-filtered Search Console data rather than global volume estimates produces a keyword priority list that is specifically calibrated to the Canadian market rather than a scaled-down version of a US plan.
Autocomplete market localisation. The autocomplete collection protocol (Stage 1 Source 4) should include collection with Canadian geographic settings and, where bilingual markets are targeted, separate collection with French-language browser and location settings. Canadian autocomplete reflects Google.ca’s index and the search behaviour of Canadian users, which diverges meaningfully from Google.com results for commercially significant queries in regulated industries, regional services, and ecommerce.
The literature consistently and explicitly treats keyword management as a continuous process, not a one-time deliverable. This is one of the clearest points of divergence between the research literature and conventional SEO practice, where keyword research is typically a project phase that concludes with a spreadsheet handoff. For businesses working with an external partner, this distinction has practical implications: a one-time keyword research deliverable without a built-in review cadence is structurally incomplete by the standard the research supports. Whether you manage SEO in-house or engage the best SEO agency in Canada for the work, the engagement model should include quarterly keyword review as a defined deliverable — not an optional add-on.
The evidence for continuous management comes from several directions. Qiao et al.’s 2017 competitive keyword discovery work showed that query-log patterns evolve as new competitors enter topics, user vocabulary shifts, and search engines update how they interpret queries. Jiang et al.’s 2017 temporal autocomplete work showed that demand patterns are dynamic and predictable on seasonal cycles. Li and Yang’s 2022 paper on keyword targeting optimisation showed that the same keyword has different opportunity characteristics under different page-assignment and matching strategies — and those strategies need to be revisited when page performance data becomes available.
A practical quarterly review process: pull Search Console performance for all content pages; identify pages that are ranking in positions 5–15 for queries not in the original keyword plan (these represent discovered opportunities the model missed); check autocomplete for the current state of suggestion patterns on core seeds; and run a YAKE pass on any competitor page that has entered or risen in the top ten since the last review. This process typically identifies two to four new content opportunities per quarter per topic cluster, and it catches performance degradation from competitive encroachment before it becomes significant.
This section is absent from most keyword research guides because it requires acknowledging that current practice has documented failure modes. The academic literature makes that documentation explicit.
Already addressed in Part 1, but worth restating as a mistake rather than just a theoretical limitation. Nagpal and Petersen (2021) showed the volume-opportunity correlation is weak enough that volume-only sorting reliably misranks opportunities. Teams that ship a keyword plan sorted by volume column descending will systematically overinvest in competitive head terms where ranking probability is low and systematically underinvest in specific, high-intent long-tail terms where probability is high.
The two 2025 failure-mode papers (Yoon et al.; Abe et al.) document this explicitly. LLM-generated keyword expansions are generated from parametric knowledge and are unconnected to real search behaviour. A term that appears in an AI-generated keyword list is a topically plausible suggestion, not a confirmed search query. Teams that skip the validation gate build content around terms that generate zero impressions.
Cai, Liang and de Rijke (2016) and Jiang, Chen and Cai (2017) document that autocomplete is time-sensitive and personal. A single collection session underrepresents seasonal variants, emerging terms, and location-specific formulations. Teams that collect autocomplete once during a keyword research project and treat the results as stable miss the early-demand signal that autocomplete monitoring over time provides.
Galletti et al. (2025) showed that extracted keyphrases do not reliably match search phrasing. Teams that extract keywords from competitor pages and add them directly to content briefs without demand validation are optimising for document vocabulary rather than search vocabulary — a meaningful distinction with direct consequences for whether organic impressions accumulate.
Chen et al. (2020) on session-based query progression and the QUIDS intent description framework both show that intent within a topic cluster is not uniform. Teams that assign an entire intent-diverse keyword cluster to a single page create content that partially satisfies multiple intent types and fully satisfies none of them. A page trying to serve “what is managed WordPress hosting” and “buy managed WordPress hosting” simultaneously is structurally incomplete for both audiences.
The competitive keyword discovery research (Qiao et al., 2017) and temporal autocomplete research (Jiang et al., 2017) together show that annual review cycles miss a substantial portion of demand evolution. Competitors enter topic areas, update their vocabulary, and change their SERP presence throughout the year. An annual review catches changes that are already fully reflected in historical data — it cannot provide lead time. Quarterly review is the minimum the literature supports; monthly autocomplete monitoring adds early-warning capability at low marginal cost.
Understanding how the research progressed helps practitioners understand why certain methods are now recommended and what principles are stable versus what is likely to change further.
2016–2018: Statistical foundations and the first embedding work. The early years of the database established the foundational frameworks: the four-box intent taxonomy was examined critically (Cai and de Rijke, 2016); statistical query expansion methods were surveyed and benchmarked (Raza, Rahmah and Noraziah, 2018); and the first embedding-based keyphrase extraction methods appeared (EmbedRank, 2018; Key2Vec, 2018; PositionRank, 2017). These papers established the vocabulary and the baselines that later work is measured against. The commercial SEO toolkit during this period was built primarily on TF-IDF frequency models and manual volume sorting — already lagging behind what the research was showing.
2019–2020: Pre-trained language models transform the baseline. BERT’s release in late 2018 was followed by a wave of BERT-based keyword extraction and query expansion papers throughout 2019–2020. BERT-QE (2020) showed that contextual passage retrieval improves query expansion substantially over static embeddings. SIFRank (2020) brought language-model representations to keyphrase extraction. YAKE (2020) provided the practical no-training baseline that remains the fastest extraction method in the cluster. This period also produced the highest-scoring pure SEO papers in the database: Nagpal and Petersen on composite keyword scoring (2021, but the research behind it was developing through this period), and the first formal work on keyphrase diversity for SEO applications.
2021–2022: Embedding maturity and portfolio thinking. The academic community spent 2021–2022 consolidating the embedding gains and beginning to examine the portfolio dimensions of keyword management. Symitsi et al.’s portfolio theory paper (2022) and Erdmann et al.’s long-run strategy paper (2022) introduced economic modelling frameworks that the SEO industry had not previously applied. The information retrieval conferences during this period produced sophisticated multi-stage retrieval pipelines that would shortly be reimplemented with LLMs.
2023–2024: The LLM inflection point. Query2doc (2023) and the LLM prompting papers (2023–2024) represented a step-change in what was achievable for query expansion and keyword generation. The Corpus-Steered QE and Analyze-Generate-Refine papers (both 2024) provided the practical frameworks for grounding LLM capability in domain-specific vocabulary. The Taobao long-tail rewriting paper (2024) demonstrated production-scale deployment of LLM-based query improvement. By the end of 2024, the academic consensus was that LLMs are valuable keyword vocabulary generators but require grounding, validation, and explicit failure-mode handling.
2025–2026: Failure modes, fairness, and fresh retrieval. The 2025–2026 papers represent the field’s maturation phase: documenting where the 2023–2024 optimism overshot (the two failure-mode papers), introducing methods that combine the strengths of LLM generation with the reliability of corpus grounding (ThinkQE, 2025; Knowledge-Aware QE, 2025), and pushing retrieval freshness to web scale (the 2026 SIGIR paper on dense next-query suggestion). The research direction points toward hybrid systems that combine LLM generation, corpus grounding, and real-time retrieval — and toward evaluation practices that are grounded in real search behaviour rather than academic benchmark proxies.
What does the research say is the single biggest mistake in keyword research?
Sorting by search volume alone and treating the highest-volume term as the highest-opportunity term. Nagpal and Petersen (2021) demonstrated this explicitly: optimal keyword selection requires a composite of demand, competition difficulty, specificity, intent alignment, and page or domain strength. Volume-only sorting ignores four of the five variables that determine whether a keyword is actually winnable and worth winning.
Is AI-generated keyword expansion reliable for SEO?
Conditionally. LLMs generate semantically plausible expansion terms effectively, particularly when grounded on domain-specific corpus content. However, two 2025 papers document specific failure modes: expansion can reflect memorised training data rather than real search demand, and it degrades significantly for niche, emerging, or ambiguous queries. All AI-generated keyword candidates must be treated as hypotheses and validated against real demand data — Search Console, autocomplete, or volume tools — before inclusion in a content brief.
What is the best keyphrase extraction method for competitor analysis?
For speed and no-setup scenarios, YAKE requires no training data and runs in milliseconds on any text. For strongest alignment with actual search phrasing, KeyBERT had the best performance on the real-search benchmark in the 2025 COLING evaluation. For long-form content where structural prominence matters, PositionRank’s positional weighting correlates well with what search engines identify as topically important. A practical workflow uses all three: YAKE for breadth, KeyBERT for search-phrase alignment, PositionRank for long-form structural analysis.
How should branded and generic keywords be managed differently?
As separate sub-portfolios with different time horizons, success metrics, and investment logic. Erdmann, Arilla and Ponzoa (2022) showed that branded terms have faster short-term conversion but bounded demand ceilings; generic terms rank and convert more slowly but compound over time without bounded demand. Evaluating both on the same three-month ranking timeline systematically undervalues generic keyword investment and leads teams to abandon generic content before it has had time to compound.
How often should a keyword plan be reviewed?
Quarterly at minimum for content programmes of any size. The research on autocomplete temporal patterns (Jiang et al., 2017) and competitive keyword dynamics (Qiao et al., 2017) both show that demand patterns evolve continuously through competitor activity, seasonal patterns, and user vocabulary shifts. Monitoring autocomplete changes every four to six weeks provides additional early-warning capability at low marginal cost.
What is the Analyze-Generate-Refine workflow and why does it outperform single-pass LLM expansion?
It is a three-stage process: first, analyze the SERP to identify topic facets currently underserved in top-ranking content; second, use a corpus-steered LLM to generate candidate keyword variants and expansion terms for those specific gaps; third, filter the generated candidates against real demand data and discard terms with no confirmed search evidence. It outperforms single-pass expansion because the analysis step prevents generating more of what already exists, and the refinement step prevents investing in fluent-but-demand-free terms.
How does working with the best SEO agency in Canada differ from using generic keyword tools?
The research makes a clear distinction between keyword generation and keyword selection. Generic tools provide volume and difficulty data for a predetermined list of terms — the generation step. The selection step — composite opportunity scoring, intent mapping, cluster design, and validation against real demand — requires analysis that tools alone cannot perform. A capable SEO agency applies the full five-stage workflow: multi-source generation that includes internal data and competitor extraction; composite scoring that weights specificity and intent match alongside volume; clustering that produces intent-coherent page assignments; demand validation before any term enters a content brief; and a defined quarterly review process. For Canadian businesses specifically, the agency should also adjust the workflow for bilingual markets, Canada-filtered Search Console data, and the regulatory vocabulary gaps that generic LLM tools systematically underrepresent.
What does “intent description” mean and how is it different from intent classification?
Intent classification assigns a keyword to a category label: informational, commercial, transactional, or navigational. Intent description generates a plain-language statement of what a specific searcher with that query probably wants — their goal, their information gap, and the outcome that would satisfy them. The description is far more useful for content planning because it determines what the page needs to say, how deep the treatment needs to be, and what the CTA should help the user do. The QUIDS research (Wang, Chen and Verberne, 2025) formalised this approach; it can be approximated with a structured LLM prompt on any target keyword.
What should Canadian businesses look for when evaluating AI SEO services in 2026?
Three things the research supports as meaningful differentiators. First, does the provider ground LLM-generated keyword suggestions in Canada-specific data — Search Console filtered to Canada, Canadian autocomplete collection, and Canadian competitor pages — rather than applying US keyword plans to Canadian markets? Second, does the service treat AI-generated keywords as hypotheses requiring demand validation, or does it present AI outputs as ready-to-use keyword lists? The 2025 failure-mode papers show that skipping validation is a documented error, not a minor shortcut. Third, is keyword research structured as an ongoing process with quarterly review, or as a one-time deliverable? The research literature is unambiguous that continuous management outperforms periodic projects. AI SEO services for Canadian businesses in 2026 should reflect the research evidence on all three dimensions.
What is corpus-steered LLM expansion and why is it better than cold prompting?
In cold prompting, you ask an LLM for keyword expansions with no additional context. The model draws on its full training distribution and generates terms common across all documents on the broad topic — generic vocabulary rather than niche-specific vocabulary. In corpus-steered expansion, you first retrieve the most relevant passages from your domain’s specific vocabulary (competitor pages, your own content, Search Console queries) and include those passages in the prompt context. The model’s output then reflects the niche vocabulary rather than generic topic vocabulary. Lei et al. (2024) showed this produces substantially better expansion quality, especially for niche topics where the LLM’s unprompted knowledge is weakest.
A decade of academic research on search behaviour, keyword selection, query expansion, and intent modelling has produced a picture of keyword research for SEO that is substantially more rigorous, and more actionable, than the conventional practitioner playbook suggests. The field has moved from frequency matching and intuition-based selection to multi-variable composite scoring, portfolio management with distinct time horizons, iterative LLM-grounded expansion with explicit validation gates, and continuous monitoring rather than periodic project delivery.
The four main revisions to standard practice that the evidence supports are: replace volume-only sorting with composite opportunity scoring that incorporates competition, specificity, intent match, and page strength; treat keyword sets as diversified portfolios with long-run time horizons and separate branded versus generic management; use LLMs as corpus-grounded vocabulary generators with mandatory demand validation, not as demand oracles; and build autocomplete monitoring into the ongoing workflow as an intent and trend signal checked every four to six weeks rather than collected once.
For Canadian businesses in particular, the framework carries additional specificity. The best SEO agency in Canada working in this environment needs to adapt the standard five-stage workflow for bilingual markets, Canada-filtered data sources, provincial regulatory vocabulary, and the specific gaps that globally-trained LLMs produce when applied to Canadian search behaviour. AI SEO services for Canadian businesses in 2026 that simply apply US keyword methodologies to Canadian domains are not applying the research — they are applying a methodology calibrated for a different market’s search behaviour, regulatory environment, and linguistic landscape.
None of these revisions requires proprietary technology, advanced data science capability, or large budgets. They require a documented process — which the research has now outlined in enough detail to implement directly. The gap between what the evidence supports and what most SEO programmes actually do is not primarily a technology gap. It is a workflow gap — and for Canadian businesses, it is also a localisation gap.
Implementing a research-grounded keyword process produces better content plans — but only if the resulting content is actually readable by the systems that determine visibility. For Canadian businesses investing in AI SEO services in 2026, the technical infrastructure underneath your content is not a separate concern from the keyword strategy; it determines whether the pages built from that strategy are indexed, rendered, and surfaced by every major search engine and AI retrieval system.
4GoodHosting provides Canadian businesses with the hosting infrastructure that supports the full research-grounded workflow: Canadian data centres that keep your content on Canadian IP addresses and within Canadian data sovereignty, server-side rendering that ensures AI crawlers can read your pages without JavaScript execution, and the fast first-byte response times that determine whether AI retrieval systems include your content in their citation pools. For any best SEO agency in Canada building content programmes on behalf of Canadian clients, hosting infrastructure is the layer that either amplifies or limits the keyword work. Talk to 4GoodHosting about the Canadian hosting infrastructure that supports your content and SEO programme.





