Rank forecasting workspace

Glossary

Every signal on the Data Coverage page, in plain English: what it means, exactly how its count is computed, and why it matters to the forecast. Nothing here is a score — every number is a count of real observations.

Ranking Velocity

Journey data3,607 / 100,000 observations
What it is
How fast a page's Google position is moving right now, for a specific keyword.
How it's counted
For every page + keyword pair, we take its average position over the last 28 days and compare it to the 28 days before that. Every pair that has data in both windows counts as one observation. Positive velocity = moving up.
Why it matters
A page already climbing tends to keep climbing; one slipping tends to keep slipping. Recent movement is one of the strongest hints about what happens next.

Topical Attainment

Ranking outcome7,397 / 25,000 observations
What it is
Proof of what a site is capable of ranking for: page + keyword pairs that made it into Google's top 20 (roughly pages 1–2).
How it's counted
We count every page + keyword pair that recorded an average position of 20 or better on at least one day it got impressions.
Why it matters
A site that has broken into the top 20 many times in a topic is far more likely to do it again — this is the raw material for “has this site done it before?”

Stable Ranking Attainment

Ranking outcome2,123 / 25,000 observations
What it is
Journeys that didn't just touch a good position for a day, but held it.
How it's counted
“Stable” means: within some 21-day stretch, the page ranked at or better than the threshold (top 20 or top 10) on at least 70% of the days it had impressions, with a minimum of 10 such days. We count journeys with at least one stable milestone.
Why it matters
Stable top 10 / top 20 is the exact event the forecast predicts — a one-day spike doesn't count as “ranking.”

Content Integrity & Decay

Journey data2,127 / 25,000 observations
What it is
Pages that used to rank better than they do now — rankings rotting over time.
How it's counted
We count page + keyword pairs whose average position over the last 28 days is more than 5 spots worse than the best position they ever held.
Why it matters
Decay is the downward force the forecast has to price in: rankings aren't only won, they're also lost.

Low-Volume Noise

Data quality3,926 / 50,000 observations
What it is
Pairs where Search Console simply doesn't have enough data to trust.
How it's counted
We count page + keyword pairs with fewer than 10 days of recorded impressions. Search Console only reports position on days someone actually saw the page in results — below ~10 such days, the “average position” can swing wildly.
Why it matters
Knowing how much of the data is noisy tells us (and you) how much to trust journey-level numbers — these pairs get flagged, not trusted.

Keyword Cannibalization

Data quality8,585 / 25,000 observations
What it is
Two or more of a site's own pages competing against each other for the same keyword.
How it's counted
For each site, we group journeys by keyword; every journey in a keyword-group that contains more than one page counts as one observation.
Why it matters
Google generally won't rank many pages from one site for one query — cannibalized pages split their chances, and it's one of the most fixable problems we can detect.

Pre-Ranked at Import

Data quality1,614 / 50,000 observations
What it is
Journeys that were already underway before our data window opened — we missed the start of the movie.
How it's counted
We count journeys whose first recorded impression falls within the first 14 days of the site's imported history window (Search Console only gives us ~16 months back).
Why it matters
If we didn't see the climb start, we can't measure how long it took — so these are excluded from time-to-rank training instead of being counted as instant wins.

Brand Query Influence

Data quality0 / 25,000 observations
What it is
Searches for the site's own name (or containing it), where ranking is automatic rather than earned.
How it's counted
We take the site's domain name minus the extension (e.g. “buildtheshelf”) and count journeys whose keyword contains that token (minimum 4 characters).
Why it matters
You'll rank #1 for your own brand no matter what SEO you do — including these would poison the training data with fake “fast wins,” so they're excluded.

Censored Outcomes

Ranking outcome6,666 / 50,000 observations
What it is
Journeys we watched for a while that hadn't reached the target by the time our observation ended.
How it's counted
We count journeys whose status is “censored”: tracked, never hit the target while we could see them, and no longer receiving data.
Why it matters
The statistics we use (survival analysis) are built to use these correctly — “hasn't ranked yet” is real information, and treating it as “failed” would bias every forecast.

Journey Corpus

Journey data11,199 / 100,000 observations
What it is
The total library of ranking stories the forecast can learn from.
How it's counted
Every page + keyword pair with at least 5 days of impressions counts as one journey — the basic unit of everything else on this page.
Why it matters
Forecasts work by finding journeys similar to yours and seeing how they played out. More journeys = more precise, more confident forecasts.

Daily Observation Density

Journey data425,998 / 1,000,000 observations
What it is
The raw daily data underneath everything: one row = one page, one keyword, one day.
How it's counted
A straight count of daily rows imported from Search Console (clicks, impressions, average position per page × keyword × day).
Why it matters
This is the substrate every other signal is computed from — its size is the closest thing to the dataset's total weight.

Intent-Classified Queries

Journey data3,175 / 50,000 observations
What it is
Which of five jobs a search is trying to do: learn something (informational), compare options (commercial), take an action like buying or booking (transactional), reach a specific site (navigational), or find something nearby (local).
How it's counted
Every keyword gets run through a transparent set of word rules (e.g. “buy/price/coupon” → transactional, “best/vs/review” → commercial, “how/what/guide” → informational, “near me” → local). If we've seen the live results page, its features can override the words — a map pack means local intent regardless of phrasing. Keywords matching no rule stay honestly unclassified. We count keywords with a classification.
Why it matters
Time-to-rank differs enormously by intent — a transactional keyword behaves nothing like an informational one. Forecasts now match network comparables by intent first, which makes “similar journeys” actually similar.

Update-Adjacent Attainments

Data quality4,782 / 25,000 observations
What it is
Ranking wins that happened suspiciously close to a confirmed Google algorithm update.
How it's counted
We keep a calendar of updates Google has officially confirmed (from their Search Status Dashboard). Any journey whose first top-20 or top-10 date lands within ±14 days of an update's rollout window counts.
Why it matters
During updates, rankings move because Google changed, not because of anything the site did — these milestones stay in the data but are flagged so the model can discount them.

Site Diversity Ceiling

SERP data1 / 10,000 observations
What it is
Search results where one website already holds two or more of the top 10 spots.
How it's counted
For each live SERP check we store, we count it if any single domain appears 2+ times in the top 10.
Why it matters
Google deliberately limits how many pages one site can show for a query. If a site already has two spots, a third page targeting that query is structurally blocked — the forecast warns about this.

SERP Snapshot Coverage

SERP data3 / 50,000 observations
What it is
How many real Google results pages we've captured and stored.
How it's counted
A straight count of SERP snapshots taken (each snapshot = the actual top ~50 results for one keyword, in one country, on one device, at one moment).
Why it matters
Search Console's “position” is a blended average, not a true rank. Snapshots are ground truth — they verify milestones and track competitors.

Live SERP Presence

SERP data2 / 25,000 observations
What it is
SERP checks where the tracked site actually appeared in the results.
How it's counted
Of all snapshots taken, we count the ones where the tracked domain was found in the top ~50.
Why it matters
The ratio of presence to coverage shows how often tracked targets are visible at all — and each appearance is a true-rank data point, better than any estimate.

Competitor Composition

SERP data147 / 500,000 observations
What it is
The individual competing listings we've observed — who occupies the positions our users want.
How it's counted
Every organic result stored inside every snapshot counts (about 50 per check): its position, URL, domain, and title.
Why it matters
Over time this reveals how crowded and how stable each SERP is — a results page that never changes is much harder to break into than one that churns.

Competitor On-Page Profiles

SERP data20 / 25,000 observations
What it is
What the winning pages actually look like inside: how long they are and how many internal and outbound links they carry.
How it's counted
During competition analysis we fetch each top-10 page (and the tracked page), strip the code, and count words plus links pointing within the same site (internal) versus elsewhere (external). Pages that block our fetch are recorded as blocked, not guessed. Each successfully read page is one observation.
Why it matters
It grounds “is my page competitive?” in measurements of the pages that are actually winning — content depth and link structure of the real top 10, not generic best-practice folklore.

Content Interventions

SEO actions0 / 10,000 observations
What it is
SEO work logged in the work log that changed page content.
How it's counted
A count of work-log entries in the “content change” category (rewrites, expansions, refreshes).
Why it matters
The long-term goal: connect what was done to what happened afterward. That's only possible if the work is recorded — these logs become the model's “actions” inputs.

Technical Interventions

SEO actions0 / 10,000 observations
What it is
SEO work logged in technical categories — fixes, redirects, structured data.
How it's counted
A count of work-log entries in the “technical fix,” “redirect / consolidation,” and “structured data” categories.
Why it matters
Technical changes often affect a whole site at once; logging them lets the model treat them as events that touch every journey on the site.