Methodology
How we measure
Everything Mapford reports comes from one activity: asking AI assistants the questions your customers ask, many times over, and logging every page they read to build each answer. This page explains why we do it that way, in plain English, with the receipts.
Why one sample lies
Ask an assistant “best accounting software for restaurants” twice and you will often get two different shortlists. The models are probabilistic: the same question can surface different pages, different names, different order. A tool that asks once and shows you the result isn’t measuring anything — it’s quoting a coin flip. Any single answer is one draw from a distribution, and the distribution is the thing that matters.
Why at least eight runs per question
We run every question a minimum of eight times per assistant. That floor isn’t arbitrary: Schulte et al. (arXiv:2604.07585) measured run-to-run variance in assistant answers and set two floors: seven samples per prompt if you only want to know whether a brand gets named, and at least eight if you need to know which pages the assistant actually read. We measure pages, not just names, so we hold to the higher floor. Below it, the rates are so wide they’re close to meaningless. In the sample report we used 8 runs per question per assistant — 240 sampled answers in total.
A full measurement multiplies that out. It is 35 buying questions, asked 8 times each on every one of 5 assistants: 35 × 8 × 5 = 1,400 sampled answers, every cited page in every one of them fetched and classified. Adding an assistant adds samples rather than dividing the ones we already have, because the floor of eight is per assistant — if it were shared out, a five-assistant measurement would be worse evidence than a three-assistant one at the same price, which is how a coverage number gets sold as an improvement.
What that buys is a floor on what we can honestly call a change. At 1,400 sampled answers the smallest month-over-month move distinguishable from the assistants’ own randomness is 4.2 percentage points, and every report prints its own version of that figure. Anything smaller we report as no detectable change, in those words, rather than as a small win.
The evidence this rests on
Four findings, all published in 2026, all checkable. We cite them because the floor we hold to is expensive, and a buyer is entitled to know whether we invented the reason.
- Schulte et al., arXiv:2604.07585 — four categories, eight questions, four assistants, sampled daily for two months, plus thousands of pairs of simultaneous re-runs. Two results decide our design. The pages cited on one day overlap those cited the next by about a third; and re-runs fired at the same moment overlap in the same range, so the variation is the model’s, not the day’s. Their recommendation: seven runs per question for brand-level detection, at least eight when you need source-level coverage, and a window of weeks rather than days before a trend means anything.
- Sielinski, arXiv:2603.08924 — measured how far citation shares wobble, and found differences of around three and a half percentage points falling inside the margin of error. Its recommendation is the one we follow on every number we print: report the interval, not the point.
- Sielinski, arXiv:2607.10341 — thirty assistant-and-topic combinations, of which most needed between thirty-three and ninety-four answers before the numbers settled, and three never settled at all after a hundred and twenty-five. The honest reading is that no fixed sample size is universally enough, which is exactly why we print the range beside every rate instead of promising a precision the sample may not have reached.
- Fishkin and O’Donnell, SparkToro, 27 January 2026 — six hundred volunteers ran twelve questions through the real consumer apps, 2,961 runs in all. The chance of getting the same list of brands twice was under one in a hundred. The same study found that a visibility rate measured across many questions and many runs is stable enough to be useful, while a position in the list is not. That is the shape of this whole product: rates yes, ranks never.
What a range means
Every rate we publish looks like 38.0% (32–45, n=210). Read it as: out of 210 sampled answers, you were named 38.0% of the time, and if we kept sampling forever, the true rate would very likely settle between 32% and 45%. The range is a 95%confidence interval (we use Wilson intervals, which is why the range isn’t symmetric around the point — small samples genuinely aren’t symmetric). A bare “38%” hides all of that. The n tells you how much evidence sits behind the number; the range tells you how seriously to take it.
Which assistants we ask, and what we are actually asking
“We measure ChatGPT” is not one claim, it is several, and they are not equally strong. Three of our five assistants are reached through a vendor’s programming interface with web search switched on; two are real Google results pages read the way a visitor sees them. Three name the model that answered; two cannot, because the vendor discloses none. The one published measurement of that gap we could find is a seller’s own: MentionsAPI, which sells this kind of data, states that across 1,000 questions the answers it got from a programming interface differed from the same question asked on chatgpt.com — in the pages cited, in the set of brands, or in their order — on 96% of them. We read that on 1 September 2026. It is a vendor’s own count about its own product’s premise, published with a base and no range, and we have not reproduced it; treat it as the reason we declare the route for every assistant in the table below rather than as a figure to lean on. The distinction itself is not a technicality — it is the difference between what we measured and what your customer saw. Surveying this market in September 2026 we found no vendor page stating which model tier its numbers came from; that is the second reason the table exists.
ChatGPT
What we ask
Reached through our search-data provider, on OpenAI's programming interface with web search turned on.
Model recorded
gpt-5.6-terra
Pin last verified
25 Aug 2026
Gemini
What we ask
Reached through our search-data provider, on Google's programming interface with web search turned on.
Model recorded
gemini-3.6-flash
Pin last verified
25 Aug 2026
Perplexity
What we ask
Queried directly on Perplexity's own programming interface — no intermediary.
Model recorded
sonar
Pin last verified
25 Aug 2026
Google AI Overviews
What we ask
A real Google results page for the question, fetched through our search-data provider, in the United States on desktop.
Model recorded
None — Google discloses no model
Pin last verified
Nothing to verify
Google AI Mode
What we ask
Google's own AI Mode results page for the question, fetched through our search-data provider, in the United States on desktop.
Model recorded
None — Google discloses no model
Pin last verified
Nothing to verify
Reading the table
Two consequences worth saying out loud. First, a pinned model is a disclosed proxy, not a claim of identity: ChatGPT serves different models to different subscription tiers and routes upward on hard questions — including, precisely, the compare-and-recommend questions we ask — so no single pin can be the same thing as the app. Gemini is the cleaner proxy, because its app does not route. Second, for the two Google surfaces, “we record which model answered” is simply false, and we would rather write none in the table than put a plausible model name in a row nobody could check. A question Google shows no answer summary for is counted as an answer that cited nothing, not as a failed sample — whether the summary appears at all is part of what is being measured.
Each assistant’s pin last verified date is the last time a person checked that pin against the vendor’s own release notes, not the last time the code changed. Pins go stale: the previous one went stale in eighteen days. If a date there is old, treat the proxy as older evidence, and ask us.
The honest funnel
“How often does the AI recommend me?” is really three questions multiplied together, and we report each one separately so you can see where you actually lose:
- Activation — of all sampled answers, how often did the assistant name specific businesses at all (rather than giving generic advice)?
- Shortlist — when it named businesses, how often did it produce a real shortlist of the kind you could appear on?
- You — when there was a shortlist, how often were you on it?
Multiplying the three gives your effective visibility. Reporting only the last stage — as most tools do — silently drops the denominator and flatters everyone. We keep all three stages visible, each with its own n and range.
What we refuse to report
Two things you will never find in a Mapford report. First, answer-position (“you were ranked #2”) and sentiment scores at low sample sizes: below n=100 those metrics are noise dressed up as insight, so we don’t show them — a point Evertune’s precision ladder makes well: the finer-grained the claim, the more samples it demands, and position and tone are the most fragile claims of all. Second, guarantees. Nobody controls what an assistant will say tomorrow, including us. What we can measure is which pages it reads, whether you’re on them, and which of them are worth going after.
Which pages decide your category
Knowing which pages the assistants read (Martinez et al., arXiv:2607.14035 showed answers lean heavily on a small set of retrieved pages) only matters if you know which ones you can act on. So the report leads with a list, not a score: the 25 most-cited pages, who controls each one, and whether you can reach it. In the sample report 11 distinct pages were cited and the top 10 carried 95.8% of the 618 citations those 240 sampled answers produced. That share carries its base and no range, and should not carry one: it is how a single period’s answers were composed, not a draw from a population. A few pages usually decide a category, and the honest question is who owns them.
Who controls a page
We do not classify by domain and guess. We fetch every cited page and read it: does it name three or more of the products in your category, does it carry an advertiser disclosure, is it selling a product of its own, could we read it at all. From that, each page gets one of these classes:
- Competitor’s own page — published by one of the products the assistants recommend, including its own “best of” lists. A rival will not add you.
- Comparison page — adjacent business — a page comparing your category, published by a business that sells something else to the same buyers. It has no reason to refuse a listing; this is the class we exist to find.
- Comparison page — independent publisher — a comparison with nothing to sell and no disclosure.
- Comparison page — affiliate — a comparison carrying an advertiser disclosure. It lists what pays it per lead.
- Single-product page — a review, an integration page or a case study about one or two products.
- Community — forums, video and question sites.
- Major publication — a large publication with a real submission path.
- Locked by structure — Google’s own surfaces, institutions, journals, big retailers, closed editorial, directories where everyone is listed already.
- Unread — a page we could not fetch, or too thin to judge.
Placeable, gettable, locked — and unknown
Those classes roll up into four buckets, and every report shows all four so a number never hides what it excludes:
- Placeable — comparison pages run by adjacent businesses or independent publishers, and community threads. You can realistically get on these: typically $0–500 and 7–90 days per placement.
- Gettable — major publications, single-product pages, and affiliate comparison pages. Reachable with editorial effort or a commercial relationship: typically $250–3,000 and 30–180 days. Affiliates sit here on purpose: a site that lists what pays it per lead is a relationship over months, not a placement a newcomer buys for a few hundred dollars — and we would rather under-promise than call it placeable.
- Locked — competitors’ own pages, big retailers, institutions, closed editorial. Nothing to buy into; we tell you so instead of selling you the attempt.
- Unknown — pages we could not read. Reported as their own share, and never counted as reachable. An earlier version of this rubric treated any unrecognised page as reachable; that flattered every category, so we stopped.
The cost and time bands are estimates by page class, not quotes — they exist so you can rank effort, not budget to the dollar. The placeable share still appears in every report with its unknown share beside it, but the thing we claim is the list: these pages, these owners, these you can reach.
How we report change
A measurement period is a calendar month, never a week. Asking the same question twice on the same day returns a different set of pages about as often as asking it a day apart, so a week of samples mostly measures the assistant's own randomness. A month clears the window the research recommends and lands when a marketing team can actually act. There is no weekly product here at any price, which is a decision rather than a limitation.
Within a month, the samples are not taken in one sitting. A run spreads them across at least three separated days, so a single afternoon — a model update, a provider outage, one unusual news cycle — cannot define the month. The research says day-to-day drift is not the largest source of variation, so this is prudence rather than necessity; it costs nothing but patience, and it removes an objection we would otherwise have no answer to.
There is one exception, and it is labelled rather than hidden. An agency pitching a prospect can run a single-day measurement, at the same floor of 8 runs per question per assistant. Every rate on such a report carries the words “Snapshot — a single-day measurement, not a period” beside it, its range is wider and printed as always, and the product refuses to compare it with anything — it cannot become the baseline for a later month, it is named and excluded from the month-by-month chart rather than quietly dropped, and it never enters a running total. A same-day number is useful for a conversation. It is not a period, and the report says so on every screen where a number appears.
When a report follows an earlier one for the same brand, it compares the two — but only if they measured the same thing: the same questions, the same assistants and models, the same competitors, and the same page rubric. If any of those changed, the report says so and prints no change at all, because a number there would be fiction.
Every difference between two rates carries its own range, computed from the two samples. Where that range includes zero the report says no detectable change, in those words, and shows the smallest move the two samples could have detected. It never shows a direction on its own, and never a colour without a word.
What we do when the engine changes under you
Assistants change how they search, and when they do the ground moves under everyone at once. On 8 August 2026 ChatGPT changed the way it fans a question out into searches; as measured by Promptwatch, Reddit's share of the pages it cited fell by 86% in a day, while the same measurement showed Google's answer summaries down 11%. Those two figures are theirs, dated and linked again in the log below; we have not re-run them. Nobody's marketing team did that. A report that showed you the resulting move and called it your month would be selling you a story.
So we keep a dated record of those changes and lay it over the comparison. Where one falls inside the period, the affected rows carry the date, what changed, and a link to the source, and the report says the move is not attributable to your work. Where the change hit an assistant carrying enough of the samples to account for the whole move on its own, the report drops the direction entirely and prints the size of the move with its range instead. “Enough” is arithmetic, not taste: if an assistant carries a fifth of the evidence, the most a change confined to it could move the pooled rate is a fifth, so it can only explain a move no larger than that.
What we never do is adjust the number. Subtracting an estimated engine effect would be a guess about the assistant's behaviour dressed up as a measurement, and you would have no way to tell the two apart. The figures and their ranges are exactly what we counted; the dates sit beside them so you can hold both at once.
The public record
The record is public, and this is it — the same rows that annotate a client’s report, read live rather than retyped here. Only we can add to it: an agency that could enter an engine change could invent one to explain away a bad month, which is the exact dishonesty the log exists to prevent. Where an entry has no primary source page we say so in its own text rather than leave a link-shaped gap.
8 Aug 2026 · chatgpt · Broke the trend line
ChatGPT fan-out switched to site-scoped retrieval
In a single day the share of ChatGPT's fan-out searches using the site: operator went from 0.37% to 16.8% (~46×) and queries per response rose from 1.08 to 1.83. Reddit's citation share fell from 3.83% to 0.52% — down 86% on ChatGPT — while the same measurement showed AI Overviews down 11% and AI Mode down 30%. Those are Promptwatch's own counts, published with a base and no range, and we have not re-run them. This is the canonical case for this whole feature: an 86% move on one engine and 11% on another, on the same day, caused by neither site.
6 Aug 2026 · chatgpt · Changed what gets cited
GPT-5.6 Luna became the ChatGPT Free and Go default
The Free and Go tiers moved to GPT-5.6 Luna while Plus, Team and Enterprise stayed on GPT-5.5 Instant, and the router escalates on exactly the comparative buying questions we measure. There is therefore no single 'ChatGPT model'; our pin is a disclosed proxy, and this is the date the consumer default moved away from it.
21 Jul 2026 · gemini · Changed what gets cited
Gemini 3.6 Flash became the default across tiers
Gemini 3.6 Flash became the default model on every tier of the Gemini app. This is the model our Gemini measurements record, as a disclosed proxy for what the app answers, so it is the date the proxy started being accurate rather than the date it drifted. Verified by us on 25 Aug 2026 against the vendor's published default; no public announcement exists to link.
7 May 2026 · chatgpt · Changed what gets cited
ChatGPT began hyperlinking brand names to homepages
Brand names in answers became links to the brand's own homepage. Referrals from OpenAI to monitored brand sites were reported to roughly double overnight. This changes which URLs appear in an answer's citation set, so it changes the citation map a report is built from.
6 May 2026 · gemini · Recorded
Google added inline links and a collapsible sources panel
More inline links beside relevant text, hover source previews on desktop, and a collapsible Sources panel; secondary reports said click-through on cited content fell, and a study by ALM Corp over 1.3 million citations found AI Mode citing Google properties in about 17% of answers. We could not find a primary page for the click-through claim, so no figure is given for it. Reported on AI Overviews and AI Mode, which we now sample separately; filed against Gemini as the Google surface this log tracked at the time, and kept at the lowest severity for that reason. Secondhand: no primary page was verified.
11 Mar 2026 · chatgpt · Recorded
GPT-5.1 retired from ChatGPT
OpenAI removed the GPT-5.1 family from ChatGPT, and states the date in its own model release notes. Recorded because the model that answers is part of what a measurement means: a run before this date and a run after it were answered by different models, even though both are labelled ChatGPT.
See it applied
The sample report is a worked example: this method applied end to end, on realistic data, to a fictional company. Every number in it carries its sample size and range — exactly as yours would. What it costs is on the pricing page, with the same arithmetic and no allowance left unstated.