Ask a GEO agency how they measure AI visibility and you will usually get one of three answers: a screenshot, a traffic chart, or a single percentage. None of the three survives contact with how AI answers actually behave. This piece publishes our method instead — the six surfaces we track, the prompt discipline behind it, why a human still runs a monthly pass by hand, and the numbers we deliberately refuse to report.
We are publishing it in full for a simple reason. Measurement is the part of this category that is easiest to fake and hardest for a buyer to check, which means an agency that won't describe its method in public is asking for trust it hasn't earned. If you read this and decide to run a lighter version yourself, that's a fine outcome — the last section explains how.
Why a screenshot proves nothing
AI answers are not stable objects. The same prompt, asked twice, can return different firms — because the model samples probabilistically, because retrieval pulls a different slice of sources, because your account has memory and personalization switched on, because you're in a different city, or because the engine shipped an update between Tuesday and Thursday.
This has a blunt consequence: a single observation carries almost no information. A screenshot showing your firm named is not evidence you are visible, and a screenshot showing you absent is not evidence you aren't. Both are one draw from a distribution.
Everything below exists to convert that instability into something you can actually read: a fixed set of prompts, run repeatedly, across named engines, recorded the same way every month, so that the movement you see is the firm's position changing rather than the sampling noise.
The six surfaces, and why they're reported separately
Our panel treats six surfaces as first-class, each with its own reported row:
- ChatGPT
- Perplexity
- Google AI Overviews
- Google AI Mode
- Gemini
- Claude
Three of those deserve a note. Google AI Overviews and Google AI Mode are frequently collapsed into "Google AI" in vendor dashboards, and they behave differently enough that collapsing them hides real movement — a firm can be present in one and absent from the other. AI Mode in particular is poorly isolated by automated tooling, which is part of why the manual pass described below exists. Claude is on the panel because it is a surface prospective clients genuinely use for research-shaped legal questions, and because we are not willing to market a surface we don't measure.
The rule underneath the panel is that each engine carries equal weight and is always shown as a per-engine breakout. We don't blend the six into one number. A blended score lets a strong showing on one surface conceal a zero on another, and the zero is usually the more actionable fact. If your firm is named consistently in Perplexity and never in AI Overviews, that's not an average — it's a diagnosis, and it points at different work than the reverse pattern would.
We have deliberately not usage-weighted the panel. Weighting engines by market share would require a credible, current source for engine traffic share, and the figures circulating in this industry don't meet that bar. Until one does, equal weight is the honest default. The pillar this feeds is described in more detail in the VERDICT framework.
The prompt set is the measurement instrument
Which prompts you track determines what your visibility number means, so the prompt set is chosen before any measuring happens and then held still.
Three rules govern it. First, prompts are buyer-intent: the phrasing a person with a live legal problem actually uses with an assistant, not keyword strings and not the firm's own name. "Who should I call after a rear-end collision in Sacramento" is a tracked prompt. "Sacramento personal injury lawyer" is a search query wearing a costume.
Second, the set is demand-weighted toward the practice areas and geography the firm actually wants cases in — a small, honest set beats a large, flattering one.
Third, and most important, the set is frozen. Once tracking starts, prompts are not swapped, tuned, or quietly retired. A prompt set that changes between reporting periods can produce any trend line you like, which makes it worthless as evidence. When a genuine change is needed — a new practice area, a market shift — it's added as a documented amendment with its own baseline, and the original set keeps reporting alongside it.
What actually gets recorded
For every prompt, on every engine, two things are logged: was the firm named, yes or no, and where did it appear relative to the other firms mentioned.
Named-or-not is the primary unit because that is the moment that matters to a prospective client — the assistant either puts your firm in front of them or it doesn't. Position is secondary but useful; being named third in a list of five is a different position than being the single firm mentioned, and the two move differently as work lands.
Naming is judged strictly. The firm has to be identifiable as itself. A generic recommendation to "consult a personal injury attorney in your area" is not a citation, and neither is a mention of a directory page the firm happens to appear on. Counting those inflates the number and teaches you nothing.
Why a human runs a pass every month
Automated AI-visibility platforms are useful and we use them. They are also incomplete, and the gap is not evenly distributed — coverage of Google's AI surfaces in particular lags what you can confirm by hand.
We know this concretely. On one client's tracked set, an automated platform reported zero presence on Google's AI surfaces for a reporting period during which manual pulls found the firm named in the top position. Had we shipped the automated read, the client would have received a report stating they were invisible on a surface where they were, in fact, leading.
So every month, a person runs the client's top prompts live across the panel and records the result by hand. The procedure is fixed: run each prompt, log named yes/no and position per engine, then reconcile against the automated read.
Where the two disagree, the manual read is authoritative, and the report carries a footnote saying so — which tool under-detected, which engine, how many prompts were manually confirmed, and on what date. There is a related rule: no surface is ever reported as zero without a manual confirmation that month. A zero is a serious claim about a firm's position, and it should never be an artifact of a tool's blind spot.
This costs roughly an hour per client per month. It is the least scalable part of what we do, and we have kept it because the alternative is publishing numbers we can't stand behind.
The numbers we refuse to report
A method is defined as much by what it excludes. Four things we don't produce:
A single composite visibility score. We considered one and rejected it. Collapsing volatile, per-engine citation data into one figure creates false precision — it looks authoritative and moves for reasons nobody can explain. Scores are reported individually.
A percentage without a denominator. "Visibility up 40%" is meaningless unless you can see the prompt count, the engine breakdown, and the baseline. Every figure we report carries the set it was measured against.
Traffic charts presented as citation evidence. Sessions and impressions measure something real, but they are not a measure of whether an assistant names your firm. Substituting one for the other is the most common way an SEO report gets relabelled as a GEO report.
Engine-behavior claims we can't source. There are several widely repeated statistics about AI search — the share of queries that trigger AI summaries, and a specific multiplier for how much certain engines favor fresh content — that trace back to vendor marketing rather than a checkable study. We don't put them in client materials. Where the underlying point is directionally sound, we state it qualitatively and say so.
How to run a lighter version yourself
You don't need an agency to get a first read. In about an hour:
- Write down eight to ten prompts a real client would type — problem-first, in your practice areas and your metro. Save the list; you'll reuse it.
- Run each one on ChatGPT, Perplexity, Google AI Overviews, Google AI Mode, Gemini, and Claude. Use a logged-out or fresh session so personalization and memory don't flatter the result.
- Record named yes/no and position in a simple grid — prompts down the side, engines across the top.
- Repeat the identical exercise in 30 days, same prompts, same engines, same conditions.
The first run gives you a baseline, and baselines are frequently uncomfortable — most firms score lower than they expect, and a score of zero across a well-built prompt set is common rather than alarming. The second run is where the information is, because it's the first comparison you can actually trust.
If you'd rather see the panel run against your firm without building the grid yourself, that's what our AI visibility audit does, and the free 60-second version is a starting point with no call attached. And if you're evaluating providers rather than doing it yourself, the questions worth asking are collected in the buyer's checklist — the shortest of which is the one this whole article answers: name your engines, and show me how you count a citation.