Kaivox
KaivoxMethodology

How Kaivox measures AI visibility

A score you can’t explain is a score you can’t defend, and these numbers go in front of the people who matter, your leadership or your clients. This page lays out exactly what we measure, how we measure it, and where the edges of the measurement are, so you can stand behind every number you show.

01 · Surfaces

We measure the assistants buyers actually use

Buyers don’t ask the benchmark champion. They ask the assistant they already have open, in its default mode. So Kaivox probes the consumer-default tier of each assistant, not the heavyweight models a buyer has to go looking for.

ChatGPTopenai/gpt-5.4-mini

The fast default tier that consumer ChatGPT routes everyday questions through. We measure the default because that is the answer a buyer who just types a question gets.

Geminigoogle/gemini-3.5-flash

The Gemini app default. Flash answers the shopping questions; the Pro tier is opt-in.

Claudeanthropic/claude-sonnet-5

Claude.ai serves the Sonnet tier by default. Here the consumer default genuinely is the strong model.

Perplexityperplexity/sonar

Free Perplexity runs Sonar, search-grounded by nature. It exists only in the live layer below.

This lineup has been the measured panel since June 1, 2026. Every scan records exactly which models answered it, and any change to the lineup or to the scoring ships as a new measurement version, listed in section 09 below with its date, so a moving trend line can always be traced to either the market or the instrument.

02 · Layers

Two layers: what AI remembers, what AI finds

Every assistant is measured two ways, because an AI answer has two sources and they move at different speeds.

Memory

What the AI remembers about you

The model answers from its training data alone. This layer reflects your long-term footprint on the web and moves slowly, when models retrain. It is the answer buyers get when the assistant doesn’t bother to search.

Live

What the AI finds when it looks

The same model with live web search, which is how the consumer apps increasingly answer buying questions. This layer moves when content gets published, in weeks rather than quarters, and it is where a fix shows up first.

A gap between your memory score and your live score is a finding, not noise: strong memory with a weak live score means your fresh content is invisible to retrieval; strong live with weak memory means AI finds you but doesn’t yet know you. The two diagnoses have different fixes, which is the point of measuring both.

03 · Questions

The questions are real buying questions

Each scan asks the assistants what a real buyer in your category would actually type, and the questions are shaped to your business. A national brand is measured on discovery and comparison questions like “best [category] for [need]” or “[rival A] vs [rival B]”. A local business is measured the way people really search for one: “who should I hire for [service] near me,” not a national best-of list. Across both, the set spans discovery, head-to-head comparisons against the rivals that truly compete with you, alternatives, problem-first questions, and trust checks.

We don’t invent these questions in a vacuum. They are grounded in real demand in your category: the keywords people actually search and the “People Also Ask” questions Google surfaces, so the panel reflects how buyers really phrase things rather than how we imagine they do. Two hard rules hold throughout: the questions never name your brand, because the entire point is whether the AI brings you up on its own, and they are grounded in today’s date, so a search-backed assistant is never steered into last year’s articles.

You see the full question set, the exact wording of every question, and approve it before any scan runs, and every scan asks that same approved set. That is what makes your score comparable scan over scan: when the number moves, it moves because the assistants’ answers changed, not because we asked different questions. And when the set does change, because you edited it, the app says so on the score and stops comparing across the change rather than passing the difference off as movement. (The free check picks its own questions and shows them on the result; the approval step belongs to the full scan.)

04 · Confidence

Asked up to five times, because once is a coin flip

The same question asked twice can return two different brand lists. A single-sample score cannot tell a real change from measurement jitter, so every question is asked up to five times per assistant (we stop early only after at least two asks, and only when every answer so far agrees on whether you were named and how prominently; a question counts the same in your score however many times we asked it), and your score is reported with a repeatability band, for example 27.8 ± 2.0. That band is calibrated the hard way: we periodically re-scan brands back to back with nothing changed and measure how far the number moves on its own (the current floor comes from three back-to-back scans of one unchanged brand; same-day repeats on three more brands came in tighter, so the published band is the conservative one), and the band you see is never allowed to claim more precision than those controlled repeats actually showed. When a scan’s own spread comes in narrower than that calibrated floor, we show the floor and mark it “calibrated minimum band”, which is why two different brands can carry the same band. The band is one scan’s wobble; the gap between two scans can wobble by roughly twice that, which is where the five-point rule comes from: a genuine move on the 0-100 score is one that clears roughly five points or holds up across scans, not a wobble inside the band. The app says the same thing in place: a small score move is labeled “within normal variation” right beside the number. Single-sample measurement dressed up as precision is the most common shortcut in this category; sampling every question, calibrating against real repeat runs, and showing you the band is how you can tell we’re not taking it.

The size of the instrument, in plain numbers: a full scan asks about 40 questions, each of them up to five times on each of seven assistant-and-mode combinations (three assistants asked from memory and again with live search, plus Perplexity, which always searches). On a recent full scan that came to 38 questions and 568 graded answers. It would be cheaper to ask hundreds of questions once each. We spend the budget on the repeats and on the calibrated noise floor instead, because the second ask is what tells a real change from a wobble, and a trend you cannot trust is not worth having more of.

05 · The score

What the score means

The Kaivox AI Visibility Score (0 to 100) blends four measured signals: how often you appear at all, how prominently you appear when you do (leading an answer counts for more than a passing mention), your share of voice against the rivals the assistants actually name, and how often your own site gets cited as a source. Each assistant is scored separately, then blended. No signal in the score is assumed, padded, or simulated; if we didn’t measure it, it isn’t in the number.

The blend is not a flat average. Each assistant’s influence on your score reflects how many real people actually use it, dampened so no single assistant dominates the number and smaller assistants still register; ChatGPT carries the largest share because it carries the most use. When we recalibrate these weights against published usage data, the change ships as a new measurement version (section 09), every scan carries its version stamp, and no trend line ever presents a recalibration as market movement.

The score is built only from questions where the assistant actually recommends a business. Pure how-to and cost questions, where the AI explains something and names no one, are measured separately as Answer Visibility, so they never drag your recommendation score down. And when an assistant declines to name anyone at all, that answer is set aside rather than counted as you being absent, because in some categories the assistant routinely refuses to name names. The number reflects the questions where being recommended was genuinely on the table.

06 · Corrections

Your corrections outrank our scan

Kaivox reads your brand from the public web, and the public web is sometimes wrong about you. Anything you correct, from your category to the names your brand goes by, becomes permanent ground truth: it overrides what we scanned, survives every rescan, and steers every future measurement. The system also audits its own data after every scan and shows you what looks wrong, so finding problems doesn’t depend on you stumbling into them.

07 · Proof

How a fix is proved

A score tells you where you stand. The proof loop tells you whether something you published changed it. When Kaivox publishes a page or post against one buyer question you were losing, it records a baseline that day: the same question, asked twelve times of each assistant, under the same conditions as the scan. Fourteen days later it asks the same question the same way, twelve times again, and compares the two ends.

Two things stop a coincidence from being sold as a win. First, two untouched questions from the same scan, questions the content did not target, are re-asked alongside it at both ends: if those moved too, the market moved, and the lift is not credited to your content. Second, a noise test: an exact permutation test at the five-percent level, which asks how often a gap this large would appear if the content had changed nothing. A move is credited only when it is positive, it clears that test, and the untouched questions stayed flat.

The verdicts are the ones the instrument can actually give: a real, attributable lift; a real lift that is small (under five points on that one question: it counts as a win, but on its own it is not enough to say “do more of this”); a move that the untouched questions also made; a change within measurement noise; and a real drop. A lift that is not proved at day 14 is checked again at day 28, and a proved lift is re-checked at day 28 to see whether it held. The fixes that did not work stay on the board next to the ones that did.

Two limits, stated plainly. The number that moves is the score for that one question, not your overall visibility score. And the range shown beside a result is the spread of the same-day re-asks at each end, not a measure of how sure we are that your content caused the change; the untouched questions are what carry that weight.

08 · Limits

What this measurement is not

  • 01

    It is not every conversation. AI answers vary by user, location, and chat history. We measure a representative, repeatable slice under controlled conditions, which is what makes scan-over-scan comparison meaningful.

  • 02

    It is not comparable across tools. Different products probe different models with different questions; a Kaivox score is built to be compared with your last Kaivox score and with competitors measured in the same scan, not with a number from another vendor.

  • 03

    It is not static. The assistants change under everyone constantly. Our lineup discipline, date grounding, and repeatability bands exist so that when your trend moves, you can trust the movement is yours.

  • 04

    It is not a margin of error. The plus-or-minus band is a repeatability band, not a statistical confidence interval: it is never narrower than the floor our controlled repeat runs demonstrated, and it widens when a scan’s own spread exceeds that floor. In a small or niche market, once off-topic questions and refusals are set aside, a score can rest on a modest number of answers, so treat a small move as noise and a sustained trend as signal.

  • 05

    It is not a cross-brand leaderboard. The 0-to-100 score is built to track ONE brand over time and against the rivals named in the same scan. Comparing the raw score of two unrelated brands, or two different categories, is not something the scale is built to support.

  • 06

    It is not immune to the models changing under their own names. We hold the assistant lineup fixed and compare like-for-like scans, but when a provider updates a model without renaming it, a score can shift for reasons unrelated to you, which is why a sustained trend matters more than any single scan-to-scan move. We review the lineup on a fixed quarterly schedule against what the consumer apps actually answer with, and every lineup change is batched into one published measurement update, never one per model, so your trend is cut as rarely as honesty allows.

  • 07

    It is not pinned to your customer’s exact location. When an assistant searches the web live, that search runs from our infrastructure, not your buyer’s town. Local questions name the place so the answers stay locally relevant, but the search origin is not your customer’s address.

  • 08

    It is not a transcript of the consumer apps. We ask each assistant through its business interface, not by driving the consumer app, with a neutral setup and no chat history. Published research shows the consumer apps can answer differently for different people, which is exactly why a controlled, repeatable instrument is the right way to measure change: your trend moves because the assistants moved, not because the measurement did.

09 · Changes

What changed, and when

Every scan is stamped with the version of the measurement that produced it. This is the current version and every change before it, in plain language. Your stored scans never change; only new scans use the new version, and when a trend line crosses one of these dates the app says so beside the number and does not read the step as market movement. Most changes so far tightened the number: harder to inflate, not easier.

CurrentMeasurement version 12since September 8, 2026

  1. An answer that only repeats the rival a question asked about no longer counts against youv12 · September 8, 2026

    When a question names another brand, the assistants name that brand back almost every time. An answer that named only the brands the question handed it is now set aside and counted apart, because there was no room in it for anyone else; an answer that brought in any other name still counts, because that name could have been you. A rival named in a question no longer earns share of voice for being named back in the answers to it, and the competitor table says how often each rival was asked about by name. Rivals your questions ask about by name will show a smaller share than before; the brands the assistants volunteer on their own will show a larger one. The Claude assistant in the lineup now answers with the model the Claude app has used by default since June 30.

    On your trend line: Scores from before this date can sit lower than the same brand measured after it, and share-of-voice numbers are not comparable across the date; the difference is the instrument, not the market. Scans on the new method are compared only with each other.

  2. A link to your site counts only when the page loadsv11 · August 28, 2026

    Assistants sometimes invent page addresses on real sites. Each link to your own site in an answer is now checked once per scan, and a link to a page that does not exist (a confirmed 404) is no longer counted as a citation. A link that could not be checked, for example because the site blocked the check, still counts exactly as before, and the citation tile says when links were not checked.

    On your trend line: This can only lower the citation part of a score, never raise it. Mentions, prominence and share of voice are untouched.

  3. A brand name must start where the match startsv10 · August 10, 2026

    Your name is matched only at the start of a word. A different company whose name merely contains yours ("Pumpjack Golf" for "Jack Golf") no longer counts as you, and an ordinary word that happens to contain a short name no longer earns you a mention.

    On your trend line: This can only remove mentions that were never yours. A handful of brands read slightly lower at this boundary and that is the truer number.

  4. Every repeat of a question counts equallyv9 · August 8, 2026

    When a question is asked several times of one assistant, each real answer now carries equal weight. Before this, the order the answers happened to be stored in could give one answer four times the weight of another.

    On your trend line: Movements at this boundary were under half a point on the headline score across our whole fleet. Repeat scans of an unchanged brand agree more closely from here on.

  5. A fix for brands whose answers name almost no rivalsv8 · July 24, 2026

    When the assistants named almost no rival brand in your answers, the mention rate was being taken over the few answers that named someone, which could publish a score in the 70s for a brand named once in 36 answers. It is now taken over all the answers when rival signal is that thin.

    On your trend line: Only scans with almost no rival signal were affected. Everything else is byte-identical.

  6. More of your real mentions and citations are countedv7 · July 8, 2026

    Citations the assistant itself declared and that resolve to your website now count alongside links found in the answer text. Punctuation variants of your name (with or without periods) match. A numbered heading that names you carries its list position.

    On your trend line: Scores can step up at this boundary because credit that was always real is now counted.

  7. Honest floorsv6 · July 5, 2026

    A brand with no tracked rivals is scored on the same scale as everyone else. Share of voice is reported only when rival mentions were actually measured. A brand whose name is also an ordinary word is matched only in its exact cased form.

    On your trend line: This can only lower a score, and only for the least-evidenced scans.

  8. Rivals judged by where they operate; local questions name the placev5 · July 4, 2026

    A same-category business in the wrong state no longer counts as a rival in a local brand’s share of voice. New local question sets name the state. New national sets carry a few questions that name the brand, and those are kept out of the headline score.

    On your trend line: Local brands may have seen wrong-state rivals drop out of their share of voice at this boundary.

  9. Memory and live answers blended by how often each answeredv4 · July 4, 2026

    Within one assistant, the memory answer and the live-search answer are combined in proportion to how often each produced a usable answer, instead of half and half. A single borderline answer had been able to swing a score by several points between two identical scans.

    On your trend line: Repeat scans became steadier. The answers underneath did not change.

  10. Assistants weighted by real-world usev3 · July 4, 2026

    Each assistant’s share of your headline score now reflects how many people use it, dampened so no single assistant dominates and the smaller ones still register. Perplexity had been carrying far more weight than its reach.

    On your trend line: The blend shifted at this boundary; the answers underneath did not change.

  11. Cleaner name matching and citation countingv2 · July 2, 2026

    Your name is matched in its exact casing, so a common word no longer counts as your brand. A link that redirects to your site counts as a citation. An answer whose brand list could not be read is set aside instead of being counted as a miss.

    On your trend line: Scores from before this date can sit a little differently from the same brand measured after it.

  12. The first versionv1 · before July 2, 2026

    Every scan completed before July 2, 2026 is version 1.

Losing questions

The order of the questions you are losing carries its own version, because it can change while your score does not. These are the changes to how that order is worked out.

  1. Losing questions are ranked by what they are worth times your chance of winning themv3 · September 9, 2026

    Two changes. First, a question counts as one to win whenever the assistants had room for a name nobody asked about and did not use it for you, even if you are already named in some of the answers; before, only questions where you were never named counted. Second, each question now carries two numbers: what winning it is worth (the old impact number, unchanged) and your chance of winning it, which combines how often the answers had room and how often the assistants already name you on questions of the same kind. The list is ordered by worth times chance, so a valuable question the assistants have never named you on sits below a smaller one they already half-know you for. The chance term borrows one question’s worth of answers from your whole service line when a kind has few answers, so one lucky mention cannot swing it.

    On your trend line: Your visibility score and its trend line are unchanged. The order of your losing questions, and which questions appear at all, are not comparable to before this date; a question’s worth number is the same scale as the old impact number.

  2. Losing questions now count whether a slot openedv2 · September 2, 2026

    The order of your losing questions now counts whether the answers left room for a third name. When a question names a rival by name, the only way you could have appeared in it is if the assistants brought in a name nobody asked about, so we now measure how often that happened and rank the question by it. A question whose answers never bring in a third name still shows, and now sits last.

    On your trend line: Your visibility score and its trend line are unchanged. A losing question’s impact number is not comparable to the one it had before this date.

See your own numbers

The free check runs a light version of this methodology on your brand in about a minute. The full scan goes deeper on every axis described above.

Kaivox · brand intelligence for the AI search era