AI Mode SEO Tracking Tools: What They Can and Cannot Measure
There is now a whole product category selling AI Mode rank tracking. The pitch is familiar and comfortable: the same dashboard you already understand, with a new column for Google’s AI answers. Before you buy into that comfort, one number is worth sitting with.
When SE Ranking ran 10,000 US keywords through AI Mode three separate times on a single day — 20 June 2025, signed out — the average overlap of exact URLs between the three sets of results was 9.2%. In 21.2% of cases there were no overlapping URLs at all. Domain-level overlap was 14.7%. Same query, same day, same method, and roughly nine in ten of the cited URLs changed.
That is the fact any AI Mode tracking product has to survive. This article is about what such a tool can honestly tell you, what it cannot, how the three observation methods differ, and what to ask a vendor before you pay. It is deliberately not a roundup — if you want the suite-versus-specialist question, that is covered separately; if you want the general problem of reading AI visibility measurement, that is here; and if you want to see AI-referred traffic in analytics, that is a different job again.
Why AI Mode is not a SERP with extra steps
AI Mode is Google’s conversational answer surface, and at I/O on 19 May 2026 Google said it “has surpassed one billion monthly users, with queries more than doubling every quarter since launch,” alongside upgrading it to Gemini 3.5 Flash as the default model globally. It is not a fringe surface any more. But it is assembled differently from a results page, and every difference degrades the meaning of the word “rank”.
From the same SE Ranking dataset — 122,617 links across 9,734 responses, citing 22,235 unique domains:
- Each response cites many sources, not ten. “On average, each AI Mode response includes 12.6 URLs (sources).” There is no first position in a meaningful sense; there is a set.
- Most citations are not in the text. 90.8% were block links placed in a separate area, 8.9% were in-text links, and 0.3% mimicked organic results. Where a link sits changes whether a human ever sees it.
- It barely resembles the organic results for the same query. Average exact-URL overlap with the top 10 organic was just 14%; with the top 20, 12%.
- It is not the same as AI Overviews either. Exact-URL overlap between AI Mode and AI Overviews responses was 10.7%, domain overlap 16%.
That last point deserves emphasis because it breaks a common assumption. AI Overviews and AI Mode are different surfaces producing substantially different citation sets. A tool that reports “AI visibility” without telling you which surface it observed is reporting a blend of two things that agree with each other about a tenth of the time. For the mechanics of the surface itself — query fan-out, how it differs from AI Overviews — we cover that in the AI Mode guide.
A necessary caveat on that study. It was published on 29 August 2025 using data collected in June 2025, which means it predates this year’s model changes to AI Mode. We looked for an equivalent 2026 replication with a disclosed sample and method and did not find one. Treat the figures as direction rather than as a current measurement — and note that the underlying mechanism producing the volatility, an answer assembled fresh from multiple sub-queries, has not changed.
What this does to the word “rank”
If nine in ten cited URLs change between runs, then a single observation carries almost no information. The only statistic that survives is appearance frequency across repeated observations — how often you show up in N runs of the same prompt — and even that needs to be read as an estimate with wide error bars.
Work through what that means. Suppose a tool checks your prompt ten times over a month and you appear three times. Here is what you may and may not say:
| You can honestly say | You cannot say |
|---|---|
| “We appeared in 3 of 10 observations of this prompt in July.” | “We rank third for this prompt.” |
| “Our appearance rate moved from 1-in-10 to 5-in-10 over three months across a fixed 20-prompt panel.” | “Our AI Mode visibility improved 40% last month.” (from ten observations) |
| “Competitor X appeared in 8 of 10 observations; we appeared in 3.” | “Competitor X outranks us in AI Mode.” |
| “We are cited most often on the qualifier prompts, least on the head prompts.” | “Our average position in AI Mode is 4.2.” |
Three appearances in ten runs is a genuinely uncertain estimate — the honest reading is “somewhere in the rough vicinity of one in three, and do not bet a strategy on the difference between 3 and 5.” Which produces a simple test for any tool or agency report you are handed: if it gives you a position number for AI Mode, ask how it was derived. If the answer is “the order the links appeared in one response”, that is a rank on a set that will be mostly different an hour later.
| Metric | Meaningful for AI Mode? | Why |
|---|---|---|
| Appearance frequency over N runs | Yes, with N disclosed | The only metric that accounts for run-to-run variance. Useless without N. |
| Share of voice across a fixed prompt panel | Yes, if the panel is fixed | Comparative and directional. Changing the panel resets the series. |
| Which competitors co-appear | Yes | Stable enough to be useful, and genuinely actionable. |
| Which of your pages gets cited | Yes | Tells you what content is doing the work. Cross-check with Search Console. |
| Average position | No | Position within a 12.6-link citation block, 90.8% of which sits outside the answer text, is not a rank. |
| Day-to-day change | No | Almost entirely noise at this variance. Monthly at the finest. |
| A single “AI visibility score” | Only if decomposable | If you cannot see the observations behind it, you cannot tell improvement from noise. |
The three ways a tool can observe AI Mode
Every product in this category uses one of three methods, and the differences matter more than the interface does.
1. Google’s own report — free, official, deliberately limited
Search Console includes a generative AI performance report showing “data about how your site performs in generative AI features on Google Search,” covering AI Overviews and AI Mode. It is the only source in this list that reports what Google itself recorded rather than what a third party observed.
Its limits are explicit in the documentation, and you should know all of them before quoting it in a board pack. Impressions only — defined as “how many times links to your site were shown to a user in a generative AI feature on Google Search” — with no clicks and no click-through rate. Grouping by pages, countries, dates (Pacific Time) and devices, and no query data at all. No split between AI Overviews and AI Mode. It excludes Search Labs and Discover. And access is staged: “We’re rolling out this report to a subset of website owners, allowing for thorough testing before rolling it further.”
Despite all that, it is where every Singapore business should start, because it is free, it is Google’s own count, and it is not subject to sampling. Set up Search Console first if you have not.
2. Automated collection from the live interface
This is what most paid AI Mode trackers do. AI Mode has no public API, so a tool that wants to see it must drive a browser and read what comes back. That is a legitimate engineering approach and it captures what a user would actually see — which is exactly why the observation conditions are the whole ballgame:
- Signed in or signed out? Personalised responses differ. The SE Ranking study used non-logged-in accounts specifically to control for this.
- Which location? A Singapore result set and a US result set are different data. If a vendor cannot confirm Singapore-localised collection, its numbers describe a market you do not sell in.
- How many runs per prompt, per period? At 9.2% run-to-run overlap, one run per week is close to a coin flip. This single number determines whether the product measures anything.
- What happens when the interface changes? Browser-based collection breaks when the page changes. Ask what the gap in your history looks like when that happens.
3. Model APIs — which are not AI Mode at all
A tool can query Gemini’s API and report the result. That is a useful thing to know, but it is not AI Mode. AI Mode is a product built on a model, with its own retrieval, its own fan-out, its own grounding and its own interface. Results from an API call and results from the live surface are not interchangeable, and any vendor implying otherwise is selling you a proxy without saying so.
| Method | Cost | What it genuinely tells you | Main weakness |
|---|---|---|---|
| Search Console generative AI report | Free | Google’s own impression count for your site across AI Overviews and AI Mode | No clicks, no queries, no surface split, staged rollout |
| Automated live-interface collection | Paid | Whether you appear for chosen prompts, and who appears with you | Sampling and locale assumptions are usually undisclosed; breaks on interface changes |
| Model API querying | Paid or metered | What the underlying model says | Is not the AI Mode product; not a substitute for observing the surface |
| Your own manual panel | Free, ~10 min/month | Exactly what a signed-out customer in Singapore sees | Does not scale beyond a few dozen prompts |
Eight questions to ask before you pay
This category is young, crowded and full of near-identical products with near-identical marketing. The roundups ranking them are, for the most part, lead generation. A short list of questions cuts through it faster than any comparison table:
- How many times do you run each prompt per reporting period? If they will not say, stop here.
- Do you report variance, or only an average? A tool that hides the spread is hiding the only thing that makes the average readable.
- Signed in or signed out, and from which location? Ask specifically whether Singapore-localised collection is available.
- Do you separate AI Mode from AI Overviews? They agreed on only 10.7% of exact URLs in the study above. A blend is not a measurement of either.
- Can I export the raw observations, not just the dashboard? Raw observations are what let you check the vendor’s arithmetic and take your history with you.
- How is any composite “visibility score” calculated? If it cannot be decomposed into observations, it cannot be audited.
- What happens to my history if Google changes the interface? Ask about past outages, not hypothetical ones.
- What is the shortest commitment? In a category this young, a twelve-month contract is a bet on a product roadmap you have not seen.
What a Singapore business should actually do
Local context changes the priority order here, and mostly in a helpful direction. StatCounter put Google at 92.46% of Singapore search in July 2026, with Bing at 3.37%. That means AI Mode and AI Overviews are, for most Singapore businesses, the AI surfaces that matter — more so than ChatGPT, which dominates the international conversation. It also means the free Google-native report covers the majority of your exposure.
A sensible sequence:
- Turn on the free report and read it for two months before buying anything. Impressions with no query data is a weak dataset, but it is Google’s own count and it is free.
- Run a fixed 12-to-20-prompt panel by hand, monthly, signed out, from Singapore. Log appeared / did not appear and who else appeared. This is your control group, and it stays useful even after you buy a tool.
- Only then consider paying, and only if you can name the decision the data will change. Good reasons exist: multiple locations, a high-consideration purchase where being named in an answer matters, or a competitive category where you need to watch someone specific. “Because we should be tracking AI” is not one.
- Whatever you buy, keep reporting it monthly, never daily, and keep the observation count on the same slide as the result.
If your reporting has a habit of promoting numbers that feel like progress without driving decisions, our piece on vanity metrics applies to this category with unusual force. And when AI-referred sessions do start arriving, a correctly configured GA4 is what turns them from an anecdote into a number.
A 30-day protocol you can run before spending anything
| Week | Action | Output |
|---|---|---|
| Week 1 | Write a fixed panel of 15 prompts: 5 category, 5 qualifier, 3 comparison, 2 problem. Run all 15 in AI Mode, signed out, from Singapore. Log appeared / not, plus who did. | Observation set 1. Do not draw conclusions. |
| Week 2 | Run the identical 15 prompts again, same conditions. | Observation set 2. Compare with set 1 — the difference is your personal volatility baseline. |
| Week 3 | Run again. Open the Search Console generative AI report and record impressions and top pages. | Observation set 3 plus Google’s own count for cross-reference. |
| Week 4 | Run again. Calculate appearance frequency out of 4 per prompt. Note which prompts were stable and which flipped. | A defensible baseline, and a shortlist of prompts worth paying to track. |
Four runs is a small sample, and it will not settle anything with statistical confidence. That is the point of doing it: you will learn how noisy your own prompts are before a dashboard smooths that noise into a reassuring line.
The verdict
AI Mode tracking tools can tell you something worth knowing: whether you appear for the questions your buyers actually ask, how often, and who appears alongside you. That is real, and for some businesses it is worth paying for. What they cannot honestly give you is a rank, a precise trend from a handful of observations, or a single score you can put in a report without a footnote.
The category will mature, and the honest vendors will be the ones publishing their observation counts and locales rather than their scores. Until then, the sequence that protects you is the cheap one: free report, fixed manual panel, then buy only against a decision you can name. A tool bought before you have a baseline cannot be evaluated, because you have nothing to check it against.
If you would rather have AI visibility measured properly alongside the rest of your search programme — with the observation counts on the page and no invented positions — our AI SEO service is built that way, our AI SEO guide is the wider map, and our case studies show the reporting standard we hold ourselves to. If you are a smaller business wondering whether any of this deserves your time yet, start here instead.
Frequently asked questions
Can you actually track rankings in Google AI Mode?
Not in the sense of a position number. When SE Ranking ran 10,000 keywords through AI Mode three times on one day, the average exact-URL overlap between runs was 9.2%, and 21.2% of keywords shared no URLs at all. The only defensible metric is appearance frequency across repeated observations, reported with the number of observations alongside it.
Is the free Search Console report enough for AI Mode?
It is the right starting point and it is not sufficient on its own. The generative AI performance report covers AI Overviews and AI Mode, but reports impressions only, with no clicks, no query data, and no split between the two surfaces, and it is still rolling out to a subset of site owners. Pair it with a manual prompt panel and you have a workable free baseline.
What is the difference between an AI Mode tracker and an AI visibility tool?
Scope and surface. AI Mode tracking observes one Google surface specifically. Broader AI visibility tools blend several assistants together, which is useful for brand monitoring but blurs an important distinction: AI Mode and AI Overviews agreed on only 10.7% of exact URLs in SE Ranking’s study, so a blended figure is not a measurement of either one.
How often should an AI Mode tracking tool check my prompts?
Often enough that the volatility averages out, and reported no more frequently than monthly. At roughly 9% run-to-run URL overlap, a single weekly check is close to a coin flip. Ask any vendor how many runs per prompt per period they perform, and treat an unwillingness to answer as the answer.
Does AI Mode matter more than ChatGPT for a Singapore business?
In most cases yes. StatCounter put Google at 92.46% of Singapore search in July 2026, and Google reported AI Mode passing one billion monthly users in May 2026. The AI surface most of your local customers encounter sits inside Google, which also means the free Search Console report covers a large share of your exposure.
Should I buy an AI Mode tracking tool now or wait?
Wait until you have a baseline and a decision the data will change. Run a fixed panel of 12 to 20 prompts by hand for a month, read the free Search Console report for two, and then buy only if you can name what you will do differently with more frequent data. Multi-location businesses and high-consideration purchases have the strongest case; short commitments are sensible in a category this young.



