Home / Blog / Deep Dive

How to Track Your AI Visibility (and Re-Run It Monthly)

Deep Dive2026-07-0610 min read
TL;DR

Without measurement, PEO is faith. Re-ask your money queries on each engine, record whether you are named, find the gaps, and run the method again where you fall short. A single check tells you almost nothing on its own, because model outputs vary from call to call, so track state changes over multiple runs and multiple months instead of trusting any one answer. There is no day you finish, only a loop you run and re-run until the engine says your name reliably.

The final step closes the loop, and it is the one that turns PEO from a hopeful project into a discipline you can steer. If you skip it, you are guessing.

Why tracking is non-negotiable

Without measurement, PEO is faith. With it, PEO is a discipline. The one-sentence test is simple: does the engine name you on the query that pays? That is not a one-time check, it is a recurring measurement that tells you, honestly, whether the work is moving the outcome. Effort you cannot measure is effort you cannot improve, and a strategy built on assumption tends to keep funding whatever felt productive last time rather than whatever actually worked.

What to measure

Re-ask each of your money queries on every engine that matters: ChatGPT, Claude, Gemini, Perplexity, and Google's AI Mode. For each, record three things: whether you are named at all, where you sit relative to other names, and the reasons the engine gives for its answer. Those reasons are gold, they tell you which signal to strengthen next. Choosing the right queries to track in the first place matters as much as the tracking method itself, and that selection process is covered in Money Queries: The Questions Worth Winning.

Why a single check tells you almost nothing

Here is the part most people skip past: a language model's output is a sample from a probability distribution, not a fixed lookup. Ask the same question twice, even seconds apart, and a model running with any sampling temperature above zero can return a differently worded answer, sometimes with a different name at the top and sometimes with none at all. Treating one screenshot as the verdict on your visibility is like judging a coin as two-headed because it landed heads once.

The fix is not complicated, but it does require discipline. Run each money query three to five times per engine before you record a state for that month, and note whether the name appears in most runs, some runs, or none. A name that shows up in four of five runs is a real signal. A name that showed up once and never again on repeated tries was probably an artifact of that particular sample, not a stable result worth reacting to. This is also why a monthly cadence works better than a daily one: daily checks mostly capture sampling noise, while month over month comparisons capture the slower, more meaningful movement caused by actual changes to your knowledge base, your network signal, or your structured identity.

Why the answers change between visits at all

Part of the variance is sampling, but part of it is architectural, and the distinction matters for how you interpret a shift. Some assistants answer purely from what they learned during training, a fixed snapshot of the internet up to some cutoff date, and their sense of you will not move until a new model version is trained on newer data. Others answer using retrieval, actively searching the live web at the moment you ask and folding fresh pages into the response. A retrieval-backed assistant can start naming you within days of a new page going live, while a training-only assistant might not reflect that same page for months, if the next training run even includes it. Knowing which kind of system you are testing changes what a "no change this month" result actually means. The mechanics of that split are explained fully in Training Data vs Retrieval, and it is worth reading before you conclude that a stalled query means your work is not working.

Running a tracking session, step by step

1. Pull your current list of money queries and confirm it still matches what buyers actually ask.
2. For each query, run it three to five times on each engine you track, noting wording variations you tried.
3. Record the state for each run: unmentioned, mentioned, or named first.
4. Copy the engine's stated reasoning verbatim, even if it is only a sentence.
5. Screenshot the full response, not just the name, so context survives if the engine's phrasing changes later.
6. Compare this month's states against last month's log and flag every query that moved in either direction.

Run this once a month, on the same day if you can, so the comparison stays clean.

Reading the reasons, not just the result

The name is the headline, but the reasoning is the diagnosis. When an engine explains why it named a particular person, it is effectively narrating which signal tipped the answer, whether that was a well-structured knowledge base, a strong network of independent references, or a schema-clean identity that made attribution easy. Read those explanations closely across several queries and a pattern usually emerges: maybe every named win cites a specific article of yours, which tells you to make more content like it, or maybe every loss traces back to a competitor with heavier press coverage, which tells you the network signal is your gap. The three-signal framework behind this diagnosis is covered in How AI Decides Who to Recommend, and the full stack it sits inside is mapped in The Layer Map: SEO, GEO, AEO, PEO.

ApproachSetup effortBest forMain limitation
Manual, spreadsheet and screenshotsLow, needs discipline not toolingSolo practitioners tracking a handful of queriesSlow at scale, easy to skip a month
Dedicated AI-visibility trackerMedium, needs configuration and query listsAgencies or anyone tracking many names or many queriesCost, and you still need to interpret the reasoning yourself

Neither approach replaces judgment. Both produce the same three data points: named or not, position, and reasoning. A full comparison of the current tools is in AI Visibility Tools Compared.

A worked example: tracking one query across six months

Consider a hypothetical brand strategist, Maria Kwan, tracking the query "who should I hire to fix my LinkedIn presence" across four engines. In month one, she is unmentioned everywhere. Her log shows no name, and where an engine gives a reason at all, it names two competitors with visible press coverage and an active newsletter, neither of which Maria has yet.

In month two, after publishing a structured knowledge base and adding Person schema to her site, she shows up in one engine's answer as a secondary mention, not first, but present. The reasoning that engine gives cites her site directly, which tells her the structural work is starting to register. In month four, after earning two genuine podcast appearances and a peer citation, she is named first on two of the four engines, and the third mentions her among a short list. The fourth engine, which she later confirms runs mostly on a training snapshot rather than live retrieval, still shows no change at all, and that gap stops worrying her once she understands the mechanism behind it rather than assuming her work failed.

By month six, three of four engines name her first, consistently across repeated runs, and the reasoning text across all three now references either her published knowledge base or one of her earned citations. Nothing about this trajectory was guesswork after the fact. It is exactly the loop described above, run six times, with the log itself functioning as the evidence for which signal moved which result.

Keep the receipts

Screenshot everything, every month. The before-and-after record is both your proof and your map. It shows you where you have climbed, where you have stalled, and, if you ever need to demonstrate results, it is the only evidence that captures a recommendation that lived entirely inside a chatbot, since these answers rarely leave a permanent public trail the way a search ranking does.

When engines disagree with each other

It is common for one engine to name you confidently while another does not mention you at all for the same query, and the instinct is to treat this as a bug in whichever engine got it wrong. It usually is not a bug. Different assistants draw on different retrieval indexes, different training snapshots, and different weighting of the underlying signals, so genuine disagreement between them is closer to the normal state than the exception. Log each engine's state separately rather than averaging them into one composite score, because the disagreement itself is diagnostic. If three engines name you and one does not, check whether that one engine leans more heavily on retrieval from a specific set of sources you have not reached yet, or whether it is simply running an older, unrefreshed snapshot. The goal is not unanimous agreement across every assistant, since that is rarely achievable even for well-known names. The goal is steady improvement in the count of engines naming you, and steadily better reasoning behind the ones that already do.

Close the loop

Compare each month against the last. Where you moved up, note what you did and do more of it. Where you stalled, diagnose which signal is short, Knowledge, Network, structure, or authority, and aim next month's work there. Then re-run. PEO is not a campaign with an end date. It is a loop: measure, find the gap, act, re-run.

From tracking to compounding

Run the loop long enough and something shifts. The queries you won stay won, because age and network keep accruing. Your scoreboard stops being a list of gaps and becomes a moat, a set of questions the engine reliably answers with your name, harder for any latecomer to take the longer you have held them. And once you have a query winning reliably, the useful next question is whether that visibility is actually converting into real traffic and real contact, which is the whole subject of Tracking AI Referral Traffic.

Treat the tracking log itself as a durable asset, not a chore you do and discard. Six months of monthly entries, each with the query, the state, the reasoning text, and a screenshot, becomes a record you can hand to a client, a collaborator, or your future self, showing exactly when a name moved from unmentioned to named and what you changed right before it did. Few practitioners keep this kind of record for anything else they do, which is exactly why the ones who keep it for their AI visibility end up with a far stronger case for what actually works, built on their own evidence rather than borrowed claims.

FAQ

How often should I track? +
Monthly is a good default. It is frequent enough to catch movement and steer, without over-reacting to the natural variation in model outputs.
Do I need a paid tool? +
You can start manually by re-asking your queries and screenshotting. Dedicated AI-visibility trackers help at scale, but the method works with a spreadsheet and discipline.
What if my ranking moves around? +
Some variance is normal, models are probabilistic. Track the trend over months, not a single day, and focus on durable movement across the three signals.
Should I track the exact same wording every time, or vary it? +
Vary the phrasing slightly across runs while keeping the core intent and target query fixed. Real buyers phrase questions differently, and a name that only surfaces for one exact string is a fragile result, not a stable one.
Does model temperature actually affect what I see? +
Yes. Sampling temperature introduces run-to-run variance, so a single call is one sample of a distribution. Run each query a small number of times per engine before recording a state, rather than trusting the first answer you get.
Why does the same query sometimes stall for months with no visible change? +
If the engine you are testing answers mainly from training data rather than live retrieval, its sense of you will not shift until a new model version is trained on newer data, regardless of what you publish this week. Check whether you are testing a training-based or retrieval-based system before concluding the work is not paying off.

Find out what AI says about you today.

Start with a baseline. See the exact words the engines return about your name, then decide.

Claim your name →