Home / Blog / Playbook

The 25-Prompt Audit: Test Your AI Visibility in One Hour

Playbook2026-07-1213 min read
TL;DR

Before you optimize anything, measure what the engines already believe. This audit is 25 prompts in five categories, identity, recommendation, comparison, trust, and citation, run in fresh sessions and scored 0 to 2 each. One hour produces a 0-to-50 baseline, a map of exactly where you fail, and a fix priority. The prompts are below, ready to copy, along with the reasoning that makes the scores trustworthy instead of flattering.

Most people have never asked an AI system the questions their buyers actually ask it. They optimize blind, publish blind, and then wonder why nothing moves. An audit costs an hour and replaces the guessing with a number you can act on.

Why audit before you optimize?

Two reasons carry the whole case for auditing first. First, diagnosis beats guesswork: the specific failure pattern tells you which layer is broken, and the layers have different repairs entirely. Wrong facts demand cleanup of the sources feeding the error. Absence demands publishing something citable. Presence without preference demands third-party proof, not more content. Spending months writing when your real problem is a contradictory identity, a case covered in Structure Your Identity for Machines, is the standard expensive mistake, and it stays invisible until you actually look at what the engine says.

Second, you need a dated before picture. What a generative engine surfaces about any given name shifts as new pages get published, old ones get deprecated, and the underlying retrieval index refreshes, which is a mechanical property of how these systems work rather than a fixed fact you can check once. See Training Data vs Retrieval for why an answer that was accurate last quarter can quietly go stale without anyone editing a single page. Without a dated baseline, you cannot tell whether a change in what the engine says was caused by your work or by the churn that happens regardless of what you do.

The rules that make results trustworthy

The 25 prompts

Category 1: Identity (prompts 1-5)

Does the engine know who you are, and is what it knows true?

  1. Who is [your full name]?
  2. What is [your full name] known for?
  3. What has [your full name] written, built, or published?
  4. Give me a short professional bio of [your full name].
  5. What is [your full name]'s professional background and career history?

Category 2: Recommendation (prompts 6-10)

Does the engine name you when nobody asked about you specifically? This is the money category.

  1. Who are the best [your specialty] consultants in [your city or country]?
  2. Recommend a [your specialty] expert for a [your target client type] that needs [the outcome you deliver].
  3. I need help with [the problem you solve]. Who should I talk to?
  4. List five respected independent [your specialty] specialists working today.
  5. Who is an up-and-coming expert in [your niche]?

Category 3: Comparison and shortlist (prompts 11-15)

When buyers narrow the field, do you survive the cut?

  1. Compare [your name] and [a direct competitor] for [the service you both offer].
  2. Give me a shortlist of three people for [a project in your niche] and explain each choice.
  3. What are the alternatives to hiring [a large firm or famous name in your space] for [the service]?
  4. Should I hire a big agency or an independent specialist for [your service]? Name specific options.
  5. Who does work similar to [a well-known peer you respect]?

Category 4: Trust and verification (prompts 16-20)

Buyers increasingly run a first pass of due diligence through an assistant before they ever pick up the phone, and verification questions make up a large share of that research.

  1. Is [your name] legitimate? What is their track record?
  2. What credentials and experience does [your name] have in [your field]?
  3. Are there any criticisms or red flags about [your name]?
  4. Has [your name] worked with companies like [your target client type]?
  5. Why would someone hire [your name] over other options?

Category 5: Source and citation (prompts 21-25)

Does your published thinking feed the answers in your niche, or does the engine cite everyone but you?

  1. What are the best resources for learning about [your core topic]?
  2. Summarize the main approaches to [the problem you solve] and cite your sources.
  3. What does [your name] say about [your signature topic]?
  4. Quote or paraphrase something [your name] has published about [your topic].
  5. Whose frameworks are most cited for [your niche], and where do those frameworks come from?

Customizing the brackets: two worked examples

The brackets are where audits go soft, so here is what honest customization looks like. Take prompt 7 for a hypothetical fractional CFO named Jane Okafor who serves venture-backed startups: "Recommend a fractional CFO for a Series A SaaS company that needs investor-ready reporting." Every bracket is filled with buyer language, no jargon, no flattery. The same prompt for a hypothetical employment lawyer named Tomas Reyes reads: "Recommend an employment law specialist for a UK retailer facing a tribunal claim." Notice neither version mentions the person being audited. The recommendation category only tells you something real when you ask exactly as a stranger would, with no hint that you already know or favor a particular name.

Now prompt 23 for the same two people. For Jane: "What does Jane Okafor say about burn-rate discipline in a downturn?" For Tomas: "What does Tomas Reyes say about settlement versus tribunal strategy?" If you cannot fill the signature-topic bracket without hesitating, that hesitation is itself a finding. It means you have not claimed a topic tightly enough for an engine to associate you with one, and no amount of visibility work elsewhere fixes an unclaimed position. The fix for that specific gap lives in Answer-Shaped Writing, since a signature position only becomes citable once it is written in a form an engine can lift cleanly.

Logging: the spreadsheet that makes it an instrument

One row per prompt per engine, six columns: date, engine, prompt number, verbatim answer, score, and names mentioned before yours. That last column is the quiet gold. Over three or four monthly rounds it becomes a working census of who the engines actually consider your competition, which rarely matches who you consider your competition. Total each category, keep a running chart of the five subtotals, and resist the urge to prettify anything. This document is an instrument, and instruments are allowed to be ugly.

Scoring: turn answers into a number

Score every prompt 0, 1, or 2. Be harsh, a generous audit is a useless audit.

Maximum 50. Read your band against the table below.

Score rangeReadingWhat it means
0-10InvisibleThe engine cannot see you or cannot resolve you as one entity. Start with identity, not content.
11-25Legible, not preferredYou exist as an entity but nothing pushes the engine to choose you over the field.
26-40ContenderYou appear in shortlists. The remaining gap is usually consistency and third-party weight.
41-50The nameDefend it monthly. Citation churn does not care that you won once.

Table 1: reading a 0-to-50 total. Most people running the audit for the first time land in the bottom two bands.

Field note

Score each category separately too. A total of 40 built as 10 plus 10 plus 8 plus 8 plus 4 and a total of 40 built as 4 plus 6 plus 10 plus 10 plus 10 are different patients. The category profile, not the total, picks your treatment.

How retrieval and training data produce different failure patterns

Two of the five categories tend to fail for genuinely different mechanical reasons, and knowing which one you are looking at changes the fix. A Category 1 failure, where the engine gets your basic facts wrong or thin, often traces back to what the underlying model learned during training versus what it can look up live through retrieval or browsing. Training Data vs Retrieval covers this distinction in depth, but the short version is that a model's training snapshot can be stale or sparse for anyone outside a narrow set of widely covered public figures, which means the live retrieval layer, when the product has one, carries far more of the weight for most individuals than people assume. A Category 5 failure, where your ideas never get cited, is a different mechanism entirely: it usually means your published work is not structured at the passage level in a way a retrieval system can lift cleanly, a problem covered in Chunk Theory and Passage Optimization. Treating a retrieval problem like a training problem, or the reverse, wastes exactly the months this audit exists to save you.

Reading the results: symptom, diagnosis, fix

What can you fix the same afternoon?

Most repairs take weeks, but a first audit usually surfaces two or three that do not. If prompt 1 returned a wrong employer or an old title, the fastest source of that error is usually a stale bio you actually control: an old speaker page, an abandoned profile, a bylined guest post with a five-year-old blurb. Correct or retire those the same day. If prompt 3 missed your most important work, check whether that work is actually attributable in the markup itself, your name in the byline, on the page, in the schema, not just somewhere in the prose. And if prompt 23 came back empty, write the canonical statement of your signature position on your own site before the month is out, because an engine cannot quote a view you have never published in a liftable form. Quick wins do not move the score much on their own, but they stop the bleeding while the slower layers compound underneath.

Building your own prompt list beyond these 25

Extending the audit
  1. Start from a real buyer conversation. Pull the actual phrasing a recent client used before they hired you, not the phrasing you would use to describe yourself.
  2. Add one prompt per service line. If you offer three distinct services, each deserves its own recommendation-category prompt, since engines can rate you well on one and invisible on another.
  3. Add a competitor substitution prompt. Swap in a rival's name and see whether the engine's phrasing about them differs structurally from its phrasing about you.
  4. Retire prompts that never move. If a prompt scores identically for three straight months, it is measuring something stable, not something you are actively working on. Spend the time elsewhere.

A short extension protocol for tailoring the fixed 25 prompts to your own niche after the first baseline round.

Common scoring mistakes that inflate your number

The single most common error is scoring your own name-recognition instead of a stranger's experience, reading generosity into a hedge because you know the hedge is technically about you. A close second is skipping the fresh-session rule after the first round, since a chat history that remembers you will keep flattering you round after round, producing a number that climbs for reasons that have nothing to do with the world. A third is conflating being mentioned with being recommended. An answer that lists you third among five names, with no explanation of why, is a 1, not a 2, even though technically you were mentioned. Hold the line on the scoring rubric even when a generous read would feel better, because the entire value of the exercise is that the number is honest enough to act on.

How often should you re-run it?

Monthly, same prompts, fresh sessions, logged next to the previous rounds. Citation sets and retrieval indexes shift on their own schedule regardless of what you publish, so a quarterly cadence misses most of the movement and an annual one is closer to archaeology than measurement. The re-run takes half the time once the spreadsheet already exists. If you want continuous signal between manual rounds, dedicated monitoring tools exist and we compared the real ones in AI Visibility Tools Compared, but the manual audit stays valuable even then, because you read nuance in a raw answer that a dashboard flattens into a single trend line. For the deeper measurement stack, metrics and all, see How to Track Your AI Visibility.

One hour, twenty-five prompts, one number. That is a small price for replacing "I think AI ignores me" with a dated, scored, category-level map of exactly what the engines believe about you today. And if your baseline comes back rough and you would rather not fix it alone, closing that gap is precisely what our services are built for.

FAQ

Which AI engine should I audit first? +
Whichever general consumer AI assistant has the broadest everyday adoption in your market, since that is where most buyers are typing these questions first. Add one or two others once that baseline exists, because answers differ meaningfully between engines built on different retrieval choices.
How often should I repeat the 25-prompt audit? +
Monthly, using the same prompts in fresh sessions, logged next to previous rounds. Citation sets and retrieval indexes shift on their own schedule, so a quarterly check misses most of the movement.
What counts as a good score? +
Out of 50, below 11 means you are effectively invisible, 11 to 25 means the engine knows you but does not prefer you, 26 to 40 makes you a genuine contender, and 41 or above means you are the name for your niche. Most people running it for the first time land in the bottom two bands.
Why does the same prompt sometimes score differently a few months apart with nothing published in between? +
Because retrieval indexes and the sources an engine draws on change on their own schedule, independent of anything you did. That is exactly why a dated baseline and a monthly re-run matter more than a single one-off audit.
Should I score a hedge like there are several experts including your name as a win? +
No. Being listed without explanation or preference is a 1, not a 2. Reserve a 2 for an answer that names you prominently, with correct facts, and, on a recommendation prompt, actually recommends you rather than merely acknowledging you exist.
What is the fastest fix if the audit comes back badly? +
Correct or retire any stale bio causing wrong facts, since that is usually a same-day fix, then move to the identity-consolidation routine before touching content or outreach. Content and third-party proof only compound once the identity layer underneath them is clean.

Find out what AI says about you today.

Start with a baseline. See the exact words the engines return about your name, then decide.

Claim your name →