Before you optimize anything, measure what the engines already believe. This audit is 25 prompts in five categories, identity, recommendation, comparison, trust, and citation, run in fresh sessions and scored 0 to 2 each. One hour produces a 0-to-50 baseline, a map of exactly where you fail, and a fix priority. The prompts are below, ready to copy, along with the reasoning that makes the scores trustworthy instead of flattering.
Most people have never asked an AI system the questions their buyers actually ask it. They optimize blind, publish blind, and then wonder why nothing moves. An audit costs an hour and replaces the guessing with a number you can act on.
Why audit before you optimize?
Two reasons carry the whole case for auditing first. First, diagnosis beats guesswork: the specific failure pattern tells you which layer is broken, and the layers have different repairs entirely. Wrong facts demand cleanup of the sources feeding the error. Absence demands publishing something citable. Presence without preference demands third-party proof, not more content. Spending months writing when your real problem is a contradictory identity, a case covered in Structure Your Identity for Machines, is the standard expensive mistake, and it stays invisible until you actually look at what the engine says.
Second, you need a dated before picture. What a generative engine surfaces about any given name shifts as new pages get published, old ones get deprecated, and the underlying retrieval index refreshes, which is a mechanical property of how these systems work rather than a fixed fact you can check once. See Training Data vs Retrieval for why an answer that was accurate last quarter can quietly go stale without anyone editing a single page. Without a dated baseline, you cannot tell whether a change in what the engine says was caused by your work or by the churn that happens regardless of what you do.
The rules that make results trustworthy
- Fresh sessions only. Open a new chat for every category, and disable memory or use a logged-out window where the product allows it. An assistant that remembers previous conversations with you will flatter you in ways a genuine stranger's query never would.
- Start with the most widely used general assistant. Whichever consumer AI chat product has the broadest everyday adoption in your market is the primary exam, because that is where most buyers are actually typing these questions. Repeat on at least one or two others if you have time, since answers differ meaningfully between engines built on different retrieval and ranking choices.
- Record verbatim. Paste each answer into a spreadsheet with the date, the engine, and the prompt number. Summaries hide the details that matter, like a competitor named ahead of you or a title that is subtly wrong.
- Customize the brackets honestly. Where a prompt says your field or your city, use the words a real buyer would type, not your internal jargon. If you have not decided which buyer questions matter most, settle that first with Money Queries: The AI Questions Worth Winning.
- Do not argue with the engine. No follow-ups, no corrections mid-audit. You are measuring the cold answer a stranger gets on the first try, not the improved answer you can coax out with three rounds of clarification.
The 25 prompts
Category 1: Identity (prompts 1-5)
Does the engine know who you are, and is what it knows true?
- Who is [your full name]?
- What is [your full name] known for?
- What has [your full name] written, built, or published?
- Give me a short professional bio of [your full name].
- What is [your full name]'s professional background and career history?
Category 2: Recommendation (prompts 6-10)
Does the engine name you when nobody asked about you specifically? This is the money category.
- Who are the best [your specialty] consultants in [your city or country]?
- Recommend a [your specialty] expert for a [your target client type] that needs [the outcome you deliver].
- I need help with [the problem you solve]. Who should I talk to?
- List five respected independent [your specialty] specialists working today.
- Who is an up-and-coming expert in [your niche]?
Category 3: Comparison and shortlist (prompts 11-15)
When buyers narrow the field, do you survive the cut?
- Compare [your name] and [a direct competitor] for [the service you both offer].
- Give me a shortlist of three people for [a project in your niche] and explain each choice.
- What are the alternatives to hiring [a large firm or famous name in your space] for [the service]?
- Should I hire a big agency or an independent specialist for [your service]? Name specific options.
- Who does work similar to [a well-known peer you respect]?
Category 4: Trust and verification (prompts 16-20)
Buyers increasingly run a first pass of due diligence through an assistant before they ever pick up the phone, and verification questions make up a large share of that research.
- Is [your name] legitimate? What is their track record?
- What credentials and experience does [your name] have in [your field]?
- Are there any criticisms or red flags about [your name]?
- Has [your name] worked with companies like [your target client type]?
- Why would someone hire [your name] over other options?
Category 5: Source and citation (prompts 21-25)
Does your published thinking feed the answers in your niche, or does the engine cite everyone but you?
- What are the best resources for learning about [your core topic]?
- Summarize the main approaches to [the problem you solve] and cite your sources.
- What does [your name] say about [your signature topic]?
- Quote or paraphrase something [your name] has published about [your topic].
- Whose frameworks are most cited for [your niche], and where do those frameworks come from?
Customizing the brackets: two worked examples
The brackets are where audits go soft, so here is what honest customization looks like. Take prompt 7 for a hypothetical fractional CFO named Jane Okafor who serves venture-backed startups: "Recommend a fractional CFO for a Series A SaaS company that needs investor-ready reporting." Every bracket is filled with buyer language, no jargon, no flattery. The same prompt for a hypothetical employment lawyer named Tomas Reyes reads: "Recommend an employment law specialist for a UK retailer facing a tribunal claim." Notice neither version mentions the person being audited. The recommendation category only tells you something real when you ask exactly as a stranger would, with no hint that you already know or favor a particular name.
Now prompt 23 for the same two people. For Jane: "What does Jane Okafor say about burn-rate discipline in a downturn?" For Tomas: "What does Tomas Reyes say about settlement versus tribunal strategy?" If you cannot fill the signature-topic bracket without hesitating, that hesitation is itself a finding. It means you have not claimed a topic tightly enough for an engine to associate you with one, and no amount of visibility work elsewhere fixes an unclaimed position. The fix for that specific gap lives in Answer-Shaped Writing, since a signature position only becomes citable once it is written in a form an engine can lift cleanly.
Logging: the spreadsheet that makes it an instrument
One row per prompt per engine, six columns: date, engine, prompt number, verbatim answer, score, and names mentioned before yours. That last column is the quiet gold. Over three or four monthly rounds it becomes a working census of who the engines actually consider your competition, which rarely matches who you consider your competition. Total each category, keep a running chart of the five subtotals, and resist the urge to prettify anything. This document is an instrument, and instruments are allowed to be ugly.
Scoring: turn answers into a number
Score every prompt 0, 1, or 2. Be harsh, a generous audit is a useless audit.
- 0: not mentioned, or mentioned with materially wrong facts. A wrong answer scores zero even if it is flattering.
- 1: mentioned accurately but weakly: late in the list, thin detail, hedged, or missing your actual positioning.
- 2: named accurately, prominently, with correct facts and, on recommendation prompts, actually recommended rather than merely acknowledged.
Maximum 50. Read your band against the table below.
| Score range | Reading | What it means |
|---|---|---|
| 0-10 | Invisible | The engine cannot see you or cannot resolve you as one entity. Start with identity, not content. |
| 11-25 | Legible, not preferred | You exist as an entity but nothing pushes the engine to choose you over the field. |
| 26-40 | Contender | You appear in shortlists. The remaining gap is usually consistency and third-party weight. |
| 41-50 | The name | Defend it monthly. Citation churn does not care that you won once. |
Table 1: reading a 0-to-50 total. Most people running the audit for the first time land in the bottom two bands.
Score each category separately too. A total of 40 built as 10 plus 10 plus 8 plus 8 plus 4 and a total of 40 built as 4 plus 6 plus 10 plus 10 plus 10 are different patients. The category profile, not the total, picks your treatment.
How retrieval and training data produce different failure patterns
Two of the five categories tend to fail for genuinely different mechanical reasons, and knowing which one you are looking at changes the fix. A Category 1 failure, where the engine gets your basic facts wrong or thin, often traces back to what the underlying model learned during training versus what it can look up live through retrieval or browsing. Training Data vs Retrieval covers this distinction in depth, but the short version is that a model's training snapshot can be stale or sparse for anyone outside a narrow set of widely covered public figures, which means the live retrieval layer, when the product has one, carries far more of the weight for most individuals than people assume. A Category 5 failure, where your ideas never get cited, is a different mechanism entirely: it usually means your published work is not structured at the passage level in a way a retrieval system can lift cleanly, a problem covered in Chunk Theory and Passage Optimization. Treating a retrieval problem like a training problem, or the reverse, wastes exactly the months this audit exists to save you.
Reading the results: symptom, diagnosis, fix
- Low Category 1 (identity): the engine holds no coherent entity for you. Fix the identity layer first: consistent naming, connected profiles, Person schema. The routine is in Structure Your Identity for Machines.
- Wrong facts anywhere: a contradiction problem. Hunt the stale bios and conflicting pages feeding the error before you publish anything new, using the approach in Contradiction Debt.
- Good identity, weak Categories 2 and 3: the engine knows you but has no reason to prefer you. That is a knowledge and network gap: publish citable depth and earn third-party references. The end-to-end sequence is in How to Get Recommended by ChatGPT.
- Weak Category 4 (trust): verification prompts return shrugs. Get proof on pages you do not own: client outcomes, press, speaking, reviews, the kind of third-party evidence covered in E-E-A-T as Data.
- Weak Category 5 (citation): your ideas are not machine-liftable. Restructure your published work so an engine can quote it cleanly and attribute it to you.
What can you fix the same afternoon?
Most repairs take weeks, but a first audit usually surfaces two or three that do not. If prompt 1 returned a wrong employer or an old title, the fastest source of that error is usually a stale bio you actually control: an old speaker page, an abandoned profile, a bylined guest post with a five-year-old blurb. Correct or retire those the same day. If prompt 3 missed your most important work, check whether that work is actually attributable in the markup itself, your name in the byline, on the page, in the schema, not just somewhere in the prose. And if prompt 23 came back empty, write the canonical statement of your signature position on your own site before the month is out, because an engine cannot quote a view you have never published in a liftable form. Quick wins do not move the score much on their own, but they stop the bleeding while the slower layers compound underneath.
Building your own prompt list beyond these 25
- Start from a real buyer conversation. Pull the actual phrasing a recent client used before they hired you, not the phrasing you would use to describe yourself.
- Add one prompt per service line. If you offer three distinct services, each deserves its own recommendation-category prompt, since engines can rate you well on one and invisible on another.
- Add a competitor substitution prompt. Swap in a rival's name and see whether the engine's phrasing about them differs structurally from its phrasing about you.
- Retire prompts that never move. If a prompt scores identically for three straight months, it is measuring something stable, not something you are actively working on. Spend the time elsewhere.
A short extension protocol for tailoring the fixed 25 prompts to your own niche after the first baseline round.
Common scoring mistakes that inflate your number
The single most common error is scoring your own name-recognition instead of a stranger's experience, reading generosity into a hedge because you know the hedge is technically about you. A close second is skipping the fresh-session rule after the first round, since a chat history that remembers you will keep flattering you round after round, producing a number that climbs for reasons that have nothing to do with the world. A third is conflating being mentioned with being recommended. An answer that lists you third among five names, with no explanation of why, is a 1, not a 2, even though technically you were mentioned. Hold the line on the scoring rubric even when a generous read would feel better, because the entire value of the exercise is that the number is honest enough to act on.
How often should you re-run it?
Monthly, same prompts, fresh sessions, logged next to the previous rounds. Citation sets and retrieval indexes shift on their own schedule regardless of what you publish, so a quarterly cadence misses most of the movement and an annual one is closer to archaeology than measurement. The re-run takes half the time once the spreadsheet already exists. If you want continuous signal between manual rounds, dedicated monitoring tools exist and we compared the real ones in AI Visibility Tools Compared, but the manual audit stays valuable even then, because you read nuance in a raw answer that a dashboard flattens into a single trend line. For the deeper measurement stack, metrics and all, see How to Track Your AI Visibility.
One hour, twenty-five prompts, one number. That is a small price for replacing "I think AI ignores me" with a dated, scored, category-level map of exactly what the engines believe about you today. And if your baseline comes back rough and you would rather not fix it alone, closing that gap is precisely what our services are built for.
FAQ
Which AI engine should I audit first? +
How often should I repeat the 25-prompt audit? +
What counts as a good score? +
Why does the same prompt sometimes score differently a few months apart with nothing published in between? +
Should I score a hedge like there are several experts including your name as a win? +
What is the fastest fix if the audit comes back badly? +
Find out what AI says about you today.
Start with a baseline. See the exact words the engines return about your name, then decide.
Claim your name →