How to Run an AI Visibility Audit by Hand
Aug 31, 2026 · 10 min read
A complete AI visibility audit you can run this week with a spreadsheet and no budget: the prompt set, the four fields to record per answer, the source list, and how to read the gaps. Written for an in-house marketer or a consultant producing a first assessment for a client.

An AI visibility audit is a fixed set of buyer questions run across the answer engines your category uses, with four things recorded per answer: whether you appeared, how you were described, which competitors appeared, and which sources were named. You can run one by hand in a week, and the baseline it produces belongs to you rather than to a vendor.
Every count in this article is a choice made here, not an industry standard. Thirty prompts, five engines, one week, one spreadsheet. Change any of them and the method survives, so long as you write down what you changed and date the run.
Key Takeaways
- The prompt set is the audit: thirty questions phrased the way your buyers phrase them beat three hundred keyword variants.
- Four analytical fields per answer: whether you appeared, how you were described, which competitors appeared, and which sources were named. Record the prompt, engine, and date separately as run metadata.
- The source list is the artifact that pays: the domains engines cite in your category are next quarter's publishing priorities, already ranked by frequency.
- A mention rate on its own tells you close to nothing. It stays trivia until you can name the sources behind the answers you missed.
- A hand-run audit expires. It is a dated snapshot, so re-run it before you quote it in a meeting.
What an AI Visibility Audit Produces
An AI visibility audit produces four artifacts and no report deck. Everything downstream is built out of these four, including the slide your leadership asks for.
- The prompt set. The questions you froze before running anything, numbered and dated.
- The answer log. The raw text of every answer, pasted rather than summarized.
- The source list. Every domain the engines named or linked across the whole set, counted.
- The gap list. The sources that never mention you, and the competitors that appear in nearly all of them.
The Labels Often Overlap
People ask what a GEO audit is as though it were always a separate exercise. In practice, GEO audits, AEO audits, and AI visibility audits often describe overlapping work. For the method used here, all three labels lead to the same four outputs. Use whichever label your stakeholders already say out loud.
The compressed version of this is a five-step section on one page. In the end-to-end discipline, auditing your current model visibility produces the baseline that the seven-step process after it works from. This article expands those five steps into a week of work. The four recorded fields stay exactly as that page defines them.
Building the Prompt Set: Category, Comparison, and Brand Questions
The prompt set is the part no tool does better than you. You know how your buyers phrase a question. A tool guesses. Thirty prompts, split three ways, is the shape used here: fifteen category questions, ten comparison questions, and five brand questions.
| Prompt type | Count used here | What it tests | Example, adult literacy nonprofit |
|---|---|---|---|
| Category | 15 | Whether you appear in the answer at all | "Which charities run adult literacy programs in Yorkshire?" |
| Comparison | 10 | Where you sit against named peers | "Compare UK adult literacy charities for a corporate volunteering partnership" |
| Brand | 5 | Whether the engine describes you correctly | "Is [organization name] a reputable adult literacy charity?" |
Write the Questions in Your Buyer's Words
A nonprofit is the cleanest place to see why this matters. Its buyers are donors, corporate partners, and referral agencies. None of them uses the sector vocabulary the organization uses about itself. Write the question the way an outsider would type it, with their geography, their budget, and their job title in it.
Weight the split toward the question your stakeholders will argue about. A team defending its place against three named rivals needs more comparison prompts. A team nobody has heard of needs more category prompts. Absence shows up in category questions first.
Freeze the Set Before You Run It
Save the prompts to a numbered tab and stop editing them. A prompt improved halfway through the run destroys the comparison. It destroys next month's comparison too, which is the more expensive loss. The same discipline sits behind building a prompt set worth tracking on one engine. It does not get easier at five.
Running the Prompts Across Engines by Hand
Most people who ask how to check AI visibility stall here. The running is dull rather than difficult. Thirty prompts across five engines is 150 answers. At roughly two minutes each to ask, copy, and paste, that is about five hours, spread over three or four sittings.
Run the set against ChatGPT, Claude, Perplexity, Gemini, and Grok. Add Google AI Mode as a sixth surface if your category gets real Google traffic. Read how AI Mode picks its sources before you read anything into that tab.
Rules That Keep the Run Comparable
- Fresh session for every prompt. Use a private window or log out. Personalization and chat memory will otherwise quietly flatter you.
- No follow-ups. Ask, copy, close. A second turn produces a better answer and a useless data point.
- Same week for every engine. Answers drift, so a run spread over a month is comparing engines to calendars.
- Paste the raw text. Summarizing as you go is the fastest way to lose the description you will need in the gap analysis.
- Screenshot anything with rendered citations. Perplexity and AI Mode show their sources on screen, and that screenshot is your evidence later.
- Date and stamp every row with the engine name and the day you ran it.
Recording Four Things Per Answer
Use one row per answer. Record the prompt number, engine, and date first, followed by the four fields below. Keep this structure identical between runs so the results remain comparable.
| Field | What you record | Format |
|---|---|---|
| 1. Appearance | Whether your brand is named anywhere in the answer | Yes or No |
| 2. Description | The sentence that describes you, pasted verbatim | Free text |
| 3. Competitors | Every rival named, in the order the answer names them | List |
| 4. Sources | Every domain named or linked in the answer | List |
Resist adding more analytical fields during the first audit. Sentiment scores, position numbers, and confidence ratings feel useful in week one. They go unfilled by week three, and a half-populated column is worse than an absent one. The four fields roll up into the metrics worth tracking as soon as you have two runs to compare.
Field 2 is the one teams skip and later regret. A brand can appear in twenty-two answers from one engine’s thirty-prompt run and still be called a regional supplier. Its buyers are shopping nationally. Absence is a publishing problem, and a wrong description is a correction problem. A mention count cannot tell you which one you have.
Building the Source List
Sort every entry from field 4 into one column. Count how often each domain appears across all 150 answers, then rank them. That ranked list is the audit’s most valuable output. It shows which domains the engines cited most often across your sample.
A Real Source List, Counted
Here is a counted one. On 31 August 2026 we queried Ahrefs Brand Radar for every answer in its United States index whose question carries the phrase "soil moisture sensor" or "soil moisture sensors", across ChatGPT, Gemini, Perplexity, Google AI Overviews, and Google AI Mode. That scope returns 26 answers to six distinct questions. These are Ahrefs figures, not Klarivo measurements. Brand Radar's own published methodology calls its metrics directional indicators rather than performance metrics.
| Cited domain | Answers citing it, of 26 | Distinct pages cited | What it is |
|---|---|---|---|
| youtube.com | 9 | 29 | Video platform |
| extension.umn.edu | 8 | 2 | University extension service |
| en.wikipedia.org | 6 | 3 | Reference encyclopedia |
| metergroup.com | 6 | 17 | Sensor manufacturer |
| mdpi.com | 5 | 5 | Academic publisher |
| dfrobot.com | 5 | 5 | Electronics supplier |
| forum.arduino.cc | 5 | 6 | Community forum |
| vegetronix.com | 4 | 3 | Sensor manufacturer |
| niubol.com | 4 | 4 | Sensor manufacturer |
| pmc.ncbi.nlm.nih.gov | 4 | 6 | Research repository |
Six of those ten domains sell no sensors. An agriculture technology company building a publishing plan from this list would investigate legitimate routes to inclusion on the extension service, forum, and video platform before defaulting to another post on its own blog. Claude is absent from that dataset, and Grok returned nothing for this scope, which is one argument for the hand-run version: you can ask them both yourself.
Read the Pages Column, Not Only the Count
Concentration changes the reading. The metergroup.com domain earns its six answers across seventeen different pages, so no single page is carrying it. The extension service earns eight from two pages, and one of those, soil moisture sensors for irrigation scheduling, accounts for six on its own. It is the most-cited single page in the sample, level with Wikipedia's entry.
Open that page and the gap stops being abstract. An extension irrigation specialist wrote it, it carries a 2019 review date, and its comparison tables name specific sensor products by model number. METER is one of those names, and metergroup.com is the most-cited vendor domain in the set. The other three manufacturers earn their citations without appearing on that page. A place in those tables is a route in, not the only one.
Where an engine renders a full link, record that URL alongside the domain. Reading a rendered citation is an observation. Inferring the page behind a domain-level report is a guess. Keep the two apart in your notes. Only the first survives a question in a meeting.
Reading the Gaps: Where Competitors Appear and You Do Not
Three gaps come out of the sheet. Each needs different work, from different people.
Source gaps. A domain cited often that carries no mention of you. Rank these by citation frequency and start at the top. In the sample above, that ordering puts a university extension service ahead of every sensor manufacturer in the category.
Description gaps. You appear, and the sentence is wrong or thin. Pull every field 2 entry into one list and read them together. Repeated wording across engines usually traces back to a single source page you can go and correct.
Engine gaps. You appear on two engines and vanish on the other three. That pattern is about where the engines read, and it says little about what you published. It is the clearest argument for running the set across all five rather than the one your team already uses.
Do the Comparison Out Loud
Use field 3 to compare brand appearances. For each competitor, count the answers that name them and compare that total with the answers naming you.
Use field 4 separately to compare citation frequency. In the soil moisture sample, YouTube appears in nine answers, the extension service in eight, and the strongest vendor domain in six. That describes the source landscape, not competitor share of voice.
Together, the two readings show who appears and which sources shape the category. Keep them separate so a frequently cited vendor domain is not mistaken for proof that the vendor was named just as often.
When a Manual Audit Stops Being Enough
A manual AI visibility audit is a snapshot, and it should be sold internally as one. Three limits matter before anyone treats it as a monitoring system.
Sample size. Thirty prompts across five engines is 150 readings taken in one week. That is enough to find the obvious gaps and not enough to detect a five-point move. Small changes between two hand-run audits are noise until proven otherwise.
Engine drift. Run the same prompt twice on the same day and the wording changes. Model updates move answers for reasons nobody outside the provider can see, which is why the date on the run matters more than the precision of the numbers.
Labor. Five hours a month, done properly, is sixty hours a year of copying and pasting. Teams do not stop because the method failed. They stop because month four lands in a busy week.
This is the point where people start searching for an AI visibility checker. The category has two ends. A free AI visibility checker returns a one-off reading across a handful of prompts. A scheduled AI search visibility checker reruns your set across several engines and charts the trend. What separates them is whether the output names the sources behind an answer or stops at a mention count. Only the first tells you what to do next. Sequencing the work that follows is a separate job, and a dated ninety-day plan covers it in more detail than belongs here.
How Klarivo Runs the Audit at Scale
Klarivo runs this same loop as a service. The difference is repetition rather than insight. Klarivo Monitor holds the prompt set, runs it across ChatGPT, Claude, Perplexity, Gemini, and Grok independently, and refreshes on a schedule the client controls. The reporting surface covers query-level mention rates, competitor share of voice, sentiment, and a top citation sources list that reports domains.
The reason to hand it over is the sixtieth hour, not the first. A team that runs this audit by hand once has proved it understands its own category. That is the hard part. What breaks is running the identical set the identical way twelve months running, then acting on the source gaps it keeps surfacing.
That second half is where the Reddit and community work sits, because a source gap closes when somebody publishes into that source rather than when a dashboard reports it again. You can see the program sequence in how Klarivo works, and the tracking side in Klarivo Monitor. Book a Klarivo discovery call. Fifteen minutes, a slot you choose, confirmed instantly.
Frequently Asked Questions About AI Visibility Audits
Is a free ai visibility checker worth running before I do this by hand?
A free AI visibility checker can reveal whether a problem exists by sampling a few prompts on one or two engines. That is useful as an initial check, but too limited to serve as a reliable baseline. The prompt set is what makes the audit specific to your market, and a free tool is unlikely to produce that set without substantial input from your team.
How often should I re-run the audit?
Monthly if you are actively publishing, quarterly if you are not. For a manual audit, running more frequently than monthly will often add more noise than useful evidence unless your brand has just changed something material. The value comes from the same prompts on the same engines, so an irregular schedule with a frozen set beats a strict schedule with an edited one.
How many prompts do I need for the audit to mean anything?
Thirty is the count used here. Twenty is the working floor, matching the range the end-to-end guide publishes. Below twenty, one unusual answer swings the picture. Past forty, our experience is that you mostly add variations of questions already in the set.
Who inside a marketing team should own this?
Whoever already owns a reporting line, which is usually SEO. The running is administrative and can be delegated. The reading is not, because deciding whether a description is wrong or merely unflattering takes someone who knows the category and the sales objections.
The audit says we appear often but the description is wrong. What now?
Trace the wording back through your source list to the pages using it. An outdated positioning line on a directory entry or a review profile travels further than most teams expect. Correct the source where possible, then reinforce the accurate information across other credible pages. Changing the evidence available to the engines is what can change later answers.
Start accelerating your Reddit presence
See how Klarivo can shift your visibility across AI, search and buyer communities.
Book a Demo

