Audit, Pilot, or Retainer: Which AI Visibility Offer Fits?
Learn how to match AI visibility audits, pilots, and retainers to client needs. Avoid mismatched offers and improve results with the right engagement type.
A practitioner's guide to matching AI visibility audits, pilots, and retainers to what the client actually needs before you quote a price.
Updated on: 2026-09-09
If a client asks "where do we stand in AI answers," they want an audit. If they ask "can this even work for us," they want a pilot. If they ask "how do we keep improving as this changes," they want a retainer. Most of the mismatched deals I've watched fall apart came from selling one of these when the client was really asking for another. An audit produces a diagnosis. A pilot tests an intervention. A retainer manages a system that keeps moving. Get that framing wrong and either the client feels oversold or you leave money and results on the table.
The three offers are levels of commitment, not interchangeable packages. Here's how I decide which one fits.
What each offer actually does
The cleanest way to keep these straight is to write down the primary job of each one and refuse to blur it.
| Offer | Primary job | Time horizon | The question it answers |
|---|---|---|---|
| Audit | Diagnose the current problem | One-time or short fixed engagement | "Where do we stand, and what's wrong?" |
| Pilot | Validate that an intervention moves the needle | 4 to 12 weeks | "Can this work for our category?" |
| Retainer | Improve and defend visibility continuously | Ongoing monthly or quarterly | "How do we keep winning as things change?" |
An audit that promises performance is a broken audit. A pilot that's just a discounted retainer wastes everyone's time. A retainer that reports activity instead of outcomes is the most common failure of the three, and the easiest to fall into.
When an audit is the right call
An AI visibility audit is a bounded diagnostic study of how selected AI engines respond to a defined set of buyer prompts about a brand, its category, its competitors, and its alternatives. The unit of measurement is the model's answer, not a search-result position. That distinction matters more than it sounds. A brand can rank fine in Google and be completely absent from what ChatGPT recommends.
Sell an audit when the client has never measured AI visibility, is skeptical that AI answers move their buyers, or needs evidence before approving any recurring budget. It's also the right move when the target audience, prompt set, or competitor list is fuzzy. You can't run a useful ongoing program without those defined, and the audit is where you define them.
A strong audit deliverable is a prioritized diagnosis, not a giant dashboard the client has to interpret. It should answer, in plain language:
- Which commercially important prompts show the biggest gap?
- Is the brand absent, mentioned, cited, or actually recommended?
- Which competitors keep showing up?
- Which third-party sources do the models cite instead?
- Is the brand described accurately?
- What can the client fix on their own site, and what needs outside authority?
- Which three actions carry the highest expected business value?
On prompt count, I've settled on a practical range of 25 to 100. Fewer than 25 and you can't see the distribution clearly. More than 100 and you're adding noise for most SMB brands without adding insight. For a first pass, 20 to 40 high-value prompts tested against two or three close competitors is usually enough to produce a defensible picture.
One thing I'd push back on: the instinct to test every AI engine because they're available. ChatGPT, Claude, and Perplexity pull from different source ecosystems and disagree with each other more than people expect. Test the engines your client's buyers actually use. Universal coverage for its own sake just spreads your attention thin.
The audit is also, frankly, the best pitch tool you have. One growth lead I know ran an audit on a Monday, sent it with a proposal Tuesday, and closed a retainer that week. The diagnosis sold the ongoing work because it made the gap concrete.
When a pilot earns its keep
A pilot sits between the audit and the retainer. It's a limited, hypothesis-driven test: use the audit baseline, change a small number of high-priority assets or authority signals, then re-measure the same prompts to see if anything moved.
A pilot is not a cheaper retainer. If you treat it that way you'll disappoint people. A real pilot has a narrow business problem, a fixed prompt and competitor set, two to five documented interventions, a set measurement window, a pre/post comparison, and a decision rule written down before you start. That last part is what separates a pilot from a hopeful before-and-after anecdote.
Choose a pilot when the client believes there's a problem but wants proof before committing to ongoing spend, can make changes during the test period, and has two or three high-value prompts where improvement would genuinely matter to the business. It fits especially well for:
- A B2B SaaS product that's missing from "best tool" and comparison prompts.
- A professional-services firm that gets mentioned but never recommended.
- An ecommerce brand whose products show up through retailers but not with the right attributes.
- Any business that AI answers describe with stale pricing, wrong integrations, or the wrong ideal customer.
Here's the discipline that makes or breaks a pilot: don't change everything at once. If you rewrite ten pages, launch a PR push, open up crawler access, and adjust pricing in the same two weeks, visibility might improve and you'll have no idea which change did it. Keep the interventions few and documented so the result means something.
A concrete example. A B2B project-management tool shows up in 4 of 20 baseline prompts, recommended in only 2. Three competitors dominate the comparison answers. Independent directories keep getting cited, while the client's own site explains features but never addresses the agency use case that buyers keep prompting about. The pilot: test the same 20 prompts across four engines, refresh two high-intent pages, publish one agency-specific comparison page, correct the product language, chase a shortlist of relevant directories and review sites, then re-run the prompts on a schedule.
Success isn't "the score went up." Success is that priority prompts show repeatable improvement, the brand gets recommended more often, the cited source ecosystem changes, or the pilot produces a clear explanation for why the original hypothesis was wrong and should be revised. A pilot that disproves your theory cleanly is still a good pilot.
When a retainer is justified
A retainer is an ongoing engagement where you repeatedly monitor a defined prompt portfolio, diagnose what changed, execute approved improvements, report what moved, and turn unresolved gaps into next period's work. The recurring loop looks like this: measure, prioritize, influence, retest, report, renew.
Recommend a retainer when the client has already done an audit or pilot, has recurring commercial value tied to being recommended, can supply ongoing content or subject-matter input, and operates in a category where competitors and cited sources shift often enough that a static diagnosis goes stale. If nobody internally can publish, revise, or approve anything on a regular basis, a retainer will stall no matter how good your process is.
A monthly cycle I'd stand behind:
- Run the stable prompt set.
- Capture the full answers and cited sources.
- Separate true negatives from provider results that were simply unavailable.
- Compare against baseline and the prior period.
- Identify priority losses and competitor gains.
- Pick content, technical, or authority actions.
- Get human approval on claims and drafts.
- Publish or execute.
- Re-test on schedule.
- Report activity and outcomes separately.
- Convert unresolved gaps into next month's scope.
That "separately" in step 10 is the whole game. Reporting that you published four pages is not the same as showing the model now recommends the client. Confusing the two is the fastest way to lose a renewal once the client gets sophisticated.
This is where the tooling matters, because doing all of this by hand across four engines every month doesn't scale past a client or two. This is the exact operating model SEOforGPT was built around. Its Monitoring Agent runs scheduled visibility checks across ChatGPT, Claude, Perplexity, and Gemini and flags movement. The Content Agent turns visibility gaps and buyer prompts into gap-driven articles and publishes to WordPress, Webflow, Ghost, Notion, or Wix. The Reddit Agent surfaces relevant buyer conversations and drafts replies for review, and the Outreach Agent finds the roundups, listicles, directories, and comparison pages where the brand should appear and prepares pitches. White-label reporting handles the part most agencies dread: proving the work to a client or a board.
The renewal argument matters too. Don't tell a client you "permanently won AI search." That claim is fragile because prompts, source pages, competitors, and model versions all change. The stronger argument is that priority prompts were measured consistently, some commercially important observations improved, the remaining gaps are understood, and next quarter has a defined backlog tied to real evidence.
The measurement caution that keeps you honest
AI answers are variable. A single response is a snapshot, not a trend. Independent research across Perplexity, SearchGPT, and Gemini found that identical queries produce different citation sets across repeated runs, and that raw citation counts differ wildly by platform. Gemini returned a median of roughly 36 to 40 citations per response in that work, Perplexity around 19 to 22, and SearchGPT around 5 to 7. You cannot compare raw citation counts across engines and call it analysis.
Sample size sets confidence. That same research found Gemini often hit a stable citation-share estimate at roughly 30 to 50 observations, while Perplexity needed closer to 90 to 100 and SearchGPT sometimes 150 or more. These aren't universal minimums, but the lesson holds: a 10-prompt pilot and a 100-prompt retainer report don't carry the same weight. Use smaller samples for directional diagnosis, repeated samples for pilot decisions, and larger or repeated samples before you claim a small percentage-point win.
Two more distinctions I'd never skip in a report. First, a URL being listed as a source isn't the same as that source shaping the answer. Report citation selection and citation influence separately, or clients will celebrate more citations even when their description and recommendation status didn't budge. Second, crawl access is a prerequisite, not proof. OpenAI's OAI-SearchBot controls whether a site can appear in ChatGPT Search, and it operates independently from GPTBot. Allowing the crawler removes one barrier. It doesn't earn a recommendation.
And Search Console won't cover this for you. Google folds AI Overviews and AI Mode into overall Search Console traffic under the Web search type, and while dedicated generative-AI reporting arrived in 2026, it still measures your exposure inside Google's features. It doesn't tell you what an answer said, which competitors appeared, or whether ChatGPT described your client accurately. For that you need answer-level monitoring across engines.
What I'd do first
If a prospect can't tell you which prompts matter to their buyers, start with an audit. Don't skip to a pilot or retainer on a hunch, because you'll be optimizing prompts nobody's actually asking.
If they've got a clear problem and can make changes but aren't ready to commit budget, run a pilot with a written decision rule. Two to five interventions, one measurement window, and a plain rule for continue, revise, or stop.
If they've already seen a diagnosis, have real commercial exposure to AI recommendations, and can feed the work each month, move to a retainer and separate activity from outcomes in every report.
The through-line across all three: measure the answer, report movement rather than a single tidy score, and never promise a guarantee on a system that changes underneath you.
FAQ
Is a pilot just a smaller retainer?
No, and treating it like one is where people go wrong. A retainer is an open-ended process for improving and defending visibility. A pilot is a closed test with a hypothesis, a fixed window, and a decision rule about whether to keep going. If your "pilot" has no pre-written rule for stopping, you've sold a discounted retainer and called it a test.
How many prompts should an audit track?
For most SMB brands, 25 to 100. Below 25 you can't see the distribution clearly, and above 100 you add noise without insight for many clients. A first pass of 20 to 40 high-value prompts against two or three close competitors is usually enough for a defensible diagnosis.
Can I promise a client they'll get recommended by ChatGPT?
You shouldn't. Prompts, source pages, competitors, and model versions all change, and answers vary even for identical queries. What you can promise is consistent measurement, specific interventions tied to evidence, and honest reporting of what moved. Guarantees on probabilistic systems don't survive contact with reality.
Does Google Search Console tell me how I'm doing in AI answers?
Only partially. Google reports AI Overviews and AI Mode traffic inside overall Search Console data, so it shows exposure within Google's own features. It won't tell you what an answer said, which competitors were named, or how you were described across ChatGPT, Claude, or Perplexity. For that you need answer-level monitoring across the engines your buyers use.
Further reading
Users also found this interesting
Continue with guides covering the same topic and workflow.
SEO Audit vs Monthly Retainer: Which One Works?
Compare SEO audits and monthly retainers to find out which fits your needs in 2026. Learn when to choose each and how AI visibility changes the game.
How to Choose an AI Visibility Platform for Clients
Learn how agencies can select the right AI visibility platform for client work, comparing features, pricing, and workflow essentials for effective delivery.
SEOforGPT vs AEO Vision for Agencies
Compare SEOforGPT and AEO Vision for agencies: features, pricing, and best fit for content production, monitoring, and AI visibility in 2026.
Ready to optimize your content for AI?
Start creating AI-native content that gets discovered and recommended by leading AI systems.