From Prospect Audit to a Scoped GEO Proposal
Learn how to turn AI visibility audit findings into a scoped GEO proposal, mapping every recommendation to real evidence for defensible client results.
A practical method for turning AI visibility findings into a proposal every recommendation traces back to real evidence.
Updated on: 2026-09-06
The fastest way to lose a GEO deal is to hand a prospect a single number. "You have a 22% AI visibility score." Then silence, because the next question is always "compared to what?" and "so what do you actually want us to pay for?"
A prospect audit earns a proposal when every line item points back to something you observed. Buyer question, AI answer, who showed up, what got cited, the likely gap, the specific work, and how you will check whether the work changed anything. If you cannot draw that line, you are guessing, and the client will smell it.
I have watched clean audits collapse under one bad habit: treating the audit as proof that a tactic will work. It is not proof. It is a diagnosis. The audit tells you where the brand is missing and who is winning instead. It does not promise that three comparison pages will make ChatGPT recommend you. Sell the diagnosis honestly and the proposal writes itself.
What a prospect audit is actually measuring
Generative Engine Optimization, or GEO, is the practice of improving how likely a brand, page, or source is to appear, be cited, or influence an AI-generated answer. The original GEO research frames generative engines as systems that retrieve documents and use a language model to synthesize an attributed answer. That framing matters because it separates the things people casually lump together.
AI visibility is the observable presence of a brand in those answers. Inside that, a few outcomes look similar but mean different things:
- Mention: the answer names the brand. It does not prove your client's site supplied the information.
- Recommendation: the answer advises choosing the brand. Different from a passing mention in a long list.
- Citation: a page or domain is linked as supporting evidence. A brand can be recommended without its site being cited, and cited without being recommended.
- Provider coverage: the brand appears across ChatGPT, Claude, Perplexity, and Gemini. A blended average can hide that it is at 50% on one and 5% on another.
Keep mention and citation separate in your head and in your report. This is where most audits get sloppy, and it is exactly the distinction that makes a proposal defensible later.
The five questions a prospect audit has to answer
Skip the vanity metrics. A useful prospect audit answers five commercial questions:
- Can the brand be discovered at all?
- Where does it appear, and where does it vanish?
- Which competitors and which sources win instead?
- What kind of work could plausibly close each gap?
- What do you monitor to know if the work moved the number?
That last one is the part agencies skip, and it is the part that protects you six weeks in when the client asks "is it working."
Build the prompt set like you will have to defend it
If you test only branded queries, you learn that AI knows the brand exists. Useful for entity checks, useless for category discovery. Mix prompts across the buying stages instead.
| Buyer stage | Example prompt pattern | What you are watching |
|---|---|---|
| Problem discovery | "Best ways to solve [problem]?" | Is the category and brand associated at all |
| Evaluation | "What to look for in a [product]?" | Does the brand appear against decision criteria |
| Comparison | "[Competitor] alternatives for [use case]" | Included, excluded, or replaced |
| Recommendation | "Best [product] for [audience]" | Position and the competitor set around it |
| Transactional | "Where to buy/book [solution] in [location]?" | Commercial and local visibility |
| Branded validation | "What is [brand] known for?" | Entity understanding only, scored separately |
Branded prompts should not count toward the core discovery score. The point of the discovery number is unprompted category presence, not the fact that a system recognizes a name you fed it. This is how SEOforGPT's audit workflow treats it: brand-name prompts are excluded from the core visibility score, brand aliases get normalized so you do not double-count, and provider-level results stay distinct instead of getting flattened into one blended figure.
Record the exact prompt text, the category, geography and language, the provider and model where available, whether web browsing was on, the date, the competitors detected, and the URLs cited. If you change the prompt set after seeing bad results, that is a new baseline, not the same one. Rewriting prompts quietly until the score looks better is the audit equivalent of moving the goalposts, and it always comes back to bite you.
Metrics that belong on the report, and how not to lie with them
Every percentage needs its numerator, denominator, prompt categories, providers, and date. A number floating on its own is weak evidence.
Visibility rate is appearances divided by eligible discovery prompts. Eighteen appearances in sixty non-branded prompts is 30%. That is 30% of your test set, not 30% of every real-world question a buyer will ever ask. Say that out loud to the client so they never misread it later.
Average observed rank only applies when the answer contains an ordered list. Report how many prompts that average is built on. A brand appearing once at position one is not the same as a brand appearing twenty times at position three. And a missing appearance should lower the visibility rate, not get assigned a fake rank like position 10.
Citation rate is first-party domain citations divided by eligible prompts. Split first-party from third-party citations. A brand can be named across dozens of answers while its own site is barely cited, which points at a very different problem than low mention counts.
Recommendation rate requires you to define "recommended" before you score anything. A neutral name-drop in a list is not an endorsement. Label each outcome: explicit recommendation, included as an option, neutral, negative, or absent.
Share of voice is brand mentions divided by all tracked brand and competitor mentions in the same sample. It is a sample-level competitive figure, not market share, and it swings hard based on how complete your competitor list is and whether negative mentions count the same as positive ones. SEOforGPT's own definition is exactly this, with the caveat that it measures observable signals and does not guarantee future visibility. Put that caveat in your proposal too.
One more thing worth being honest about with the client: these numbers move week to week even when nobody touches anything. AI search citations are volatile. A single week's wobble from a tiny prompt set is noise, not a trend. Do not sell it as one.
Why the same prompt gives different answers
Google says its AI Overviews and AI Mode use "query fan-out," running multiple related searches across subtopics and discovering more pages while the response generates, and that the two features may use different models entirely, so their answers and links differ. Their guidance on AI features is deliberately conservative: pages need to be indexed, eligible for a normal snippet, and technically compliant, with no special requirements for AI features beyond ordinary fundamentals.
Practically, that means:
- Optimizing one page for one exact keyword is not enough when the system expands the query.
- A brand can be strong at the education stage and invisible at comparison.
- Third-party pages can shape the answer even when your client's site is technically perfect.
- The same prompt shifts by date, provider, model version, and location.
This is also why "we fixed your technical SEO, so now AI will recommend you" is a claim you cannot back. Technical readiness is an enabling condition. It clears the door. It does not walk the brand through it.
Mapping findings to scope
Here is the part that turns a diagnostic into a contract. Each finding maps to evidence, a diagnosis, a work package, and a measurement.
| Audit finding | Likely diagnosis | Work package | Success evidence |
|---|---|---|---|
| Absent from high-intent non-branded prompts | Weak category association or thin external evidence | Buyer-prompt research, content gap plan, priority pages, monitoring | More appearances in the same qualified prompt set |
| Appears but consistently below competitors | Recognized but weaker on relevance or proof | Comparison and evaluation content, stronger proof points | Better observed position or recommendation frequency |
| Mentioned but own site rarely cited | AI leans on third-party descriptions | Evidence pages, internal linking, clearer entity signals | Higher first-party citation rate and cited-page breadth |
| Competitors dominate directory and review citations | Third-party source gap | Source mapping, profile corrections, editorial outreach | More qualified third-party coverage |
| Important pages fail readiness checks | Crawl, index, render, or canonical problem | Technical remediation and validation | Blockers resolved, recrawl confirmed |
| Visibility differs sharply by provider | Provider-specific retrieval gap | Provider-specific monitoring and source analysis | Improvement by provider, not just blended average |
| Results swing between runs | Non-determinism or underspecified prompts | Expand sample, repeat, report ranges | Stabler baseline with documented variance |
If competitors win from directories, reviews, and forums, the answer is not "more blog posts." It is off-site source work. Selling monthly articles against a third-party citation problem is how retainers quietly fail.
A worked example
Say a B2B software prospect gets tested on 80 non-branded prompts across four providers:
- Appears in 16 of 80 prompts: 20% visibility
- Recommended in 7 of 40 recommendation prompts: 17.5% recommendation rate
- Competitor A in 42 prompts, Competitor B in 35
- Own domain cited in 8 prompts; third-party review and directory pages cited in 31
- Strong at educational prompts, nearly absent at comparison prompts
- Much stronger on Perplexity than on Google AI features
- Several priority product pages missing from the sitemap and weakly linked
The scoped proposal follows the evidence, not a package menu:
Phase 1, measurement and technical correction. Freeze the 80-prompt baseline, build provider-specific reporting, fix the sitemap and internal linking, validate indexability on priority pages, re-run the same set after recrawl.
Phase 2, commercial content. Build comparison pages for the clusters where competitors repeatedly win. Add decision criteria, limitations, implementation detail, and verifiable evidence. One use-case page for the biggest audience gap. Link it all from product pages.
Phase 3, external source work. Review the 31 third-party cited sources, find which ones materially influence answers, correct inaccurate profiles, and skip mass directory submissions and manufactured reviews.
Phase 4, validation. Repeat the same 80 prompts, add a small expansion set only once the baseline is stable, and report changes in visibility, recommendation rate, and first-party citations by provider.
Notice what this does not promise: a guaranteed percentage lift. The GEO research reported improvements up to roughly 40% under its own experimental conditions, but the authors were clear that results varied by domain and method, that the engines are black boxes, and that their metric was citation impression, not leads or revenue. Use that paper to argue optimization effects are measurable, never to promise a client a fixed number.
Where the tooling fits
Running 80 prompts across four providers by hand, normalizing aliases, deduping, archiving answer excerpts, then repeating it monthly is not sustainable at agency scale. This is the honest case for a platform. SEOforGPT runs the free prospect audit, tracks visibility across ChatGPT, Claude, Perplexity, and Gemini, separates citations from mentions, and produces white-label reports you can attach to a proposal. The free Agency Prospecting tier gives 10 pitch workspaces and a full visibility audit per prospect with no card, which is enough to test the pitch-to-close motion before paying for anything.
On the internal knowledge side, one operational note that trips people up: Additional Knowledge accepts DOCX, TXT, Markdown, and pasted text. PDF files are not supported, and PowerPoint files are not supported either, so convert those before you upload.
FAQ
Should the audit include branded prompts at all?
Yes, but score them separately. Branded prompts tell you whether the AI understands the entity. They tell you nothing about category discovery, which is what most buyers actually experience. Blend the two and your visibility number becomes meaningless.
How many prompts is enough for a credible baseline?
Enough to be defensible for the specific market, and honest about what a small set proves. A handful of prompts on one provider is a teaser, not a market read. If the sample is thin, sell a baseline and validation period rather than claiming comprehensive visibility. Padding the count by re-running the same prompt just manufactures false confidence.
Does fixing technical readiness improve AI recommendations?
It removes blockers. Google's guidance treats standard indexability and normal SEO fundamentals as the eligibility basis for its AI features, with nothing extra required. That makes readiness necessary, not sufficient. If you promise recommendations off the back of a technical fix, you are writing a claim you cannot measure or defend.
Is AI referral traffic the right success metric?
It is a downstream signal, not a stand-in for visibility. Google folds AI-feature traffic into overall Search Console traffic under the "Web" type rather than breaking it out cleanly. Visibility, citation, click, and conversion are four different things. Report them separately or you will confuse yourself and the client.
Further reading
Users also found this interesting
Continue with guides covering the same topic and workflow.
How to Structure Content So AI Assistants Cite You
Learn how to format and structure your content so AI assistants like ChatGPT, Claude, and Perplexity cite your brand in their answers.
How to Structure Website Content for AI Trust
Learn how to structure website content to earn AI trust and citations from ChatGPT, Claude, and Perplexity. Practical steps for brands in 2026.
How to Choose an AI Content Optimization Platform for Brand Authority
Choose an AI content platform that builds entity authority and earns citations—not just keyword-optimized drafts for traditional search.
Ready to optimize your content for AI?
Start creating AI-native content that gets discovered and recommended by leading AI systems.