Authority Engine

90 Day AI Visibility Audit: A Marketer’s Revenue Roadmap

Marketer reviewing an AI visibility audit

An AI visibility audit measures whether your brand is mentioned, cited, and accurately described when people ask ChatGPT, Gemini, Perplexity, or Claude questions related to your category. The fastest way to start is to assemble a small set of 50 to 200 prompts and run them through one or two engines today. That first pass gives you a baseline report and usually surfaces a few quick wins before you build anything more formal.


TL;DR:

  • Focus unbranded queries higher in your audit, as they reveal the true opportunity for visibility and the risk of being overlooked in AI responses.
  • Maintain a prompt library of 200 to 600 questions, categorized into head-term, comparison, and buyer-intent prompts, for consistent trend analysis.
  • Cross-engine disparities mean that strong positioning in one AI platform does not guarantee similar results elsewhere, so treat each as a separate measurement challenge.
  • Prioritize fixing high-impact factual errors and outdated content on high-traffic pages first, before building new content for unrepresented queries.
  • Measure progress through branded-search lift over 30 days, citation tracking, and AI referral data, setting clear ownership and governance for ongoing visibility efforts.

Authority-engine
See Where Buyers Find You
Authority Engine helps established businesses assess competitive AI visibility and strengthen how buyers discover, understand, and consider them.
Explore AI visibility support

Table of Contents

Define the audit scope: goals, queries, and engine selection

Before you run a single prompt, decide what the audit needs to prove. An AI visibility audit that ignores business goals turns into a vanity metrics exercise, mentions for the sake of mentions. Tie every query you test back to a stage of the funnel: awareness queries show whether AI engines know your brand exists, consideration queries show whether they describe you accurately against alternatives, and conversion queries show whether they would actually point a buyer toward you.

Query mix matters as much as query count. Branded queries (“what is [your brand]”, “[your brand] reviews”) tell you how engines currently frame your company. Unbranded queries (“best [category] for mid-size teams”, “how to choose a [category] provider”) tell you whether you show up at all when the buyer hasn’t named you yet. Most audits should weight unbranded queries higher, since that’s where the real opportunity and the real risk of invisibility sit.

Engine selection and cadence should match your resources, not an ideal.

  • Start with ChatGPT and Perplexity, since they’re widely used for research-style questions and easy to query manually.
  • Add Gemini once you have a working prompt set, since Google’s generative results increasingly shape top-of-funnel visibility.
  • Add Claude last unless your buyers skew technical or enterprise.
  • Run a lightweight pass weekly during the first month, then settle into a biweekly or monthly cadence once your baseline is stable.

This scoping work takes an afternoon and prevents the single most common mistake in AI visibility work: testing everything at once and drowning in data nobody acts on.

Build a prompt library and prompt-management workflow

A prompt library is the backbone of a repeatable audit. Without one, every test run becomes a one-off exercise that can’t be compared to the last. Build it from sources you already have rather than guessing at questions buyers might ask.

  1. Pull branded queries directly from Google Search Console, since those are the exact phrases people already type when looking for you.
  2. Mine your best-performing organic keywords for the unbranded topics where you already rank, since AI engines often cite pages that already have search traction.
  3. Collect the questions your sales team hears most often during discovery calls, since those map directly to buyer intent.
  4. Review your support and FAQ logs for recurring confusion points, since AI engines frequently surface the same misunderstandings.

Organize prompts into three working types once you’ve gathered them. Head-term prompts are broad category questions (“what is [category]”). Comparison prompts pit you against the field without naming a specific rival (“how do [category] providers differ”). Buyer-intent prompts mirror a late-stage decision (“what should I look for before hiring a [category] provider”). A balanced library usually runs 200 to 600 prompts once you’ve covered all three types across your priority products or services; fewer than 200 makes trend detection unreliable, and more than 600 becomes difficult to review by hand on a recurring basis.

Track everything in a single sheet with consistent columns: prompt text, prompt type, engine, date tested, brand mentioned (yes/no), position in the answer, cited domains, sentiment, and a notes field for anything that needs follow-up. The schema matters more than the tool. A plain spreadsheet with disciplined columns beats a fragmented set of screenshots every time.

Platform testing: how to run prompts across AI engines and capture results

Running the actual tests is the most mechanical part of the audit, but small inconsistencies here quietly wreck your ability to compare results over time. Decide upfront whether you’re capturing responses through each platform’s consumer interface or through an API, since the two can return different answers for the same prompt and you don’t want to mix methods inside a single tracking period.

  • UI capture is free and fast but harder to scale past a few hundred prompts per session.
  • API access gives you consistent formatting and timestamps, which matters once you’re comparing month over month.
  • Whichever method you choose, keep it consistent within a testing cycle so differences you see reflect real visibility changes, not a change in how you collected the data.

For each prompt, log the fields that actually drive decisions later: whether your brand was mentioned at all, which domains and specific pages were cited as sources, the sentiment or framing used to describe you, which engine produced the response, and the exact timestamp. Citation data is the most valuable and the most perishable field, since engines update their source selection frequently and a page cited today may not be cited next month.

A few practical notes save hours of frustration. Rate limits on consumer interfaces mean you’ll want to space out large prompt batches rather than firing them in rapid succession. Reproducibility is limited by design, since generative engines don’t return identical answers to identical prompts, so plan to run each prompt two or three times and record the pattern rather than treating a single response as ground truth. Cross-engine consistency is rare: a 2026 analysis of 126 million AI search prompts found that agreement between engines on which brands to cite as top sources sits at roughly a third, which means a strong showing in one engine tells you very little about your standing in another, according to Semrush’s 2026 AI Visibility Index. Treat each platform as its own measurement problem rather than assuming visibility transfers between them.

Platform testing: how to run prompts across AI engines and capture results — overview diagram

Analyze branded and unbranded responses for accuracy and sentiment

Once you have responses logged, the real analytical work starts: deciding whether what the engines are saying about you is actually correct. Build a simple three-tier rubric before you start reading responses, so you’re not making judgment calls prompt by prompt.

  • Correct: the response accurately describes your offering, positioning, or facts about your business.
  • Omission: the response is technically accurate but leaves out a detail that matters for the buyer’s decision, like a certification or service area.
  • Misleading: the response states something false or materially distorts how you operate.

Misleading classifications deserve immediate attention regardless of how often they occur, since a single wrong claim repeated across thousands of buyer conversations compounds quickly.

Sentiment and framing are harder to measure at scale than a simple mention count, and they’re worth treating with more skepticism. Sample a subset of responses for human review rather than trying to score sentiment on every single one; a reviewer reading 20 to 30 responses per engine per month will catch tonal shifts that automated scoring tends to miss. Sentiment also moves more than raw mention rate does: a brand can be mentioned consistently while the framing swings from neutral to favorable to lukewarm across different sessions, because generative responses are shaped by phrasing, recent training updates, and the specific sources an engine happens to cite that day. Mention rate is a more stable trendline; treat sentiment as a qualitative signal you check periodically rather than a number you chart weekly.

Across large prompt studies, brand visibility follows a stature ladder: global household brands appear most often in AI answers, mid-market brands appear less, and niche or local brands appear least, with gaps of roughly 30 percentage points between tiers. According to research on generative engine optimization at scale, for established but non-household brands, this means closing the gap with better-known competitors in generic prompts is a longer project than winning visibility on specific, narrower buyer-intent queries where brand size matters less.

Find top-cited pages and which page formats earn citations

Every time an AI engine answers a prompt with a source attached, it’s telling you something concrete: this is a page it trusts enough to cite. Aggregating those citations across your full prompt run turns scattered observations into a prioritized list of pages worth protecting or improving.

Start by ranking cited pages, both yours and third-party, by how often they show up across your prompt set and how persistent that citation is across repeated testing cycles. A page cited once might be noise; a page cited across multiple engines and multiple months is doing real work for whoever owns it.

Page format matters more than most marketers expect.

  • Listicles and other list-structured pages are consistently among the most-cited formats in citation datasets.
  • FAQ-style pages with clear, direct question-and-answer structure also perform well, since they map closely to how buyers phrase prompts.
  • Data-dense pages with specific figures, comparisons, or named criteria tend to outperform generic marketing pages that make broad claims without specifics.

This pattern holds across large-scale studies of AI citation behavior, which find that structural formats like listicles and FAQs earn citations more reliably than long-form narrative content, reinforcing a shift from a pure ranking mindset toward one focused on structural authority and verifiable information.

Once you’ve identified your top-cited owned pages, cross-check them against Google Search Console and your analytics platform. A page earning AI citations but showing flat organic traffic might be reaching buyers through a channel your standard reporting doesn’t capture, which is exactly the kind of signal worth digging into before you decide where to invest next.

Benchmark share-of-voice against category peers without naming competitors

You can’t judge your visibility in isolation. A peer benchmark tells you whether a low mention rate reflects a category-wide pattern or a problem specific to your brand.

Build a peer set using generic tiers rather than a list of named rivals: top-tier national players, regional or mid-market providers, and category specialists who serve a narrower niche than you do. Testing your prompt library against all three tiers shows you where you sit on the same stature ladder that shapes mention rates across large studies, and whether the gap is a size problem or a content problem.

  • Run your full prompt set and tag each response by which tier of peer, if any, got mentioned alongside or instead of you.
  • Calculate a rough share-of-voice per engine: what percentage of relevant prompts mention your brand versus any peer tier at all.
  • Flag the specific queries where peers appear consistently and you don’t, since those are your highest-opportunity targets.

Once you have that gap list, turn it into two kinds of work. Content briefs address the queries where your own pages should be competing but aren’t yet structured to earn a citation. Outreach briefs address the queries where third-party sources, not your own site, are the ones getting cited, meaning the fix involves earning mentions on pages you don’t control rather than fixing pages you do.

Turn audit findings into a prioritized action plan and 90-day roadmap

An audit that ends in a spreadsheet nobody acts on isn’t worth running. The findings need to become a prioritized plan with owners and deadlines attached within a week of finishing the analysis.

Score every finding on impact and effort before you build the roadmap. A factual error on a high-traffic page that’s getting cited across multiple engines is high impact and low effort: fix it first. A content gap on a competitive unbranded query is high impact but higher effort, since it requires new content and possibly outreach.

  1. Week 1 to 2, quick wins: correct factual errors identified in your accuracy review, refresh existing pages that are already earning citations but contain outdated information, and add direct FAQ answers to pages that currently lack them.
  2. Week 3 to 6, mid-term plays: pursue third-party mentions on pages you don’t own but that showed up repeatedly in your citation analysis, build out listicle-style or comparison content for your highest-opportunity unbranded queries, and publish executive-authored content addressing buyer-intent questions where your brand was absent.
  3. Week 7 to 12, consolidation: re-run the full prompt panel, compare citation and mention rates against your baseline, and reprioritize the next cycle based on what actually moved.

Assign a single owner to each workstream, usually content, PR, or a dedicated AI visibility lead, and set a review checkpoint at the 30, 60, and 90 day marks so the roadmap doesn’t quietly stall after the initial push.

Pro Tip: Fix factual errors on your highest-traffic pages before chasing new content; a wrong claim repeated across engines undermines every other citation you earn.

Measurement and KPIs: the operational stack to prove progress

AI engines don’t expose citation impressions or click data the way traditional search does, so proving the audit’s work paid off requires stitching together signals from several sources rather than relying on one dashboard.

The most reliable lagging indicator is branded-search lift measured in Google Search Console on a rolling 30-day basis. Brands cited in AI-generated answers see branded search volume rise by roughly 23% over the following 30 days, according to Semrush’s 2026 AI Visibility Index, which makes this the single clearest signal that a citation event translated into actual buyer interest rather than just a mention.

Alongside branded-search lift, track these supporting signals on a monthly cadence:

  • Panel-based citation tracking across your prompt library, since this is the only direct way to see whether your pages are actually being cited over time.
  • AI referral traffic filtered inside GA4, which shows visits arriving from AI platforms even though click-through from citations tends to run low overall.
  • Citation persistence, meaning whether a page cited last month is still being cited this month, since engines rotate sources more than most teams expect.
  • Downstream conversion lift on pages that gained new citations, to confirm the visibility is reaching the right stage of the funnel.

Set up the Search generative AI control in Search Console as part of your baseline technical review. Google’s guidance confirms this setting lets you include or exclude site content from certain AI features, and that changes can take several days to propagate, so check it early rather than assuming it’s configured correctly by default. Report on this full stack monthly rather than weekly; AI visibility signals are noisy enough that weekly swings are rarely meaningful, while a month of data usually separates real movement from measurement noise.

Tools, templates, and how a managed program organizes ongoing AI visibility work

Running this audit by hand is realistic for a first pass, but sustaining it month after month usually means building out a lightweight stack rather than relying on memory and a scattered folder of screenshots.

Your prompt-tracking sheet needs a consistent set of columns to stay usable past the first few cycles: prompt text, prompt type, target engine, test date, mention status, cited domains, sentiment tag, and an action-needed flag for anything requiring follow-up. Keep the schema identical across testing cycles so month-over-month comparisons actually mean something.

Three categories of tools are worth considering as your volume grows.

  • Panel runners that automate sending your prompt library across multiple engines on a schedule, saving the manual UI work described earlier.
  • Citation monitors that track which domains and pages are being cited over time without requiring manual logging for every prompt.
  • Analytics connectors that pull AI referral data into your existing GA4 and Search Console dashboards so you’re not checking four separate places.

Managed programs that run this work continuously tend to combine all three categories with ongoing content and outreach execution, since the audit itself is only the diagnostic step: the harder, recurring work is producing the pages and third-party mentions that move the needle between audit cycles. Authority Engine’s managed AI visibility work is built around exactly that combination, coordinating buyer-question research, content publishing, contextual backlinks, and ongoing monitoring under a single strategy rather than treating the audit as a one-time report. Strategic direction and quality oversight are provided across client engagements, with doctoral research into AI-assisted buyer discovery informing how the program prioritizes its work.

Governance matters as much as measurement

The biggest failure mode in AI visibility work isn’t a bad prompt library. It’s treating the audit as a one-time project with no owner, so the findings sit in a shared drive while the next quarterly content calendar gets built without them.

A single owner needs clear authority here, not just a mandate to report findings. A simple RACI works: marketing owns the audit and reporting, content and SEO are responsible for execution, PR and comms are consulted on any third-party outreach, and leadership stays informed on the branded-search lift numbers each quarter. Without that structure, editorial and public relations teams end up describing the company differently across channels, and AI engines pick up on exactly that inconsistency when they synthesize an answer from multiple sources. Treat this as a standing discipline with a recurring calendar slot, not a report that gets filed once and revisited only when someone asks why a competitor seems to be winning more AI mentions.

— Patrick

A faster path: Authority Engine’s AI Visibility & Revenue Opportunity Report

Running this audit internally takes real time from a team that likely already has a full content calendar. Authority Engine’s AI Visibility & Revenue Opportunity Report gives you a baseline read on where your brand currently stands, how you compare to relevant peers, and illustrative revenue scenarios that show what stronger visibility could mean for your pipeline, without requiring your team to build the prompt library and tracking sheet themselves.

Authority-engine

For businesses that want the audit findings turned into ongoing execution rather than a static report, Managed AI Visibility coordinates the research, content, backlinks, and monitoring described throughout this article as a continuous program rather than a one-time exercise. Visit Authority Engine to see how a managed engagement could start with your business.

FAQ

What does an AI visibility audit actually measure?

An AI visibility audit measures whether your brand appears in responses from AI engines like ChatGPT, Gemini, Perplexity, and Claude, how accurately those responses describe you, and which pages the engines cite as sources. It typically combines a prompt panel, citation tracking, and branded-search data to build a full picture rather than relying on a single metric.

How many prompts do I need for a reliable baseline?

Most practitioners build a working library of 200 to 600 prompts to get a reliable trend, though a first pass with as few as 50 prompts across one or two engines can surface obvious gaps worth fixing immediately. Fewer prompts make it harder to tell a real shift from normal variation in generative responses.

Why do engines disagree so much about which brands to cite?

Cross-engine agreement on top-cited brands is limited, with one large-scale study finding only about a third agreement across engines on which sources to cite, according to Semrush’s 2026 AI Visibility Index. Each engine draws on different training data and live retrieval sources, so strong visibility on one platform doesn’t predict strong visibility on another.

What’s the best KPI for proving AI visibility work is paying off?

Branded-search lift measured in Search Console on a rolling 30-day basis is the most widely used indicator, since citation events have been shown to lift branded search volume by roughly 23% over that window, according to Semrush’s research. Pair it with panel-based citation tracking and AI referral data in GA4 for a fuller measurement stack.

Do I need special structured data for my site to appear in AI answers?

No special structured-data files are required specifically for AI features; eligibility depends on standard indexing and technical SEO fundamentals, according to Google Search Central’s guidance. You can also manage whether your content is eligible for certain AI features using the Search generative AI control in Search Console.

Sources

Created with BabyLoveGrowth tools