Attrifast
ProductAI visibilityPricingDocsBlog
Log inStart free trial
ProductAI visibilityPricingDocsBlogLog in
Product
  • Track Website Traffic
  • Attribution Software
  • Website Visitor Tracking
  • SEO Dashboard
  • Analytics for SaaS
  • Revenue by Source
  • Traffic Source Tracking
  • Revenue Channel Attribution
  • Revenue Attribution
  • Privacy-First Analytics
  • Cookieless Analytics
  • UTM to Revenue
  • AI Visibility Score
  • Share of Voice (AI)
  • Prompt Tracking
  • AI Citation Tracking
  • ChatGPT Rank Tracker
  • AI Revenue Attribution
  • Pricing
Track AI Traffic
  • Track ChatGPT Traffic
  • Track Perplexity Traffic
  • Track Claude Traffic
  • Track Gemini Traffic
  • Track AI Overviews
  • Track Copilot Traffic
  • ChatGPT Revenue Attribution
  • Perplexity Revenue Attribution
  • Claude Revenue Attribution
  • Gemini Revenue Attribution
  • AI Visibility to Revenue
Use Cases
  • Stripe Analytics
  • Shopify Analytics
  • Stripe Attribution
  • For Bootstrapped SaaS
  • Affordable Attribution
Compare
  • vs Profound
  • vs Loamly
  • vs Peec AI
  • vs Otterly
  • vs Cometly
  • vs Segment
  • vs Google Analytics
  • vs Plausible
  • vs Fathom
  • vs Simple Analytics
  • vs PostHog
  • vs Matomo
  • vs Umami
  • vs Pirsch
  • vs Mixpanel
  • vs Amplitude
  • vs Heap
  • vs Hyros
  • vs AnyTrack
  • vs DataFast
  • vs Similarweb
  • All comparisons
Resources
  • GEO Hub
  • AEO Hub
  • AI Search Hub
  • Research
  • Best Conversion Tracking Software
  • ChatGPT vs Google Traffic
  • Mixpanel Alternative
  • Track Channel Revenue
  • First vs Last Touch
  • Cookieless Conversion Tracking
  • GA4 Attribution Limits
  • CAC by Channel
  • Stripe Conversion Tracking
  • Stripe Revenue Tracking
  • AEO vs SEO 2026
  • AI Traffic Benchmark
  • How to Rank in ChatGPT
  • Best AEO Tools 2026
  • Measure GEO ROI
  • Schema for AI Search
  • Dark AI Traffic in GA4
  • What Is Referral Traffic?
  • What Is Direct Traffic?
  • What Is Cookieless Analytics?
  • What Is Conversion Attribution?
  • AI Share of Voice
  • Documentation
  • View all posts
  • Multi-Touch Attribution
  • Free Tools
  • UTM Builder
  • UTM Checker
  • ROI Calculator
  • AI Readiness Checker
  • AI Visibility Checker
  • AI Crawler Directory
  • SEO + GEO Workflow
Company
  • About
  • Contact
  • Return Delay Penalty
  • Backlink RPV Scoring
  • AI Instructions
  • Live Demo
  • FAQ
  • Log in
© 2026 Attrifast · built by Vincent Ruan & Jessica Huang
AboutContactTermsPrivacy
Blog / AI Search

AI Visibility Prompts: The 104-Prompt Library, and the Citation Ladder That Decides Which Ones Work

18 min readPublished Aug 2026
Vincent Ruan
Vincent RuanFounder, Attrifast · August 18, 2026 · 18 min read

We ran 120 prompt-engine combinations against our own domain on one day. Branded comparison prompts got us cited 94% of the time. Narrow capability prompts, 40%. Broad category prompts — 'best tools for X' — went 0 for 84. Here is the full prompt library, the ladder that explains the gap, and the run methodology that survives a 52% flip rate.

AI visibility prompts: the citation ladder measured on one domain in a single scan — branded comparison prompts cited in 15 of 16 runs (94%), narrow capability prompts in 8 of 20 (40%), and broad category prompts such as best tools for X in 0 of 84 runs (0%) — with a 52% run-to-run flip rate meaning every prompt must be run three to five times before its result is trustworthy

Part of the AI Search Hub, AEO Hub, and the generative engine optimization guide.

TL;DR

  • Your measured AI visibility is mostly a property of which prompts you chose. Same domain, same four engines, same day: 94% citation rate on branded comparison prompts, 0% on broad category prompts.
  • We ran 120 prompt-engine combinations on our own domain. The results sort into a clean three-rung ladder: branded comparison 15/16 (94%) → narrow capability 8/20 (40%) → broad category 0/84 (0%).
  • The pattern is not ours alone: an independent study of 175 brands across 8 AI platforms found 96% described accurately when asked by name, but 89% never surfaced at all in category-research questions.
  • Answers are not deterministic. 52% of our ever-cited prompt-engine pairs did not hold the citation on every run. Run each prompt 3–5 times and track appearance rate, never rank.
  • Below: the 104-prompt library organized by rung, the selection framework, and the run methodology. Score your own prompt set with Attrifast — free trial, no card required.

There is a genre of AI-visibility advice that hands you a list of prompts and wishes you luck. This is not that, because the list is the easy half. The hard half is that the prompts you pick determine the number you get, to a degree that makes most published AI visibility scores close to meaningless as comparisons.

Here is the evidence, from our own domain, and it is not flattering.

On 16 August 2026 we ran a scan of 30 prompts across four engines — ChatGPT, Claude, Gemini and Perplexity — for a total of 120 prompt-engine combinations against attrifast.com. Overall citation rate: 19%. That single number is useless. Split by the kind of prompt asked, the same scan says three completely different things:

Prompt typeRunsCitedRate
Branded comparison — [our brand] vs [alternative]161594%
Narrow capability — a specific job our product does20840%
Broad category — "best tools for…", "top platforms for…"8400%

Eighty-four consecutive runs with zero citations. Same domain, same engines, same hour as the 94%. We publish that cell because a prompt library demonstrated only on the prompts that worked is an advertisement, and because the gap between those rows is the single most useful thing we have learned about prompt tracking.

Quick facts

QuestionAnswer
What is an AI visibility prompt?A full question you send to an AI engine on a schedule to check whether its answer includes your brand or domain
How many should you track?25–40 for one product line, weighted toward commercial intent
How many times per prompt?3–5 runs per engine — 52% of our ever-cited pairs did not hold the citation across runs
Which prompts get cited most?Branded comparison (94% in our scan), then narrow capability (40%), then broad category (0%)
What metric should you record?Appearance rate across repeated runs, never rank within one answer
How often to re-run?Monthly for the full set; weekly for the prompts attached to converting pages
What moves broad-category visibility?Third-party presence — 99.99% of citations in one 49,391-citation study pointed to sites other than the brand's own
Citation rate by prompt type — attrifast.com, 120 runs across 4 engines, 16 Aug 2026
Citation rate by prompt type — attrifast.com, 120 runs across 4 engines, 16 Aug 2026

Source: Attrifast AI visibility scan, attrifast.com — single scan, four engines, 30 prompts

What an AI visibility prompt actually is

A keyword is a fragment aimed at a ranking system. An AI visibility prompt is a complete question aimed at an answering system, and you monitor it to find out whether the answer contains you.

Three practical differences follow, and each one changes how you build the list.

Length and grammar. Our own tracked prompts average 8.5 words for commercial intent and 9.1 words for informational, against the two-to-three-word norm most people still type into a search box. Prompts are sentences. They contain qualifiers — a buyer type, a constraint, a use case — and those qualifiers are the part that decides whether you appear.

The unit of success. There is no position three. Either the generated answer includes your domain or brand, or it does not. And because the same question asked twice can produce different answers, the honest metric is an appearance rate across repeated runs, not a rank in a single answer. We go into how that differs from rank tracking in what is prompt tracking.

Citation and mention are two different events. This is the trap almost every new setup falls into. A July 2026 analysis of 12,000 AI responses across roughly 2,967 prompts and 200 brands found that only 23.1% of brand mentions came with a citation — in the other 76.9%, the answer named the brand and gave the reader nothing to click. Running the relationship the other way, 69.9% of citations did name the brand in the response text, leaving roughly three in ten links that pass credit no reader ever sees attached to you. Track one event and you are measuring a fraction of your presence.

Of brand mentions in AI answers, how many come with a citation
Of brand mentions in AI answers, how many come with a citation

Source: BuzzStream, July 2026 — 12,000 AI responses, ~2,967 prompts, ~200 brands across Google AI Mode, AI Overviews, Gemini and GPT

The citation ladder

Sort any prompt set by how much of your brand is already inside the question, and citation rate falls off a cliff in a very specific shape. We call it the citation ladder. Every rate below is measured, from the single clean scan described above.

Rung 1 — Branded comparison: 94% cited (15 of 16 runs)

[Your brand] vs [alternative] for [specific job]. Four such prompts, four engines each. Three were cited in all four engines; the fourth in three of four.

This rung is defence, not discovery. The question already contains your name, so the engine has an obvious reason to fetch your domain. What you learn is not reach but accuracy: whether the engine describes your product correctly, whether it is quoting a stale price, whether the comparison it draws is one you would make. Those answers reach people at the last moment before a purchase decision, which is why a wrong one is expensive.

What it cannot tell you: whether anyone who has never heard of you will ever meet you.

Rung 2 — Narrow capability: 40% cited (8 of 20 runs)

Unbranded, but narrow enough that only a handful of products in the world are plausible answers. From our set, the prompts that earned citations here read like:

  • "how to set up revenue attribution from AI search engines" — cited in 3 of 4 engines
  • "best platform for attributing revenue to AI search channels" — 2 of 4
  • "best conversion tracking software for AI-driven website traffic" — 1 of 4
  • "cookieless analytics platforms for tracking multi-channel revenue attribution" — 1 of 4
  • "how to measure ChatGPT and Perplexity traffic separately" — 1 of 4

This is the rung that matters most, and the one most prompt sets under-weight. It is unbranded — so a citation here is genuine discovery by someone who did not know your name — but specific enough that the field of candidate answers is small. This is where a new AEO effort shows movement first, months before anything happens on rung 3.

Rung 3 — Broad category: 0% cited (0 of 84 runs)

"What is the best tool for tracking AI search visibility". "What are the top AEO tools for 2026". "best tools for monitoring AI citations across all search engines". "top revenue attribution tools for SaaS companies 2026". Twenty-one prompts of this shape, four engines each, zero citations.

These are the prompts everybody wants to win, and they behave differently from the other two rungs because the answers are not assembled from vendor domains at all. A July 2026 study of 175 brands across eight AI platforms put hard numbers on exactly this gap. The same engines that described 96% of brands accurately when asked about them by name left 89% of those brands out entirely when asked a category-research question. Of the 49,391 citations analyzed, 99.99% pointed to third-party websites. And brands with fewer than 2,000 indexed web pages mentioning them appeared in AI answers just 3% of the time.

That 96%-versus-89% pair is the citation ladder restated in someone else's data: engines know who you are, and still do not bring you up unprompted.

Which third parties? The 5W Citation Source Audit for Q1 2026, synthesizing nine independent datasets, puts Wikipedia at 13.15% and Reddit at 11.97% of U.S. ChatGPT citations — over a quarter of all citations between just those two — with YouTube, LinkedIn and Forbes close behind, and the major business newspapers absent from the top 20 entirely. Our own aggregate scan data agrees: after the engines' own redirect infrastructure, the most frequently cited domains across all scans on our platform were reddit.com, youtube.com, github.com and arxiv.org.

The implication is uncomfortable but freeing: you do not win rung 3 by writing a better page on your own site. At 99.99% third-party citations, the page you publish is almost never the artifact that gets cited for a category question. You win it by being present on the surfaces the engines already trust — a PR, community and indexed-footprint problem on a quarters-long clock. Track rung 3 as a long-term scoreboard. Do not judge a three-month AEO programme by it.

The citation ladder: how much brand is in the question vs how often you are cited
The citation ladder: how much brand is in the question vs how often you are cited

Source: Attrifast AI visibility scan, attrifast.com, 16 Aug 2026 — 120 prompt-engine runs

The commercial-intent lift replicates

If the ladder were a quirk of one domain it would not be worth a framework. It is not. The same directional finding appears in three independent datasets, measured three different ways.

DatasetCommercial / comparativeInformationalLift
attrifast.com, single scan, 120 runs27.9% cited8.3% cited3.4×
Three other domains on our platform, 440 runs, aggregated44.1% cited16.8% cited2.6×

Two first-party samples, different industries, and the lift lands between 2.6× and 3.4×. Commercial-intent prompts are roughly three times likelier to surface a brand than informational ones.

Independent research finds the same cliff using different metrics entirely, which is the strongest kind of corroboration — the pattern survives a change of instrument. BuzzStream's analysis of 12,000 AI responses measured how often a brand mention also carried a citation and found 39% for single-brand queries, 35.8% for head-to-head comparisons, and 7.2% for list and category queries. Victorious, testing 175 brands across eight platforms, found brands named in 0.10% of answers to problem-awareness prompts against roughly twelve times that rate for category-research prompts. Different metrics, different samples, same shape: intent and specificity decide whether a brand appears at all.

Mention-to-citation rate by query type — an independent replication of the ladder
Mention-to-citation rate by query type — an independent replication of the ladder

Source: BuzzStream, July 2026 — 12,000 AI responses, ~2,967 prompts, ~200 brands across 10 industries

Commercial vs informational prompts across three independent datasets
Commercial vs informational prompts across three independent datasets

Source: Attrifast first-party scans — attrifast.com (120 runs) and three other platform domains (440 runs, aggregated and anonymized)

One caution on reading that table: the middle row is aggregated across three other domains on our platform, anonymized, with no per-site breakdown — different industries, different competitive densities. It is corroboration, not a benchmark to compare yourself against.

Run it more than once, or do not bother

Everything above is worthless if you run each prompt a single time, because AI answers are not deterministic.

In our own data, of the 25 prompt-engine pairs that were cited at least once across repeated runs, 13 — 52% — did not hold that citation on every run. Half the wins were coin flips.

Independent research is blunter still. SparkToro's January 2026 study had 600 volunteers execute 2,961 prompt runs across 12 queries, each repeated 60–100 times. Their finding: there is under a 1-in-100 chance that ChatGPT or Google's AI returns the same brand list when asked the same question repeatedly, and roughly 1 in 1,000 for identical ordering. What was stable was appearance rate — a top brand showed up in 85 of 95 responses, another in 69 of 71.

Of prompt-engine pairs ever cited, how many held the citation on every run
Of prompt-engine pairs ever cited, how many held the citation on every run

Source: Attrifast first-party scan data — 25 prompt-engine pairs with 2+ runs, buggy scan window excluded

Three rules follow, and they are not negotiable:

  1. Run each prompt 3–5 times per engine. Below three you are reporting a dice roll as a measurement.
  2. Record appearance rate, never rank. "Cited in 7 of 10 runs" is a real metric. "Ranked #2 in ChatGPT" is a screenshot.
  3. Re-run the full set after any major model release. Historical baselines break on those dates; a drop the week of a model update is usually the model, not you.

Attrifast runs your prompt set across ChatGPT, Claude, Gemini and Perplexity on a schedule, records appearance rate across repeated runs instead of a single screenshot, and joins the engines that cite you to the revenue they actually send.

Score your prompt set free →

How to choose your 30: five filters

The library below has 104 prompts. You should not track 104. Run every candidate through these five filters and keep 25–40.

1. Intent first. Weight the set toward commercial intent, given the 2.4–3.4× lift. A useful starting split for a single product line: 40% narrow capability, 25% broad category, 20% branded and comparison, 15% informational. The broad-category quarter is a long-term scoreboard you expect to score zero on at first — include it anyway, so you can prove movement later.

2. Buyer language, not category language. Write the prompt the way a buyer describes their problem, not the way your industry names it. A buyer types "how do I tell which channel my Stripe payments came from"; the category says "multi-touch attribution platform". Both belong in a set, but the first predicts discovery and the second mostly measures whether you have won a taxonomy nobody outside your market uses.

3. Answerable specificity. If a hundred products could plausibly answer the prompt, it belongs on rung 3 and will score zero for a long time. If three to ten could, it is rung 2 and will move. This filter is the single highest-leverage edit you can make to an inherited prompt set.

4. Revenue traceability. For each prompt, name the page you would want the engine to cite, and know whether that page converts. A prompt whose ideal landing page has never produced a signup is a vanity prompt — real, trackable, and worth nothing when it improves. This is the filter that turns prompt tracking from a scoreboard into a budget input.

5. Stability enough to trend. Prompts that name a year ("best X in 2026") or a fast-moving feature will churn on their own schedule and pollute your trend line. Keep a few for freshness coverage; do not build the core of the set from them.

The library: 104 AI visibility prompts

Replace every [bracket] with your own terms. [product] is your product category as a buyer would say it, [job] is the specific outcome someone hires you for, [ICP] is the customer type, [constraint] is the requirement that rules competitors out.

Each block is tagged with the rung it sits on, so you know what a score there is telling you. Several blocks share a rung — three of them sit on rung 2, because that is where most of the prompts worth tracking live. If you would rather not run these by hand across four engines, Attrifast tracks a set like this on a schedule and records appearance rate across repeated runs.

Rung 1 — Branded and comparison (12 prompts) · defence · expect a high rate

  1. [Your brand] vs [alternative] for [job]
  2. Is [your brand] or [alternative] better for [ICP]?
  3. What is [your brand] and who is it for?
  4. [Your brand] pricing — what does it actually cost for [ICP]?
  5. What are the main criticisms of [your brand]?
  6. Is [your brand] worth it for a [ICP] doing [job]?
  7. What does [your brand] do that [alternative] does not?
  8. Is [your brand] a good fit if I need [constraint]?
  9. [Your brand] alternatives for [ICP]
  10. Has anyone switched from [alternative] to [your brand]? What changed?
  11. Does [your brand] integrate with [key integration]?
  12. Is [your brand] safe / compliant for [regulated context]?

Read these for accuracy before reach. A stale price or an invented limitation on this rung costs you deals directly.

Rung 2 — Narrow capability (28 prompts) · discovery · this is where movement shows first

  1. How do I [job] without [common blocker]?
  2. Best tool for [very specific job] for [ICP]
  3. What software lets me [specific capability] and [second capability] in one place?
  4. How do I measure [specific metric] separately from [adjacent metric]?
  5. Which platform connects [data source A] to [data source B] for [outcome]?
  6. How do I set up [specific workflow] for a [ICP]?
  7. Tool for [job] that works without [thing your buyers reject]
  8. What is the fastest way to [job] for a small team?
  9. How do I track [entity] across [channel A] and [channel B]?
  10. Best [product] for teams that need [constraint]
  11. How do I attribute [outcome] to [source] accurately?
  12. What can I use instead of [generic incumbent] for [narrow job]?
  13. Which [product] handles [edge case] properly?
  14. How do I get [specific report] without building it myself?
  15. Lightweight [product] for [ICP] who only needs [narrow job]
  16. How do I know if [problem] is happening on my [asset]?
  17. What is the simplest setup for [job] on [platform]?
  18. [Product] that supports [specific technical requirement]
  19. How do I automate [manual process] for [ICP]?
  20. Best way to [job] when [constraint] rules out the obvious option
  21. Which tools give [specific data] at the [granularity] level?
  22. How do I join [dataset A] to [dataset B] without engineering help?
  23. What should a [ICP] use to monitor [specific signal]?
  24. How do I [job] if I am already using [common stack component]?
  25. Tool that shows [outcome] per [dimension] rather than in aggregate
  26. How do I validate that [system] is measuring [metric] correctly?
  27. Cheapest way to [job] properly for a [company stage]
  28. What breaks when you try to [job] with [generic approach]?

Every prompt here should map to a page you own. If it does not, that is a content brief, not a tracking gap.

Rung 3 — Broad category (18 prompts) · long-term scoreboard · expect zero at first

  1. Best [product] in [year]
  2. Top [product] for [industry]
  3. What is the best [product] for [ICP]?
  4. [Product] ranked by features and price
  5. Most popular [product] right now
  6. [Product category A] vs [product category B] — which should I use first?
  7. Alternatives to [dominant generic incumbent]
  8. Best free [product]
  9. Enterprise [product] options
  10. Open-source [product] alternatives
  11. What do [ICP] actually use for [job]?
  12. Best [product] for beginners
  13. Most accurate [product]
  14. [Product] with the best support
  15. Which [product] has the easiest setup?
  16. Fastest-growing [product] in [year]
  17. [Product] shortlist for a [company stage] company
  18. What should I look for when choosing a [product]?

Score these quarterly. Movement here follows third-party presence, not on-site publishing.

Rung 2 — Problem and symptom prompts (20 prompts) · the highest-conversion discovery prompts

These describe a pain rather than a product. They are how people who do not yet know your category actually search, and they convert unusually well because the reader is mid-problem.

  1. Why is my [metric] dropping even though [expected cause] looks fine?
  2. Why does [system A] show a different [metric] than [system B]?
  3. [Symptom] — what causes it and how do I fix it?
  4. My [asset] is getting [traffic type] but no [outcome] — why?
  5. How do I find out where my [outcome] is actually coming from?
  6. Is it normal for [metric] to be [surprising value]?
  7. How do I stop [undesirable outcome] without losing [desirable thing]?
  8. What does it mean when [signal] but [contradictory signal]?
  9. I cannot tell whether [channel] is working — how do I check?
  10. Why is most of my [category] showing up as [catch-all bucket]?
  11. How do I prove [channel] drove [outcome] to my [stakeholder]?
  12. [Tool] is missing [data] — what are my options?
  13. How do I diagnose [problem] step by step?
  14. What should I check first when [failure mode] happens?
  15. Is [worrying pattern] a problem or normal variation?
  16. How do I get [stakeholder] to trust our [metric]?
  17. We are spending on [channel] but cannot see returns — how do we measure it?
  18. What is the minimum setup to answer "[core business question]"?
  19. How do I audit our current [system] for [failure mode]?
  20. Everyone says to use [popular approach] — does it work for [constraint]?

Rung 2 — Integration and stack prompts (12 prompts) · high intent, low competition

  1. How do I connect [your category] to [popular tool]?
  2. Does [popular tool] support [capability] natively?
  3. Best [product] for a [named stack] setup
  4. How do I get [outcome] into [BI/reporting tool]?
  5. [Popular platform] plugin for [job]
  6. Can I use [product] alongside [incumbent] or do I have to switch?
  7. How do I migrate from [incumbent] to a [product] without losing history?
  8. What is the setup effort for [product] on [platform]?
  9. Does [product] work with [framework/CMS]?
  10. How do I send [data] from [source] to [destination] reliably?
  11. Which [product] has an API for [specific operation]?
  12. Self-hosted options for [product]

Rung 3 — Informational and definitional (14 prompts) · low rate · include a few, do not lead with them

  1. What is [core concept]?
  2. How does [core concept] work?
  3. [Concept A] vs [concept B] — what is the difference?
  4. Why does [concept] matter for [ICP]?
  5. Best practices for [discipline] in [year]
  6. What are the common mistakes in [discipline]?
  7. How has [discipline] changed since [inflection point]?
  8. Is [old practice] still relevant?
  9. What metrics matter for [discipline]?
  10. How do I explain [concept] to [non-expert stakeholder]?
  11. What does [industry jargon term] actually mean?
  12. How do I get started with [discipline] from scratch?
  13. What is a realistic benchmark for [metric] in [industry]?
  14. How long does [process] usually take?

Our own informational prompts scored 8.3%, and the two that earned citations were the two narrowest — the ones that named a specific capability rather than a concept. That is the pattern to copy if you keep informational prompts in the set: specific how-to beats definitional every time.

Turning a prompt set into a budget decision

A prompt set that produces a percentage and nothing else is a vanity dashboard. Three joins make it an operating input.

Join 1 — prompt to page. Every rung 2 prompt should name the page you want cited. When a citation appears, you learn which page earned it; when it does not, you have a content brief. This is also how you discover that engines are citing a page you did not intend to be your answer for that question.

Join 2 — engine to traffic. Citation rate and traffic are not the same thing, and they diverge wildly by engine. In the scan above our best engine by citation rate was ChatGPT at 26.7%, then Perplexity at 23.3%, Claude at 16.7%, Gemini at 10%. Part of that spread is structural rather than a verdict on us: engines differ in how many sources they attach to an answer and how tightly they couple naming a brand to linking it. BuzzStream measured that coupling per platform and found 28.3% of mentions cited on GPT against roughly 22% on Google's surfaces, with 92.7% versus about 68% running the other way. Compare your rate against that engine's own baseline, never across engines. Our traffic split does not match the citation ordering at all — and the mismatch is the point, because most AI-referred visits arrive with the referrer stripped and land in Direct unless something fingerprints them. We measured 65–82% of ChatGPT-originated visits landing in Direct in GA4-style setups, with the mechanics in dark AI traffic.

Join 3 — traffic to revenue. The only join that settles arguments. An engine that cites you constantly and sends nobody who buys deserves less attention than one that cites you rarely and sends buyers. We have written about that inversion, with our own numbers, in the measure-first recovery playbook.

Citation rate by engine — attrifast.com, 30 prompts each, 16 Aug 2026
Citation rate by engine — attrifast.com, 30 prompts each, 16 Aug 2026

Source: Attrifast AI visibility scan, attrifast.com — 120 prompt-engine runs, single scan

Do those three joins and the prompt set stops answering "how visible are we?" and starts answering "which prompt, on which engine, is worth a quarter of work?" — which is the only version of this question with a budget attached.

A 30-day starting plan

DaysDo thisYou should end with
1–3Draft 60 candidate prompts from the library; run the five filters; cut to 30A set weighted ~40/25/20/15 across the rungs
4–7Baseline all 30 across four engines, 3 runs eachAn appearance rate per prompt, not a screenshot
8–14Map every rung 2 prompt to a target page; note which pages do not existA content brief list ranked by prompt intent
15–21Instrument AI traffic separately from Direct; connect paymentsRevenue per engine, however small
22–30Publish or upgrade the two pages behind your highest-intent uncited promptsA re-run baseline and a first movement signal

Expect rung 2 to move first, rung 1 to be stable and occasionally wrong in ways worth fixing, and rung 3 to sit at zero for a quarter or more. Ours did.

Limitations of this data

The ladder is a strong pattern, not a law, and the honest caveats matter more than the headline:

  • One domain, one day, one scan. The 94/40/0 split comes from 120 prompt-engine runs on attrifast.com on 16 August 2026. It is a small sample from a young site in a competitive B2B category. A large brand with an established third-party footprint would very likely see rung 3 above zero.
  • The rung-3 zero is partly about us. Our domain does not yet appear on the review sites, forums and listicles engines pull from for category questions. That is a specific weakness of ours, not proof that every vendor scores zero there.
  • The corroborating 440 runs are aggregated across three other domains on our platform, anonymized, with no per-site breakdown and no industry controls. Treat that row as directional support, not a benchmark to measure yourself against.
  • Citation, not mention. Our first-party numbers count whether an engine cited our domain. The independent studies quoted here measure brand mentions in the answer text, which is a related but distinct event — the two overlap only partially, as the BuzzStream figures show.
  • The 52% flip rate rests on 25 prompt-engine pairs. It is directionally consistent with much larger independent work, but it is not itself a large sample.
  • Engine behaviour changes without notice. Every number here has a shelf life measured in months. That is an argument for running your own scans on a schedule, not for trusting anyone's published figures — including ours.

Rather than take our numbers for it: point Attrifast at your own domain, load your prompt set, and get your own version of the table above across ChatGPT, Claude, Gemini and Perplexity.

Run your own scan free →

Run this prompt set against your own domain

Attrifast scans your prompts across ChatGPT, Claude, Gemini and Perplexity on a schedule, records appearance rate across repeated runs rather than one screenshot, splits AI engines out of Direct traffic, and joins each engine to the revenue it produced. Every chart in this post is that product, running on our own account, unedited.

  • ✓One script tag and a Stripe key — live in minutes
  • ✓Cookieless, so no consent banner for analytics
  • ✓Every AI referral matched to the payment it produced
Start your free trial →

7-day free trial · $0 due today · then $9.99/mo · cancel anytime

Attrifast dashboard: prompt-level AI visibility with estimated value, revenue split by channel across ChatGPT, Google, Perplexity, Claude and Direct, competitor position tracking, and per-engine scan settings.

FAQ

How many prompts should I track for AI visibility?

Between 25 and 40 for a single product line, weighted toward commercial intent. The number matters less than the spread: a set that is all broad category prompts will report near-zero visibility no matter how good your content is, and a set that is all branded prompts will report near-perfect visibility while telling you nothing about whether new buyers can find you. Our own 30-prompt set produced a 94% citation rate on one rung of the ladder and 0% on another during the same scan. If you can only afford to watch a handful, watch the branded comparison prompts for defence and the narrow capability prompts for growth, because those are the two rungs where the number actually moves.

How many times do I need to run the same prompt before the result means anything?

At least three, and five is better. AI answers are not deterministic: in our own data, 13 of the 25 prompt-engine pairs that were ever cited across repeated runs — 52% — did not hold that citation on every run. Independent research points the same way. SparkToro's January 2026 study had 600 volunteers run 2,961 prompts and found under a 1-in-100 chance that ChatGPT or Google's AI returns the same brand list when asked repeatedly, and roughly 1 in 1,000 for identical ordering. A single run is a screenshot of a dice roll. Track appearance rate across runs, not rank in one answer.

Why do broad prompts like 'best tools for X' never mention my brand?

Because those answers are assembled from the sources engines trust most for category-level questions — established listicles, forums, review aggregators and encyclopedic pages — and a vendor's own domain is rarely among them. In our clean scan, broad category prompts went 0 for 84 runs on our own domain across four engines. This is not a content-quality verdict; the same domain was cited in 94% of runs on branded comparison prompts the same day. Broad prompts are worth tracking as a long-term scoreboard, but they are the wrong place to look for early progress, and the wrong prompts to judge a new AEO effort by in its first quarter.

Do branded prompts count as real AI visibility?

They count as defence, not discovery. A branded comparison prompt tells you what an AI engine says to someone who already knows your name and is checking you against an alternative — a late-funnel moment where a wrong or outdated answer costs you a deal directly. That is genuinely worth monitoring, and it is the rung where our own citation rate was highest at 94%. What branded prompts cannot tell you is whether anyone who has never heard of you will encounter you at all. That is what the category and capability rungs measure, which is why a healthy prompt set spans both.

What is the difference between a keyword and an AI visibility prompt?

A keyword is a fragment optimized for a ranking system; a prompt is a full question optimized for an answering system. The practical differences are length, intent and the unit of success. Our own tracked prompts average 8.5 words for commercial intent and 9.1 for informational, against the two-to-three-word norm of traditional search queries. And success is not a position — it is whether your domain appears in a generated answer at all, measured as an appearance rate across repeated runs rather than a rank on a page.

Should I track prompts that mention competitors by name?

Yes, and they will probably be your highest-scoring prompts, which is exactly why they need care in interpretation. Comparison prompts naming a specific alternative were cited at 94% in our scan versus 0% for broad category prompts — but that gap partly reflects the fact that a comparison prompt hands the engine your brand name in the question. Treat them as a monitor for accuracy and framing (is the engine describing you correctly, and favourably, next to the alternative?) rather than as evidence of reach. Score them separately from unbranded prompts so a healthy branded number cannot mask a weak category number.

How often should I re-run a prompt set?

Monthly for the full set, weekly for the ten prompts closest to revenue. AI answers shift with model updates, index refreshes and your own publishing, and the volatility between individual runs means weekly full-set scans mostly buy you noise at real cost. The exception is any prompt attached to a page that converts: those deserve a tighter loop, because a citation gained or lost there shows up in revenue rather than in a dashboard. Re-run the whole set immediately after a major model release, since that is when historical baselines break.

Does getting cited by an AI engine actually mean the engine names my brand?

Often not, and the reverse is even more common. BuzzStream's July 2026 analysis of 12,000 AI responses across roughly 2,967 prompts and 200 brands found that only 23.1% of brand mentions came with a citation — in the other 76.9% the answer named the brand and gave the reader nothing to click — while 69.9% of citations did name the brand in the response text. So being cited and being named are two separate events that overlap only partially, and they break differently by platform and by query type. Any prompt-tracking setup that measures one and calls it visibility is reporting a fraction of your presence.

Sources

Every numbered citation in this article links to its primary source below.

  1. [1]New Research: AIs are highly inconsistent when recommending brands or products — SparkToro (2026).
  2. [2]How AI Mentions and Cites Your Brand (New Study) — BuzzStream (2026).
  3. [3]AI Recognizes 96% Of Brands But Mentions Almost None, New Study Finds — Search Engine Journal (2026).
  4. [4]Who AI Cites Now — Citation Source Audit Q1 2026 — 5W Public Relations (2026).
  5. [5]What Is Prompt Tracking? The 2026 Operator's Definition — Attrifast (2026).
  6. [6]We Re-Ran Our Own AI Visibility Scan After Finding a Bug In It: The Raw Data — Attrifast (2026).
  7. [7]Which Brands Does ChatGPT Recommend in 2026? — Attrifast (2026).
  8. [8]Overcoming AI Traffic Loss: The Measure-First Recovery Playbook — Attrifast (2026).
  9. [9]ChatGPT Referral Analytics: Why 70% of AI Traffic Hides in Direct — Attrifast (2026).
Reading this with an AI assistant?Ask Perplexity about this article →Read this article as markdown →

About the author

Vincent RuanFounder, Attrifast

Vincent Ruan is the founder of Attrifast, an analytics platform for website traffic, customer-level revenue and AI brand visibility. The prompt-level data in this post comes from scans run on his own domain and, in aggregate, from the platform itself — including the 84 consecutive runs on broad category prompts that returned zero citations for his own product, published unedited because a prompt library demonstrated only on the prompts that worked is a brochure.

  • X
  • vince-ruan.com
  • LinkedIn

Related reading

AI Search24 min
How to Measure GEO ROI: Proving AI Search Optimization Pays in 2026
Measure GEO ROI with a practitioner's method: baseline first, detect AI traffic by engine, join it to Stripe revenue, and compute (revenue minus cost) / cost.
AI Search14 min
Does GEO Actually Drive Revenue? An Honest Answer
GEO can drive revenue, but proving it requires architecture most teams don't have. The 4 evidence layers between AI citations and Stripe payouts.
AI Search16 min
Overcoming AI Traffic Loss: The Measure-First Recovery Playbook (2026)
When an AI summary appears, only 8% of Google visits click a result — down from 15%. Semrush's answer is brand marketing. Ours starts one step earlier: split AI traffic out of Direct, re-baseline the loss in revenue instead of sessions, then reallocate by what pays. Run live on our own dashboard: 292 visitors, 9.2% already AI-referred, and the engine that cites us most sending almost none of them.
Attribution19 min
Cross-Channel Marketing Attribution in 2026: Ten Channels, One Payment
Cross-channel marketing attribution when ChatGPT, Perplexity, Claude, and Gemini are real revenue channels: the double-count arithmetic, models, and setup.
AI Analytics14 min
What Is AI Visibility? The 2026 Definition, and How to Measure It
AI visibility is how often AI engines surface your domain across a defined prompt set. What it measures, how to score it, and why it means little on its own.

See which AI engines actually send you paying customers

Attrifast splits ChatGPT, Perplexity, Claude and Gemini into their own revenue lines — joined to real Stripe payments, not estimates.

  • ✓One script tag and a Stripe key — live in minutes
  • ✓Cookieless, so no consent banner for analytics
  • ✓Every AI referral matched to the payment it produced
Start your free trial →

7-day free trial · $0 due today · then $9.99/mo · cancel anytime

Attrifast dashboard: prompt-level AI visibility with estimated value, revenue split by channel across ChatGPT, Google, Perplexity, Claude and Direct, competitor position tracking, and per-engine scan settings.