An AI assistant names your brand for one reason: it appeared in the handful of sources, usually five to eight, that the engine retrieved before writing its answer. If your brand is absent from those retrieved documents, no amount of on-page optimization will surface it. And here is the part most teams get wrong: those retrieved sources are mostly not your own website. Muck Rack's analysis of more than one million links found that roughly 82% of AI citations come from earned media rather than owned properties. Ahrefs, studying around 75,000 brands, reported that brand web mentions correlate with AI visibility at r ≈ 0.664, while backlinks correlate at only r ≈ 0.218, a moderate correlation, so trust the ordering rather than the exact number, but the ordering says mentions outweigh backlinks roughly 3 to 1. The work of getting cited is largely the work of being talked about somewhere other than your domain.

The engine → index map

Every assistant runs on a retrieval layer. Knowing which one tells you where to be present.

Assistant Retrieval index
ChatGPT Browse Bing (Microsoft–OpenAI)
Copilot Bing
Gemini + Google AI Overviews Google Search
Claude web search Brave Search (per Anthropic subprocessor list, Mar 2025)
Perplexity Its own index, ~200 billion URLs; dropped the Bing API Aug 2025
Meta AI Bing

The coverage math follows directly. Ranking on Google plus Bing reaches roughly 70–75% of AI retrieval surfaces. The remainder, Claude on Brave, Perplexity on its own crawler, does not need ranking so much as clean, unblocked crawlability; a correct robots.txt is the price of entry. But ranking is not the whole game. Averi's study of 680 million citations found that only about 11% of citation domains overlap between ChatGPT and Perplexity. Breadth across sources beats winning position one on any single engine.

The channels that feed the models

Each channel below carries its own citation weight, and each favors particular engines.

Community and Q&A (Reddit, Quora, Stack Exchange)

Profound's 680-million-citation study found Reddit accounts for roughly 40% of all AI references, and 46.7% of Perplexity citations versus about 11% on ChatGPT. Community discussion dominates Perplexity and Google AI Overviews.

Encyclopedic (Wikipedia, Wikidata)

Wikipedia is roughly 26% of references and ChatGPT's single largest source at about 7.8% of its citations. Wikidata feeds Google's Knowledge Graph and LLM training and carries no notability bar, which makes it the highest-leverage realistic entity move for a small or mid-size business.

Video (YouTube)

YouTube is the single strongest observed correlate of AI visibility (ρ ≈ 0.74 in the Ahrefs 75k study) and about 23% of references. OtterlyAI's YouTube Citation Study (Mar 2026, 100M+ citation instances) found YouTube AI citations concentrate in Perplexity (38.7%) and Google AI Overviews (36.6%), that 94% go to long-form video, and that views barely matter (r ≈ -0.03), 40.8% of cited videos had under 1,000 views. Relevance, not reach, gets the citation.

B2B and trade directories and marketplaces (IndiaMART, TradeIndia, ExportersIndia, ThomasNet, Kompass)

These are the lever for Google AI Overviews and Perplexity, but only when the listing carries the product or identifier at product level, not just a company-category entry.

Review platforms (G2, Trustpilot, Capterra)

Industry analyses associate review-platform presence with roughly 3× higher ChatGPT citation odds, and the effect appears even at around five genuine reviews.

Professional (LinkedIn)

LinkedIn is the number-one B2B citation source, and personal profiles tend to outperform company pages.

The citation paths

A company gets named only if it lands in the retrieved five most-cited to eight. There are four ways to get there.

  • Path A, your own page ranks. Works only in thin fields where little else competes for the query.
  • Path B, a directory ranks and lists you at product level. The engine cites the directory; you are named because your product line, not just your company name, is present in that listing.
  • Path C, shipment and trade-intelligence records. Customs and trade-data aggregators surface you through documented transactions rather than marketing pages.
  • Path D, repeated pre-cutoff mentions. The model recalls you from training rather than retrieval. Rare, and not something you can reliably engineer.

Why on-site alone isn't enough

The citation source differs by engine, and that is the whole argument for working both sides. ChatGPT, sitting on Bing, tends to cite the brand's own website. Google AI Overviews and Perplexity are often directory-gated, they reach for third-party listings and Reddit threads before they reach for your homepage. Optimize only your site and you win ChatGPT while staying invisible on the two engines most likely to route through off-site sources. On-site and off-site are not competing strategies; they serve different engines, and you need both.

One honest nuance closes the loop. A directory you fill in yourself establishes company-level eligibility, it makes you retrievable. An independent third party talking about you is corroboration, and corroboration is what actually moves a model's answer. Self-made listings get you into the index; earned mentions get you into the sentence.

Map the engines to their indexes, then earn presence across the off-site channels that feed them, because the answer is written from the sources retrieved, not from the site you own.