How to Get Cited by AI: Showing Up in ChatGPT, Perplexity, Gemini and Claude
AI assistants recommend the sources they can parse and trust. How original data, clean structure, schema, and entity signals get your business quoted by name in ChatGPT, Perplexity, Gemini, and Claude.
Why getting cited is the new ranking
For twenty years the goal of search work was a position. You wanted the blue link at the top of the page, and everything you did rolled up to that one number. That goal has not disappeared, but it has been joined by a stranger and more valuable one: being the source an AI assistant names when it answers out loud.
The shift, stated directly. When a buyer asks ChatGPT for the best ergonomic chair, or asks Perplexity which invoicing tool fits a two-person studio, the model does not hand back ten links and walk away. It writes an answer. Inside that answer it mentions specific brands, quotes specific numbers, and, increasingly, drops a citation footnote pointing at the page it pulled from. If your business is the one named, you have won a placement that no amount of paid search can buy back later. If a competitor is named instead, you are invisible in the exact moment the buyer was deciding.
I want to be precise about the word cited, because people use it loosely. There are three distinct things happening and they matter for different reasons:
- Mentioned: the model names your brand in prose, with no link. This still drives recognition and downstream branded search.
- Cited: the model attaches a visible source link to your page, the way Perplexity and AI Overviews do. This drives referral traffic you can measure.
- Recommended: the model puts you forward as the answer to a decision-stage question. This is the one that converts.
All three come from the same underlying cause, so I am going to treat them together for most of this piece and separate them only where the tactic differs. The cause is simple to state and hard to fake: models surface sources they can read cleanly, verify against other sources, and quote without embarrassing themselves. Clarity, specificity, and credibility are what make a page trustworthy to a human reader. They are also, not by coincidence, exactly what make a page citable to a model.
The reason this is worth your attention right now is timing. Most brands have not adjusted. They are still writing vague, image-heavy marketing copy that a model cannot extract a single quotable fact from, and they are still blocking the crawlers that would have read them. That gap is your opening. The work is not exotic. It is mostly disciplined structure applied to genuinely useful information, and the sites that do it early are compounding an advantage while everyone else argues about the reality of AI search.
There is a second reason the shift matters more than a channel change, and it is about intent. The queries that migrate to assistants first are the high-consideration ones: the research, comparison, and decision-stage questions where a person is genuinely weighing options. Casual navigational searches still go to the box, but when someone is trying to actually choose, they increasingly ask a model to reason with them. That means the traffic you win on this surface is not idle; it is people at the moment of decision. Losing it is not losing impressions, it is losing the deals.
I also want to kill a comforting myth early. Plenty of marketers assume that if they rank well on classic Google, they are automatically safe here, and it is not true. I have watched sites that own the number-one blue link go completely unnamed by ChatGPT and Perplexity for the same query, because their page, strong as it is for human skimming, offers a model no clean unit to lift and no signal it can trust. Ranking and getting cited are correlated, not the same. The overlap is real, but the gap is where this article lives, and closing it is deliberate work rather than a byproduct of good SEO.
How large language models choose what to cite
To earn citations on purpose, you have to understand the mechanism you are optimizing for, and most advice skips this. A model deciding what to cite is running two very different processes depending on how it was asked, and the tactics change with each.
When a model answers from memory alone, with no live web access, it is drawing on patterns baked in during training. It learned those patterns by reading an enormous slice of the public web months or years before your question. Your brand is present in that answer only if your content, and other people's content about you, was in the training data and was distinctive enough to leave a trace. This is the slow loop. You cannot influence it this week, but it compounds, and it is why long-lived, widely-referenced pages matter so much.
When a model answers with live retrieval, which is now the default in Perplexity, in AI Overviews, in ChatGPT with search on, and in Gemini and Claude when they browse, something closer to real-time search happens. The system runs a query, pulls a handful of candidate pages, reads them, and synthesizes an answer with citations attached to the pages it actually used. This is the fast loop. A page you publish or fix today can be cited within days because it competes in the live retrieval set, not in frozen training data.
Inside the fast loop, the selection funnel looks roughly like this, and every stage is a filter you can pass or fail:
- Retrieval: does your page even come back for the query? This depends on classic relevance and on your page being crawlable and indexed by whatever search backend the assistant uses.
- Parsing: can the model cleanly extract text, headings, tables, and facts? JavaScript-only pages and image-trapped text fail here silently.
- Trust: does the source look credible? Author, dates, citations, and corroboration from other sites all feed this.
- Quotability: is there a clean, self-contained claim the model can lift without misrepresenting you? A buried, hedged point rarely survives.
- Corroboration: do independent sources agree? Models prefer facts they can see echoed, which is why off-site mentions matter as much as your own page.
A page that clears all five gets cited. A page that fails any one of them drops out, and the failure is usually invisible to you because nobody tells you why you were not chosen. Most of this article is really about clearing those five gates deliberately instead of by accident.
One more mechanic is worth understanding, because it explains a lot of otherwise-confusing behavior: models synthesize, they do not rank. A classic search engine hands you a list and lets you pick. A generative model reads across a handful of sources and writes one blended answer, then decides which of those sources to name. That blending is why a single source rarely dominates and why corroboration matters so much. The model is more comfortable stating a fact it saw in three places and will happily cite all three, while a lone claim on one page, however true, is riskier for it to repeat and easier to leave out. You are not competing for a slot; you are competing to be one of the few sources the model felt safe leaning on.
It also means representation matters as much as inclusion. Being named is good, but being named accurately, with your positioning and numbers intact, is the real goal, and it depends on how cleanly you stated things. Sloppy or ambiguous claims get paraphrased into something you did not mean. Precise, self-contained claims get repeated close to verbatim. The cleaner your writing, the more control you keep over how the model describes you, which is a kind of control most people never think to ask for.
The two loops: training data and live retrieval
I keep the two loops separate in my own head because they reward different work on different timelines, and conflating them leads to wasted effort. Let me put them side by side.
The training loop is slow, compounding, and mostly out of your direct control. Foundation models are retrained or refreshed periodically, and each cycle ingests a new snapshot of the crawlable web. To be in it, your content needs to have existed, been reachable, and been distinctive enough that the model retained something about you. You do not get a report telling you it worked. You find out months later when a browsing-disabled model already knows who you are. The levers here are patience and durability: publish flagship pages that stay up for years, earn references from other sites so your name appears in many contexts, and avoid the churn of deleting and reslugging URLs that resets your history.
The retrieval loop is fast and testable. The model runs a query against a live index the moment someone asks, so a page you ship today can be cited this week. The levers here are the ones you can pull immediately: be indexable, be parseable, answer the question directly near the top, and be freshly dated. This is where I tell clients to spend the first ninety days, because you get feedback fast and the wins fund the slower work.
The important thing is that the two loops reinforce each other. A page built to win live retrieval, structured, sourced, dated, is exactly the kind of page that survives into the next training snapshot and teaches the model your brand. You are never choosing between them. You build for retrieval because it pays now, and you get training inclusion as a compounding bonus for having done it consistently.
One caution about the training loop that trips people up. Because inclusion is frozen at a snapshot, a model with browsing off can confidently tell someone your old pricing, your former positioning, or a product you discontinued. You cannot edit the model's memory. Your only defense is to make the live, correct version so easy to retrieve that any browsing-enabled answer overrides the stale one, and to keep your canonical facts consistent across the web so the next snapshot learns the current truth.
There is a strategic way to think about the split that I find clarifies budgets. Treat live retrieval as your performance channel and training inclusion as your brand channel. The retrieval work is measurable, fast, and attributable, so it earns its budget the way paid search does: you can point at referrals and citations within weeks. The training work behaves like brand advertising, diffuse, slow, and hard to attribute to a single action, but responsible for the ambient recognition that makes every other channel cheaper. Nobody funds only one of those in a mature marketing program, and the same logic applies here. Fund the retrieval work because it pays the bills, and fund the durability work because it is what makes a model already know you before anyone clicks.
A concrete example of the two loops disagreeing: I have seen a rebranded company where browsing-off ChatGPT still used the old name months after launch, while Perplexity, retrieving live, used the new one correctly. Both were reading the same company; they were just reading different snapshots of it. The fix was not to argue with the model. It was to flood the live web with the new name consistently, update every profile and citation, and wait for the next training cycle to catch up, while relying on retrieval to carry the correct answer in the meantime. That is the two-loop model in action, and once you see it you stop being surprised by contradictory answers.
The five signals that earn citations
Strip away the surface tactics and citation-worthiness comes down to five signals. I have watched these hold across industries, and when a page is not getting cited, it is almost always missing one or two of them rather than failing at all five.
**Structure and schema.** JSON-LD tells a model exactly what your content is, who wrote it, and when it changed, without the model having to infer any of it from layout. Article, Organization, Person, Product, and FAQPage schema turn ambiguous prose into parseable facts. Structure also means the visible thing: real headings, short paragraphs, tables for anything comparative, and a question phrased as a heading with the answer right beneath it.
**Direct answers.** Lead the relevant part of every page with a self-contained 40 to 60 word response to a real question. Models lift clean, standalone answers because an extracted snippet has to make sense with no surrounding context. A correct point buried in paragraph nine, wrapped in qualifiers, almost never gets quoted. The answer-first habit is the single highest-impact change most sites can make.
**Original information.** Specific numbers, named frameworks, and points of view that exist nowhere else. A model has no reason to cite the tenth page that repeats the same generic advice; it cites the one page with a proprietary statistic or a defined method. If you run a survey of 500 customers and publish the result, you become the primary source everyone else has to point back to, and primary sources get cited by definition.
**Consistent identity.** The same brand name, spelling, description, and sameAs links across your site, your social profiles, and the knowledge graph. Identity is how a model decides two mentions are the same trustworthy entity rather than two unrelated strings. Inconsistent naming fractures your authority into pieces the model never reassembles.
**Freshness and trust.** Visible dateModified stamps, a named author with a real bio, citations to credible sources, and HTTPS. Models prefer current sources and confident, sourced claims over anonymous, stale ones. Freshness is more than a date in your schema; it is a genuinely updated page, because a model that browses can tell the difference.
Think of these as five gates rather than a menu. You do not get to pick two. A page with brilliant original data that a crawler cannot parse fails as surely as a perfectly structured page that says nothing new.
If I had to rank them by effort-to-impact for a site starting from scratch, I would put structure and direct answers first, because they are cheap, fast, and clear two gates at once. Original information comes next; it costs more to produce but it is the difference between being one interchangeable source and being the source. Consistent identity is low effort and high value, mostly a matter of discipline rather than creativity, and yet it is the one teams most often neglect because it is nobody's exciting project. Freshness and trust are ongoing hygiene rather than a one-time task, which is exactly why they get dropped: they require a habit, not a launch.
What makes the five-signal frame useful in practice is diagnosis. When a page is not getting cited, you rarely need a rewrite. You need to figure out which specific gate it is failing and fix that one thing. Is the answer buried, or missing entirely? Is the crawler blocked? Is there anything original on the page at all, or is it a paraphrase of the same advice on ten other sites? Is the author anonymous? Nine times out of ten the silence traces to one or two of these, and naming the gate turns a vague I am not getting cited into a concrete, cheap task. That is the entire value of a checklist: it converts anxiety into a to-do list.
The content patterns models reuse
Some formats get quoted far more than others, and the pattern is consistent enough that you can write to it deliberately. Models reuse content that is easy to extract as a clean, verifiable unit. Prose that meanders gives them nothing to grab; structured claims hand them a ready-made quote.
| Pattern | Query it wins | Why models reuse it |
|---|---|---|
| Direct answer (40-60 words) | Any question query | Self-contained, needs no context |
| One-sentence definition | What-is queries | Clean source for a term |
| Original statistic | Data and proof queries | Primary source others cite |
| Comparison table | X vs Y, best-for queries | Already structured, parses cleanly |
| Numbered steps | How-to queries | Maps directly to HowTo schema |
The direct answer is the workhorse. Put a heading that matches the question the way people actually ask it, then answer in the first two sentences, then expand. A definition is a close cousin and it is wildly underused: a crisp one-sentence definition of a term you own is one of the most-cited artifacts on the web, because assistants constantly answer what-is questions and need a clean source for the definition.
Statistics get pulled constantly, and specificity is the whole game. Nine percent average savings gets cited; significant savings does not. A number with a clear subject, a unit, and ideally a date and a source is a magnet. When the number is yours and original, you become the citation everyone else has to use, which is the strongest position there is.
Tables are quietly the highest-value format for comparative and sequential content. Answer engines and generative models both parse tables reliably and reproduce them readily, because a table already is structured data. Any time your content is X versus Y, best X for each use case, or spec-by-spec, a table will out-cite the same information written as paragraphs. Here is a quick map of the patterns and what surface they win:
- Direct answer under a question heading, 40 to 60 words: wins featured snippets, AI Overviews, and every assistant.
- One-sentence definition of a term: wins what-is queries across ChatGPT, Gemini, and Claude.
- Original statistic with subject, unit, and source: wins citations and becomes a reference other sites reuse.
- Comparison or spec table: wins X-versus-Y and best-for-use-case queries in Perplexity and AI Overviews.
- Numbered steps for a process: wins how-to queries and HowTo-schema rich results.
- Named framework or method: wins concept queries and gets your term repeated back to users.
The meta-point is that you are not writing for a reader who reads top to bottom. You are writing for an extractor that scans for the cleanest self-contained unit that answers the query. Give it obvious units, label them with headings and schema, and you make its job trivial, which is precisely when it rewards you.
There is a craft to writing the direct answer that is worth spelling out, because most people write it wrong. The answer has to survive being ripped out of the page and dropped into a model's response with zero surrounding context. That means no pronouns that reference an earlier sentence, no as we mentioned above, no this depends on your setup. State the subject, the claim, and the key qualifier in one clean unit. Read it aloud as if it were the only thing a stranger would ever see from your page, because that is often literally true. If it still makes sense standing completely alone, it is extractable. If it needs the paragraph around it, rewrite it.
The same self-containment rule is why definitions are so powerful and so underused. Every assistant fields a constant stream of what-is questions, and it needs a crisp, authoritative sentence to answer them. If you coin or clarify a term in your space and define it in one clean sentence, you become the default source for that definition, and the model repeats your framing back to thousands of people. Naming and defining your own concepts is one of the highest-return moves in this entire discipline, because a good definition is compact, distinctive, and quotable, which is exactly the profile of content that travels furthest into both the retrieval set and the training data.
Getting cited in ChatGPT, with and without browsing
ChatGPT is really two systems wearing one interface, and you optimize for each differently. With search off, it answers from training memory. With search on, which is now common, it runs live queries, reads pages, and attaches citations to the ones it used. Knowing which mode you are targeting tells you which loop to pull.
For the browsing mode, the fast loop rules apply and they pay quickly. ChatGPT's search leans on a live web index, so your page has to be indexable and it has to be readable when fetched. That means server-rendered HTML, not content that only appears after JavaScript runs, because the fetch-and-parse step frequently sees the pre-render state. It means an answer near the top of the page that resolves the query in a couple of sentences. And it means the page has to actually be about the specific question, not a broad category page that mentions it in passing. When ChatGPT browses, it tends to cite three to six sources; your job is to be the cleanest, most on-point one in that shortlist.
For the memory mode, you are playing the training loop, and the work is durability and distinctiveness. ChatGPT knows brands it saw referenced widely and consistently across the web. Original data, a defined method with a name, and mentions on sites other than your own are what leave a durable trace. You will not move this in a week, but a page that has been up for a year, is referenced by other sites, and states something specific and quotable is how you end up in the answer when browsing is off.
A few things I have seen matter more than people expect with ChatGPT specifically:
- Match the conversational query, not the keyword. People ask ChatGPT full questions in natural language. A heading phrased as best crm for a solo consultant beats one phrased crm software solutions.
- Keep the quotable claim early and unhedged. ChatGPT's synthesis favors confident, self-contained statements. It it, it depends and many argue phrasing gets skipped in favor of a source that just says the thing.
- Do not block OAI-SearchBot or GPTBot if you want to be read. I have watched sites wonder why they are never cited while their robots file quietly disallows the exact crawler doing the reading.
- Corroboration helps you survive. When two independent sources say the same number, ChatGPT is far more comfortable repeating it, so earning a second mention elsewhere protects your citation.
The honest summary is that ChatGPT rewards the same fundamentals as everything else, with an extra premium on natural-language question matching and on not accidentally blocking the crawler that would have read you.
Worth knowing: ChatGPT uses distinct crawlers for distinct jobs, and conflating them causes real mistakes. GPTBot gathers content that can inform training. OAI-SearchBot and the ChatGPT-User agent handle the live browsing and retrieval that produces cited answers in the moment. If you block GPTBot for training-rights reasons but leave the search agents allowed, you can still be cited in browsing mode while staying out of training, which is a perfectly reasonable posture for some businesses. The failure I see is a blanket block that catches all of them, killing your live citations as collateral damage. Decide crawler by crawler, on purpose.
The other thing that separates cited-in-ChatGPT pages from ignored ones is scope discipline. ChatGPT's browsing tends to reward the page that is precisely about the question over the sprawling pillar page that covers the topic broadly. If someone asks about a narrow sub-question, a focused 1,200-word page that answers exactly that, with a direct answer and a table, frequently beats a 6,000-word mega-guide where the same answer is paragraph forty. This is genuinely good news for smaller sites: you do not need the biggest page, you need the most on-point one. Break broad topics into focused pages that each own one question cleanly, and you give ChatGPT many precise targets instead of one diffuse one.
Getting cited in Perplexity
Perplexity is the purest citation engine of the four, and that makes it the best place to learn what works, because it shows its sources openly on every answer. Every response is a synthesized paragraph or two with numbered footnotes pointing at the pages it used. If you want to see your citation strategy working or failing in real time, Perplexity is your test bench.
Because Perplexity retrieves live for essentially every query, the fast loop dominates and results come quickly. It runs a search, pulls a set of candidate pages, and reads them, then writes an answer that quotes and links the ones it found most useful. Three things decide if you make the cited set. First, you have to rank in the underlying retrieval, so classic relevance and a crawlable, indexed page are the price of entry. Second, your page has to contain a clean, liftable answer to the specific question, because Perplexity is aggressive about extracting concise spans. Third, structure wins ties: tables, clear headings, and a direct answer make you the easiest source to quote, and the easiest source usually gets the footnote.
Perplexity has a strong appetite for comparison and best-of queries, which is where a lot of commercial intent lives. Questions like best project management tool for agencies or X versus Y pricing pull tables and spec lists directly. If you publish an honest, specific comparison with a real table, you are handing Perplexity exactly the structure it wants to reproduce. Vague we are the best framing gets ignored; a table that names competitors and states tradeoffs gets cited, even when it is on your own site, because it reads as useful rather than promotional.
A practical routine I use: pick the ten questions your buyers actually ask, run each through Perplexity, and record who gets cited and how they are described. The sources it names are your real competition for that surface, and the gaps, questions where the cited pages are weak, are your fastest wins. Do not allow PerplexityBot to be blocked in robots.txt, publish a tight answer plus a table for each of those ten questions, and re-check in a few weeks. Perplexity's speed of re-crawl means you often see movement faster here than anywhere else, which makes it the ideal proving ground before you scale the same pages to the slower surfaces.
One behavior specific to Perplexity is worth designing for: it frequently issues several sub-queries behind a single question and stitches the results together. Ask it for the best invoicing tool for freelancers and under the hood it may separately look up pricing, features, and reviews, then assemble them. That means a page that cleanly answers one facet, just the pricing comparison, or just the freelancer-specific features, can get pulled into an answer even if it does not cover the whole topic. You do not have to own the entire question to be cited; you have to own one facet of it cleanly. This rewards a portfolio of focused pages over a single sprawling one, the same lesson that shows up on every surface.
Because Perplexity is so transparent, it is also the fastest way to audit how you are being described, not just if you appear. Run your brand and your key questions, and read the sentence it writes about you. If the description is stale, wrong, or unflattering, that tells you which source it trusted and what you need to correct or out-publish. I treat that sentence as a live reputation readout. Fixing the underlying page or earning a better third-party mention changes what Perplexity says within a crawl cycle or two, which is a tighter feedback loop than any classic reputation channel offers.
Getting cited in Gemini and Google AI Overviews
Gemini and AI Overviews sit on top of Google's index, which changes the calculus in a useful way: the groundwork you already do for classic Google search is most of the groundwork for both. If you rank well and your pages are structured, you are already in the candidate pool these systems draw from. Around a third of Google queries now trigger an AI Overview, so this is not a fringe surface; it is the default experience for a large and growing share of searches.
Because the retrieval backbone is Google's, the entry ticket is classic: be indexed, be relevant, have healthy technical SEO, and earn the authority signals Google already rewards. On top of that baseline, AI Overviews shows a strong preference for a few things. It pulls direct answers that sit high on the page and match the query intent. It reproduces tables and structured lists readily. And it favors pages with clear FAQPage and Article schema, because the markup removes ambiguity about what each block is. Speakable schema on your answer block composes here too, since the same clean answer feeds voice results.
The control you use with Google is Google-Extended. It is the token that governs if your content can be used to ground Gemini and AI Overview responses, and it is separate from Googlebot indexing you for regular search. This is the subtle trap I see most: a site can rank perfectly in classic results while quietly opting out of the generative surfaces because someone set Google-Extended to disallow, or because a blanket AI-crawler block swept it up. If you want to be cited in AI Overviews and Gemini, confirm Google-Extended is allowed and that your pages render server-side so the extraction step sees real content.
One more nuance worth stating plainly. AI Overviews frequently cites pages that are not the number-one classic result, because it optimizes for the cleanest answer to the specific sub-question, not the strongest overall page. That is good news. A focused page that nails one question with a direct answer and a table can get cited above a broad, higher-authority page that buries the same information. Depth on a specific question beats breadth here, which is exactly the kind of page a smaller site can win with.
It is worth being clear-eyed about the tension AI Overviews creates, because it is real. The Overview answers the question at the top of the page, and that can reduce the clicks you would have gotten from the classic blue link below it. Being cited in the Overview is still worth more than being ignored, because the citation carries your brand into the answer and a meaningful share of users click through for depth, but you should measure the net effect honestly rather than assume every citation is pure upside. The pages that keep earning clicks from Overviews are the ones that promise more than the summary can contain: the full table, the calculator, the original dataset, the detailed how-to. Give the Overview enough to cite you and a clear reason to send the reader onward for the rest.
A practical note on the Google-Extended control, since it confuses people: it is decoupled from indexing on purpose. Googlebot can index you for classic search while Google-Extended governs, separately, if that content grounds Gemini and AI Overviews. So you can rank normally and still be absent from the generative layer if Google-Extended is disallowed. Check it explicitly. I have found more than one site that opted out inadvertently through an over-broad AI-crawler policy and never realized the generative traffic it was forfeiting, because classic rankings looked perfectly healthy the whole time.
Getting cited by Claude
Claude answers from training memory by default and browses the web when the task calls for it or when a connected tool provides live access. In its browsing and tool-connected modes, and increasingly inside agentic workflows where Claude is reading pages to complete a task, the same fundamentals decide if your content gets used and represented accurately.
When Claude reads your page, through its own browsing or through an MCP connection or a retrieval tool an app has wired up, it wants clean, well-structured HTML it can parse without a headless browser executing scripts. It handles long, nuanced content well, which rewards depth, but it still favors a clear answer it can lift and attribute. The premium Claude puts on not misrepresenting a source works in your favor if you write carefully: an unambiguous, self-contained claim is safer for the model to repeat than a hedged or context-dependent one, so precise writing gets quoted more.
For the training-memory mode, the levers are the familiar slow-loop ones: be present across the web, be distinctive, and be consistent about your identity so the model retains a coherent picture of who you are. Original frameworks and named methods travel especially well into a model's memory because they are compact and distinctive, which is part of why defining and naming your own concepts pays off across every assistant, Claude included.
There is a forward-looking reason to care about Claude specifically, and it is agents. As assistants move from answering questions to completing tasks, they read your pages not just to quote you but to act: to compare options, fill a cart, or shortlist vendors. A page that is machine-readable, states its facts plainly, and exposes structured data is one an agent can use confidently. This is where llms.txt and PotentialAction schema start to matter, because you are no longer only trying to be quoted, you are trying to be usable by software acting on a person's behalf. The sites that structure for that now will be the ones agents can actually transact with later, and that is a much bigger prize than a footnote.
Claude's tolerance for depth changes how you should write for it specifically. Where some surfaces reward a tight snippet and nothing more, Claude will happily read and reason over a long, nuanced page, follow a careful argument, and represent tradeoffs faithfully rather than flattening them into a one-liner. That means your considered, thorough content is an asset here rather than something to trim. The move is to give Claude both: a clean, liftable answer near the top for when it needs a quick quote, and the genuine depth beneath it for when it is reasoning through a complex task. You are not choosing between snappy and thorough; you are layering them, headline answer first, full reasoning below.
The agent angle deserves one more concrete beat, because it reframes the whole exercise. When Claude is acting as an agent, reading pages to complete a job rather than to answer a chat, it behaves less like a reader and more like a program consuming an API that happens to be written in HTML. It wants unambiguous structure, stable identifiers, clearly stated prices and availability, and, ideally, declared actions it can take. A page built only to look good to humans forces the agent to guess; a page that also exposes clean structured data lets it act with confidence. Building for that agent reader now is the same work as building for citation today, which is the happy part: you do not need a separate strategy for the agent era, you need to do the citation work well and expose your actions, and you are already most of the way there.
Schema and structured data that make you quotable
Schema is the highest-return technical work for citations because it removes guesswork. Instead of hoping a model infers that a block is your author bio or your price or your answer, JSON-LD states it as a fact the parser reads directly. It is invisible to human visitors and decisive for machines, which is exactly the trade you want.
| Schema type | What it tells the model | Where to use it |
|---|---|---|
| Article | What this is, when it changed, who wrote it | Every editorial page |
| Person | Author is a real, resolvable entity | Author byline on articles |
| Organization | Your brand identity and profiles | Every page, sitewide |
| FAQPage | Explicit question-answer pairs | Pages with Q&A blocks |
| Product | Specs, price, rating | Commercial and comparison pages |
| Speakable | Which block to read aloud | TL;DR and FAQ blocks |
Start with the types that carry the most citation weight and layer them on every relevant page:
- Article with headline, datePublished, dateModified, and a full author object. This establishes what the page is, when it changed, and who stands behind it.
- Person on the author, with name, jobTitle, description, and sameAs links to their real profiles. This is how a model resolves your author to a known entity rather than an anonymous string.
- Organization on your brand, with name, url, logo, and sameAs to your verified profiles and, ideally, a Wikidata entry. This anchors your identity across the web.
- FAQPage on your question-and-answer blocks, with each question phrased the way people ask and each answer under fifty words. This maps your content to the exact shape answer engines extract.
- Product with name, description, offers, and aggregateRating on commercial pages, so comparison and best-of answers can pull your specifics.
- HowTo on genuine step-by-step processes, and Speakable on your direct-answer and FAQ blocks so voice surfaces read them aloud.
The rule that keeps schema working is that it must describe what is actually on the page. Marking up an FAQ that does not visibly exist, or a rating you do not display, is the fast way to get your rich results suppressed and your trust score dinged. Schema is a mirror of the page, not a substitute for content. Validate everything in Google's Rich Results Test and Schema.org's validator before you ship, because a single malformed block can invalidate the rest.
The part people forget is dateModified, and it matters more than its size suggests. Models and answer engines both prefer fresh sources, and a dateModified that is genuinely current, backed by a real update to the page, is one of the cheapest freshness signals you can send. The trap is stamping a new date on a page you did not actually update; a browsing model can compare the claimed date to the content and is not fooled for long. Keep the stamp honest and keep the page genuinely current, and you get the freshness credit for real.
One subtle schema point separates people who get value from it and people who cheat themselves out of it: the sameAs array on your Person and Organization objects is doing identity resolution, not decoration. Each sameAs link is you telling the model this profile and that profile are the same entity as this page. When you link your Organization to your verified LinkedIn, your Crunchbase, your X account, and your Wikidata QID, you are handing the model a resolved identity graph instead of a pile of loose strings it has to guess are related. Skip sameAs and your author is just a name; include it and your author becomes a known entity the model can attribute to with confidence. It is a few lines of JSON doing some of the heaviest lifting on the page.
A word on maintenance, because schema rots quietly. Prices change, authors leave, products get discontinued, and stale JSON-LD that contradicts the visible page is worse than no schema at all, since it actively teaches the model something false and can get your rich results penalized. Treat schema as living data tied to the page, not a set-and-forget block you paste once. The teams that win with structured data are the ones who update the markup in the same motion as the content, so the two never drift apart. If you cannot commit to keeping a piece of schema accurate, it is better not to add it than to let it lie on your behalf.
Ship an llms.txt and an ai.txt file
Two small files at your site root do outsized work for AI discovery, and almost nobody has them yet, which is precisely why they are worth the twenty minutes. They are the robots.txt of the AI era, and being early is a real advantage.
The first is ai.txt, which is your training-crawler policy. It states, per crawler, if the big AI companies may use your content. If you want to be in the training data that powers the memory-mode answers, you opt the training crawlers in explicitly. The relevant agents to name today are GPTBot for OpenAI, ClaudeBot for Anthropic, Google-Extended for Gemini and AI Overview grounding, PerplexityBot for Perplexity, and CCBot for Common Crawl, which feeds many models indirectly. Blocking these is a legitimate choice for some businesses, but do it on purpose, not by accident, because a blanket disallow is how sites make themselves invisible without realizing it.
The second is llms.txt, a plain-text, machine-readable description of your business written for a model. It reads like a clean taxonomy: a one or two sentence statement of what you do, your primary URL patterns, your key pages, and, if you support them, the actions an agent can take. It gives a model a reliable, low-noise summary instead of forcing it to reconstruct your business from marketing pages. A useful llms.txt has a few sections:
- What we do: one or two sentences, written plainly for a model, no slogans.
- Primary URL patterns: the shape of your important sections, like /blog/{slug} or /products/{category}.
- Key pages: direct links to your flagship, most-citable content.
- Available actions: any PotentialAction endpoints an agent can call, described simply.
- Canonical reference: your exact brand name, legal entity, domain, and Wikidata QID so identity resolves cleanly.
I want to be honest about what these files do and do not do. They are not a magic ranking lever, and no crawler is strictly obligated to honor them. What they do is remove ambiguity and signal intent. When a crawler or agent is deciding what your site is and if it may use it, a clean llms.txt and a clear ai.txt make you legible and cooperative, and legibility is a real edge when most sites offer neither. Pair them with the schema work above and you have covered both the machine-readable-page layer and the machine-readable-site layer, which is the full surface a model reasons over.
A good way to think about llms.txt is as the executive summary you would hand a new analyst on their first day. If a smart stranger had to understand your business in ninety seconds and then answer questions about it accurately, what would you tell them? That is the content: what you do, in plain language; where the important things live, as URL patterns; which pages are the flagship references; and what a person or agent can actually do on your site. Written that way, it doubles as a forcing function. If you cannot describe your business clearly and briefly for a model, that is usually a sign your site does not describe it clearly for humans either, and fixing the file often surfaces positioning problems worth fixing everywhere.
I will temper the enthusiasm with realism, because hype around these files runs ahead of the evidence. Adoption of llms.txt by the major model providers is still uneven, and none of these files is a guaranteed instruction that every crawler obeys. So do not treat them as a substitute for the real work of good pages, clean schema, and genuine authority. Treat them as cheap insurance and a clarity exercise. They cost almost nothing to ship, they can only help your legibility, and if the standards solidify the way robots.txt eventually did, the sites that already had a thoughtful one will be ahead. Low cost, plausible upside, and a useful side effect of forcing you to state plainly what you are: that is the honest case for shipping both today.
The step-by-step playbook
Enough principles. Here is the order I actually run this in, because sequence matters. You want the fast, cheap, high-impact work first so you get cited quickly, then the slower authority work that compounds. This is roughly a ninety-day arc for a site that starts with reasonable content.
**Week one, make yourself readable.** Confirm your key pages render server-side, not JavaScript-only. Check robots.txt and ai.txt and make sure GPTBot, ClaudeBot, Google-Extended, PerplexityBot, and CCBot are allowed if you want to be read. Fix anything that traps text inside images. This is the parsing gate, and failing it makes everything downstream pointless.
**Weeks two and three, add direct answers and schema.** Take your top twenty pages by intent and add a 40 to 60 word direct answer near the top of each, under a heading phrased the way people ask. Layer Article, Person, Organization, and FAQPage schema. Add honest dateModified stamps. Validate everything. This is the work that gets you cited fastest, because it clears the quotability and structure gates at once.
**Weeks four and five, build the citable assets.** Publish two or three flagship pieces with something original: a small survey, a proprietary statistic, a named framework, a genuinely useful comparison table. These are your citation magnets, the pages other sites and models point back to because you are the primary source.
**Weeks six and seven, ship the site-level files.** Write llms.txt with your taxonomy and key pages, finalize ai.txt, and add sameAs links everywhere. Create a Wikidata entity if you can support it. This is the identity and legibility layer.
**Weeks eight through twelve, build corroboration.** Pursue earned mentions in the roundups and press your buyers read, seed your original data where others will cite it, and keep your identity consistent across every placement. This is the slow loop, and it is what carries you into the next training snapshot.
Run a simple loop the whole time: pick the ten questions your buyers ask, test them weekly in Perplexity and ChatGPT, and record if you are named. Treat every gap as your next content task. The playbook is not complicated. It is disciplined sequence applied to real information, and the discipline is the part most sites skip.
Why this exact order, and not some other? Because it front-loads the work that pays fastest and de-risks the rest. Readability and direct answers cost little and clear the two gates that block everything else, so doing them first means every later investment actually gets seen. Original assets come before the site-level files because there is no point advertising your taxonomy in llms.txt if the pages it points at have nothing citable on them. And corroboration comes last not because it matters least, it may matter most, but because it is slow and outside your direct control, so you want the fast internal wins banked and generating momentum while the slow external work grinds forward. Doing it in reverse, chasing press before your pages are readable, is how teams burn a quarter with nothing to show.
A note on scoping the effort realistically. You do not do this to your entire site at once, and trying to is how the whole program stalls. Pick the twenty to thirty pages that map to real buyer intent, the questions with money behind them, and run the full playbook on just those first. A tight set of pages done completely beats your whole catalog done halfway, both because focus produces better pages and because it gives you a clean read on what worked. Once that first cohort is getting cited, you have a proven template and the confidence to roll it out wider. Start narrow, prove it, then scale. The teams that try to boil the ocean on day one are the ones who conclude, wrongly, that none of this works.
The mistakes that keep you uncited
When a site is not getting cited, the cause is usually a short list of avoidable mistakes rather than anything mysterious. I see the same ones over and over, and most are cheap to fix once you know to look.
The biggest is putting the information a model needs inside things it cannot read: images, video, JavaScript-rendered content, and PDFs that are really scans. If your key facts only exist as pixels or only appear after scripts run, a fetch-and-parse crawler sees an empty page. You wrote content the model never received.
The second is blocking the crawlers you want to be read by. A blanket AI disallow in robots.txt or ai.txt, or a Google-Extended opt-out, quietly removes you from the exact surfaces you are trying to win. I have audited sites baffled at their invisibility whose own config was the reason.
The third is sameness. Publishing the same generic advice as everyone else gives no model a reason to pick you over the nine other pages that say it. Without an original number, a named method, or a first-hand detail, you are interchangeable, and interchangeable sources do not get cited.
The rest are quieter but just as costly:
- Burying the answer under 800 words of throat-clearing, so the extractable span is nowhere near the top.
- Hedging everything with it depends and many argue, so there is no confident claim to quote.
- Anonymous team bylines and no author schema, so the model cannot attribute the source to anyone credible.
- Stale content with an old date, or a fake-fresh date on a page you did not actually update.
- Inconsistent brand naming across your site and profiles, so your authority never consolidates into one entity.
- Marking up schema that does not match the visible page, which gets your rich results suppressed and dents trust.
- Paywalling or gating the very content you want cited, since a model cannot quote what it cannot reach.
None of these are exotic. They are the difference between a page that does everything right except one thing and a page that gets cited. Run the checklist against your top pages and you will usually find two or three of these hiding in plain sight.
There is a deeper, more strategic mistake that sits underneath the tactical list, and it is worth naming because it is the one that wastes the most effort: optimizing for the machine before you have anything worth citing. Some teams get so absorbed in schema, crawler config, and llms.txt that they never stop to ask if their page actually says anything a model would want to repeat. All that structure is plumbing. It carries value from your page to the model efficiently, but it cannot manufacture value that is not there. A perfectly marked-up page of generic advice is a well-plumbed empty house. The signals get you read; the substance gets you cited. If you only have budget for one, spend it on having something original to say, and structure it as well as you can.
The other trap worth flagging is chasing every surface equally instead of the ones that matter for your buyers. It is easy to read a piece like this and feel you must simultaneously conquer ChatGPT, Perplexity, Gemini, and Claude this quarter. You do not. The fundamentals overlap so heavily that doing the core work well, readable pages, direct answers, original data, clean schema, consistent identity, makes you eligible everywhere at once. Pick the one surface where your buyers actually are, prove the approach there, and let the shared fundamentals carry you onto the rest. Spreading thin across all four before you have won one is a good way to do everything at 60 percent and get cited nowhere.
How to track when you are being cited
You cannot improve what you do not measure, and AI citation is measurable if you set up a simple, repeatable process. It is less automated than classic rank tracking today, but the signal is clear enough to run a program on.
| Signal | How to capture it | What it tells you |
|---|---|---|
| Named in answers | Manual weekly prompts across 4 assistants | Share of voice per surface |
| Crawler hits | Server logs for GPTBot, PerplexityBot, etc. | Which pages are being read |
| Referral traffic | Analytics segments per AI domain | Direct proof of citation |
| Competitor mentions | Record who appears alongside you | Your real competition per surface |
Start with manual sampling, because it is the ground truth. Take the ten to twenty questions your buyers actually ask, the decision-stage ones like best X for Y and who makes the best Z, and run each through ChatGPT, Perplexity, Gemini, and Claude on a fixed cadence, weekly or biweekly. For each, record three things: are you named, how are you described, and which competitors appear. Keep it in a simple sheet. Over a few weeks you get a share-of-voice picture per surface and a running list of gaps that becomes your content backlog. This is unglamorous and it is the most useful thing you can do.
Then add server-side evidence, which is more objective. Your logs show the AI crawlers hitting you: GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot, Google-Extended, and CCBot. Watching those hits tells you which crawlers are reading which pages and how often, which is a leading indicator that the fast loop is working. A page that just started getting crawled by PerplexityBot is a page about to be eligible for citation, so it tells you where to look next.
Layer in referral and analytics signals where you can see them. Perplexity and other citation engines send real referral traffic when they link you, so a growing trickle of referrals from those domains is direct proof you are being cited, not just crawled. Set up the segments in your analytics and watch the trend. A handful of purpose-built tools have also appeared to monitor brand mentions across assistants at scale, and they are worth trialing once your manual process shows the pattern, but do not outsource the understanding to them before you have felt the signal yourself.
The cadence I recommend is weekly manual sampling on your priority questions, weekly a glance at crawler hits on your flagship pages, and a monthly rollup of share-of-voice per surface with the gap list refreshed. Treat it exactly like a rank-tracking program: measure, find the gaps, ship the fix, re-measure. The surfaces move faster than classic search, so the loop is tighter and the wins show up sooner.
A measurement caveat that keeps you honest: assistant answers are non-deterministic, so the same prompt can name you today and skip you tomorrow, and reading too much into a single run will drive you crazy. Do not treat one answer as a verdict. Sample each priority question a few times, note the pattern rather than the instance, and watch the trend across weeks. Share of voice is a distribution, not a binary. What you are looking for is a durable shift, you went from named in two of ten runs to named in seven, not a single lucky mention. Treating it statistically rather than anecdotally is the difference between a measurement program and a superstition.
It also helps to define what a win even looks like before you start, because the metric depends on your goal. If you are after awareness, being mentioned anywhere in the answer counts. If you are after traffic, only a linked citation that drives referrals counts, so you weight the analytics segments most heavily. If you are after conversion, the metric is being the recommended option for decision-stage questions, which is a narrower and more valuable bar than a passing mention. Decide which of the three you are optimizing for, instrument that one, and judge the program against it. A team measuring mentions while actually needing recommendations will feel successful and sell nothing, so name the real target up front and let it dictate what you count.
A worked example, and where this goes next
Let me make it concrete with a case I use often, then look at where the whole game is heading, because the destination changes how you should build today.
Say you sell ergonomic office chairs and you want ChatGPT and Perplexity to recommend you for back pain. The vague version of your page says we make premium chairs designed for comfort, wrapped in hero images, with no author, no date, and no numbers. A model has nothing to grab and no reason to trust it, so it cites a competitor. Now rebuild it the way this article describes. Lead with a direct answer: the best ergonomic chair for lower-back pain pairs adjustable lumbar support with a seat depth of 17 to 20 inches and a recline that unloads spinal pressure, typically priced between two and six hundred dollars. Add a spec comparison table across your models and two named competitors, honestly. Publish original data: a survey of 500 owners reporting a 41 percent drop in end-of-day back pain after eight weeks, credited to your brand. Mark it all up with Product, FAQPage, and Article schema, a named author with a real bio, and an honest dateModified. Allow the AI crawlers, add llms.txt, and earn two mentions in furniture roundups your buyers read.
Watch what happens across the five gates. The page is now retrievable because it is indexed and relevant, parseable because it is server-rendered HTML with real text and a table, trustworthy because it has an author, dates, and a real study, quotable because the direct answer and the 41 percent stat are clean liftable units, and corroborated because two third-party roundups name you the same way. It clears every gate, so ChatGPT names it, Perplexity cites the survey, and AI Overviews pulls the table. Same product, same company; the only thing that changed was making every signal a model needs explicit instead of hidden. That is the whole method in one page.
Now the horizon. The near future is agents. Assistants are moving from answering questions to completing tasks, reading your pages not just to quote you but to compare, shortlist, and eventually transact on a person's behalf. The site that is machine-readable, states its facts plainly, exposes structured data, and declares its actions in llms.txt is the site an agent can actually use. Getting cited is the first rung; getting acted on is the next one, and it is worth far more. The other durable trend is that original, first-hand, genuinely useful content keeps winning as models get better at detecting the generic and the synthetic. Everything in this article points the same direction: publish things worth citing, make them trivial for a machine to read and verify, keep your identity consistent, and stay the primary source. Do that, and you are not chasing the algorithm. You are building the thing every algorithm is trying to find.
Let me push the worked example one step further, into the agent future, because it shows why the structure is not busywork. Picture the same chair page a year from now, being read not by a person asking ChatGPT but by an agent a person delegated the whole task to: find me an ergonomic chair under 400 dollars for lower-back pain and add the best one to my cart. That agent does not skim your hero image or admire your brand voice. It parses your Product schema for price and availability, reads your direct answer and your survey stat to justify the recommendation, checks your third-party mentions for corroboration, and, if you declared a purchase action in llms.txt, it acts. The page that wins that moment is the exact same page that wins the citation today: readable, sourced, structured, honest about its facts. You did not need a separate agent strategy. You needed to do the citation work well and expose your actions, and the agent readiness came free.
The last thing I will leave you with is a way to hold all of this in your head, because the tactics are many and the principle is one. Every gate, every signal, every file comes down to a single instruction: be the clearest, most trustworthy, most verifiable source on the specific thing you want to be known for, and make that legible to a machine. Models are not trying to be tricked and cannot be tricked for long; they are trying to find sources they can safely rely on and quote. Stop thinking of this as gaming a system and start thinking of it as removing every reason a model has to doubt or overlook you. Do the substance, structure it cleanly, prove your identity, keep it fresh, and earn corroboration. The citations are just what happens when you have made yourself impossible to responsibly leave out.
Frequently asked questions
How do I get my business cited by ChatGPT?
Publish server-rendered pages with a clear 40 to 60 word direct answer near the top, mark them up with Article, FAQPage, and Organization schema, include original data and a named author, and make sure GPTBot and OAI-SearchBot are not blocked. ChatGPT cites sources it can read cleanly and trust.
Why is my competitor recommended by AI and I am not?
Usually their content is more citable: cleaner structure, a direct answer, original numbers, and stronger off-site corroboration. Models name the source they can parse and verify against other sites. Audit your top pages against the five signals and you will normally find the gap.
Does schema markup help with AI citations?
Yes. JSON-LD tells a model exactly what your content is, who wrote it, and when it changed, so it does not have to infer any of it.
What is an llms.txt file and do I need one?
It is a plain-text file at your site root that describes your business for AI crawlers and agents: what you do, your key URL patterns, flagship pages, and any actions an agent can take.
What is the difference between ai.txt and llms.txt?
ai.txt is a policy file that opts specific AI training crawlers like GPTBot, ClaudeBot, and Google-Extended in or out. llms.txt is a descriptive map of your business and content written for a model. One controls access; the other improves legibility. You want both.
How long until AI assistants start citing my content?
On live-retrieval surfaces like Perplexity and AI Overviews, a fixed and re-crawled page can be cited within days to a few weeks. Training-memory citations, where a model knows you with browsing off, take months and depend on durable, widely-referenced content. Build for the fast loop first.
Can paywalled or gated content get cited?
Rarely. A model cannot quote what it cannot reach, and training crawlers will not pay or fill a form. If you want a page cited, keep it crawlable and unpaywalled. Gate lead-magnet material you do not need cited, and keep your citable flagship content open.
Do I need original data to get cited, or is good writing enough?
Good structure gets you into the running; original information wins the citation. A model has little reason to cite the tenth page repeating the same advice.
How do I know if AI crawlers are even reading my site?
Check your server logs for user agents like GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot, and CCBot. Their hits tell you which pages are being read and how often, and rising crawls on a page are a leading indicator that it is about to become eligible for citation.
Is getting cited by AI different from ranking on Google?
It overlaps but is not identical. Classic Google ranking is most of the entry ticket for Gemini and AI Overviews, but generative citation adds emphasis on direct answers, original data, off-site corroboration, and clean parseability. A focused page can get cited above a higher-authority page that buries the same answer.
Should I block AI crawlers to protect my content?
It is a legitimate choice, but make it deliberately. Blocking removes you from the surfaces you might want to win. Many businesses opt training crawlers in for the discovery value while reserving rights on specific sections.
Which AI surface should I optimize for first?
Start with Perplexity. It retrieves live for almost every query, shows its sources openly, and re-crawls fast, so it is the best test bench to prove your citation strategy. Once your fixes work there, the same structured, sourced pages carry over to ChatGPT, Gemini, and Claude.
I'm Frederick Sona, and I've spent most of my career chasing one question: why do some brands break through while others, often the better ones, don't? I've looked for the answer as a marketer, a designer, a technologist, a salesperson, and a founder, and the honest answer is that it takes all of it: being easy to find, easy to trust, and easy to buy from. Search Everywhere Optimization is one piece of how I think about that, but this blog covers the whole picture, from search and technology to brand, design, and the work of turning attention into revenue. If any of this was useful, come say hello at fredericksona.com.