
The Structured-Truth Method: One Source of Truth for Every Search Surface
Stop optimizing each search surface separately. Maintain one body of structured, verifiable, machine-readable truth, then syndicate it everywhere. The full method, step by step, from Frederick Sona.
Why optimizing each surface separately fails
Most teams treat discovery as a list of channels. There is an SEO person, maybe an agency running ads, someone who posts to LinkedIn, a founder who occasionally answers on Reddit, and a product page that nobody has touched since launch. Each of those efforts optimizes its own surface in its own tool, on its own schedule, with its own copy. It feels like coverage. It is actually fragmentation, and fragmentation is why brands with better products keep losing to brands that are simply easier for a machine to read.
The mechanism is simple. When you optimize surface by surface, you produce a different version of the truth in each place. The founding date on the About page says 2019. Crunchbase says 2020. The LinkedIn company page never got filled in. The pricing on the marketing site is a range, the pricing in the AI answer was scraped from a two-year-old blog post, and the pricing a sales rep quotes is a third number. None of these are lies. They are drift. And drift is poison to the systems that now decide who gets found, because those systems are cross-referencing you against yourself and discounting whatever they cannot reconcile.
The old world tolerated drift because a human read one page at a time. A person landing on your site did not also open Wikidata, your schema, three review sites, and a language model in parallel and compare the numbers. Machines do exactly that. An answer engine assembling a response about your category pulls a fact from here, an entity reference from there, a rating from somewhere else, and it trusts the version that is consistent, sourced, and repeated. If your truth is scattered and contradictory, the machine either picks the wrong version or, more often, skips you and cites a competitor whose story lines up.
Fragmented optimization also does not compound. Every surface you tune in isolation is a sandcastle. The SEO work does not make the AI answer better. The social post does not strengthen the knowledge panel. You are paying full price for each surface and getting single-surface returns, when the whole promise of modern discovery is that one strong signal can pay off in ten places at once. I wrote the pillar on this, the 19 ranking surfaces, to map the terrain. This piece is about the operating system that makes all nineteen tractable without hiring nineteen specialists.
The symptoms are easy to spot once you know them:
- The same fact appears three different ways across your own properties, and nobody owns which one is correct.
- A new page ranks, an old page contradicts it, and neither gets fixed because they live in different teams.
- You cannot answer "what does ChatGPT say about us" because you have never defined what it should say.
- Every surface has its own content calendar and none of them share a source.
- When a core fact changes, updating it everywhere is a manual scavenger hunt that never finishes.
If two or more of those are true, you do not have a content problem or a channel problem. You have a truth problem, and no amount of per-surface tuning fixes a truth problem. You fix it by changing what you maintain: not nineteen optimized surfaces, but one source of truth that every surface reads from.
There is a deeper reason fragmentation is getting more costly, not less. The discovery layer used to be a set of separate audiences you could address separately, because Google's crawler and a shopper's eyeballs and a marketplace's listing did not talk to each other. That is over. The systems that now decide who gets found are increasingly the same systems wearing different masks: a language model behind the AI Overview, the same class of model behind the shopping agent, the same entity graph behind the knowledge panel and the voice answer. Optimize for one and contradict yourself on another, and you are not confusing two different audiences. You are confusing one audience that reads all of your properties at once and remembers the discrepancy. The penalty for inconsistency compounds precisely because the readers have merged.
I have watched this play out with real companies, and the pattern is depressingly consistent. The team that wins is rarely the one with the biggest content budget. It is the one whose facts are boring and identical everywhere, whose entity resolves cleanly, whose answers are easy to lift. The better product with the messier truth loses the citation to the adequate product with the cleaner truth, over and over, because the machine cannot tell how good your product is until it can first tell what is true about it. Structured truth is upstream of everything, which is exactly why fixing it beats optimizing anything downstream of it.
What structured truth actually means
Structured truth is a deliberately plain name for a specific thing: a single, maintained body of facts about your business, expressed in a form both humans and machines can read without guessing. It is not a document. It is not a brand guideline PDF. It is the authoritative set of claims you are willing to stand behind, held in one place, marked up so software can consume it directly, and syndicated outward so every surface is drawing from the same well.
Break the word in half. "Truth" is the content: the facts that are actually true about you and that you want the world, and the machines that increasingly speak for the world, to repeat. "Structured" is the form: organized into entities and fields and marked up in a schema a parser understands, rather than buried in prose a machine has to interpret and might get wrong. Prose is lossy. "We have served thousands of customers since our early days" is a sentence a machine cannot use. "foundingDate: 2019, customersServed: 4,200" is data a machine can cite. Same truth, different structure, wildly different outcome on a discovery surface.
Five kinds of things make up your structured truth, and it helps to name them because you will maintain each one differently:
- Entities. The nouns of your business: your organization, your people, your products, your locations, your categories. Each is a thing with a stable identity, ideally a canonical identifier like a Wikidata QID, so every reference resolves to the same object.
- Canonical facts. The specific attribute values attached to those entities: founding date, headquarters, price, capacity, dimensions, hours, certifications, headcount. One agreed value per fact, with a source.
- Direct answers. The clean, quotable responses to the real questions people ask about your category, written to be lifted verbatim by an answer engine or a voice assistant.
- Schema markup. The machine-readable layer, Schema.org JSON-LD, that wraps entities, facts, and answers in types a crawler parses without ambiguity.
- Original data. Proprietary numbers only you can produce: your own study, your benchmark, your survey. These are citation magnets, the things models quote because no one else has them.
When those five are consistent with each other and consistent across every place they appear, you have structured truth. When they are not, you have marketing. The difference matters because the discovery layer has shifted from ranking documents to resolving facts. A modern engine does not just ask "which page is most relevant?" It asks "what is true about this entity, and who can I trust to tell me?" Structured truth is your answer to that question, prepared in advance, in the format the question is actually asked in.
There is a useful mental test. Pick any fact about your business and ask: if a language model, a voice assistant, and a shopping agent each fetched this fact independently, would they get the same value, and would they know where it came from? If yes, that fact is structured truth. If the three would disagree, or if none could find a sourced value at all, that fact is a liability sitting on your books waiting to cost you a citation. The method that follows is how you turn every fact that matters from a liability into an asset.
The method at a glance: one loop, seven moves
The Structured-Truth Method is a loop, not a launch. You run it once to stand the system up, then you run it forever at a slower cadence to keep it honest. Seven moves make up the loop, and the order matters, because each move depends on the one before it. Skip ahead and you end up marking up facts you never agreed on, or distributing answers that contradict your own schema.
Here are the seven moves in sequence:
- Define. Decide your canonical facts and entities. Write down the one true value for every fact that matters, name every entity, and pick the identifiers that make each entity resolvable.
- Mark up. Wrap those entities and facts in Schema.org JSON-LD so machines read them as data, not prose. Validate that the markup parses and matches what a human sees on the page.
- Answer. Publish direct answers to the real questions in your category. Lead with the answer, keep it short enough to lift, and phrase the questions the way people actually ask them.
- Expose. Open the doors machines use. Ship llms.txt, ai.txt, a clean sitemap, feeds, and an API where it fits, so crawlers and agents can find and fetch your truth on purpose.
- Distribute. Push the same truth to every surface: your site, your knowledge-graph entries, your profiles, your marketplaces, your social accounts. Same values, everywhere, syndicated from the source.
- Measure. Track citations and coverage, not just rankings. Ask the engines your target questions and record which they answer with: your truth, a competitor's, or a stale version of yours.
- Refresh. When a fact changes, change it at the source and let it propagate. Retire stale numbers, restamp modified dates, and rerun the loop on a schedule.
Notice what is not on that list: "optimize surface X." There is no step where you sit down and tune your Reddit presence in isolation from your schema. Every surface-level result is a downstream effect of the same upstream truth. That is the entire point. You are not doing nineteen jobs. You are doing one job, defining and maintaining truth, and letting it express itself across nineteen surfaces.
The loop has a natural rhythm. Define, mark up, answer, and expose are heavy the first time and light afterward, because you are building assets that persist. Distribute and measure are ongoing and roughly weekly. Refresh is event-driven, triggered by a fact changing, plus a standing quarterly sweep to catch drift you missed. Most teams over-invest in distribution and under-invest in define and refresh, which is exactly backwards. A clean source refreshed on time distributes almost for free. A messy source refreshed never turns every distribution channel into another place your contradictions are on display.
The rest of this article walks each move in order, with the specific tactics, the schema, and the mistakes that sink people. Read it as a build sequence, not a menu. If you only have time for two moves this quarter, do Define and Mark Up, because everything downstream inherits their quality. You cannot syndicate, answer, or measure your way out of a truth you never actually agreed on.
Move one: define your canonical facts and entities
Before you touch a line of schema, you need to know what is true. This sounds obvious and is almost never done. Most companies have never written down, in one place, the authoritative value for the facts a machine will ask about them. They have a marketing site that implies things, a sales deck that claims other things, and a founder's memory that overrides both. Define is the unglamorous move where you replace all of that with a single sheet: one fact, one value, one source.
Start with entities, because facts hang off entities. List the nouns that matter: your organization, each founder and key person, each product or service line, each physical location, and the handful of categories you want to be known for. For each entity, decide on a canonical name and spelling, and pin down an identifier. The organization gets a Wikidata QID if you can support one, or at minimum a canonical URL that everything points back to. People get a canonical bio and a set of sameAs links. Products get a stable SKU or slug. The rule from the knowledge-graph playbook is blunt and correct: inconsistent naming fragments your entity, so "Acme", "Acme Inc", and "acme" cannot be three different things in three different places. Pick one primary and list the rest as known aliases.
Then enumerate facts. For each entity, write the attribute values a person or a machine would actually ask for. For an organization that is founding date, headquarters, headcount, industry, official website, and contact. For a product it is price, availability, dimensions, materials, warranty, and whatever specs define the category. For each fact, record three things: the value, the source that proves it, and the owner who is allowed to change it. The source is what turns a claim into structured truth. "Founded 2019" is a claim. "Founded 2019, per the incorporation filing" is a fact you can defend when a machine or a journalist checks.
A few disciplines separate a real canonical sheet from a nice-looking spreadsheet nobody trusts:
- One value per fact. If a fact genuinely has a range, the range is the value, stated as a range, not two different point numbers in two different places.
- Every fact has an owner. Facts without owners rot, because no human is accountable for keeping them true.
- Every fact has a source. If you cannot cite where a number comes from, you are not ready to publish it as truth.
- Prefer specific over vague. "4,200 customers" is structured truth. "Thousands of happy customers" is noise a machine discards.
- Flag the volatile ones. Price, availability, and headcount change; mark them so refresh knows where to look.
Scope this sheet ruthlessly on the first pass. The instinct is to catalog everything, and that instinct is how the exercise dies in a swamp of two hundred rows nobody maintains. You do not need every fact about your business. You need the facts a machine will ask about and the facts a customer decides on. For most companies that is a surprisingly short list: a dozen organization facts, a handful of attributes per product, the answers to the ten questions your category actually asks. Get those genuinely clean and sourced before you add a single row beyond them. A short canonical sheet that is completely trustworthy beats a comprehensive one that is half-verified, because the whole value of the sheet is that everything on it can be trusted without checking.
Keep this sheet in one place that the whole organization treats as authoritative. The specific form barely matters: a well-structured spreadsheet, a content model in your CMS, a database, or a git-tracked YAML file all work. What matters is that there is exactly one, that changes go through it, and that everything downstream, your schema, your answers, your profiles, is generated from it or reconciled against it. This sheet is the source in "single source of truth." Every later move is just an act of publishing it faithfully.
Move two: mark up the truth so machines read it
Once the facts are agreed, you make them machine-readable. This is Schema.org JSON-LD, and it is the single highest-value technical work in the whole method, because every surface downstream consumes it. Marking up turns your content from prose a machine has to interpret into data a machine can read directly. A person reads your About page and infers you were founded in 2019. A machine reading Organization schema with foundingDate 2019 does not infer anything. It knows.
| Schema type | Carries | Surfaces it feeds |
|---|---|---|
| Organization | Entity, logo, sameAs, founding | Knowledge panel, LLMs, brand SERP |
| Product | Price, availability, specs | Shopping, marketplaces, visual, AI answers |
| FAQPage | Direct answers | AI Overviews, voice, featured snippets |
| Article + Person | Author, dateModified, headline | SEO, GEO citations, E-E-A-T |
| HowTo | Ordered steps | AI answers, voice, rich results |
| ImageObject | Caption, credit, alt | Visual search, Lens, Pinterest |
You do not need every schema type. A small set carries most of the load. Organization schema on your homepage and, ideally, on every page, is the backbone: it declares your entity, your logo, your founding date, and, critically, your sameAs array linking to every profile and knowledge-graph entry that reinforces the same identity. Product schema carries price, availability, and specs, and one correct Product block feeds shopping results, marketplaces, visual search, and AI shopping answers at the same time. FAQPage schema wraps your direct answers so answer engines can lift them. Article schema with an author Person and a dateModified carries your editorial credibility. HowTo, LocalBusiness, ImageObject, and BreadcrumbList fill in the rest as your content demands.
The table below maps the workhorse types to what they feed, so you can prioritize by coverage rather than by whatever your CMS makes easy.
Two rules make schema work and their absence makes it worthless. First, validate. Broken markup helps nothing, and a surprising amount of schema in the wild does not parse. Run every type through a validator until it is clean, and keep it clean as pages change. Second, match the visible content. The structured data has to say the same thing the human sees, because mismatches get penalized and, worse, erode the trust that makes a machine willing to cite you at all. Schema that claims a price the page does not show is not clever, it is a flag.
The move that ties everything together is sameAs. Your Organization schema should link out to your Wikidata entry, your LinkedIn, your Crunchbase, your verified social profiles, and any other canonical reference to the same entity. This is how the structured web resolves "these are all the same company." Without sameAs, the graph never connects your properties and every profile floats alone. With it, a signal earned in one place reinforces your identity in all of them. A few specifics worth doing on the first pass:
- Put Organization schema on the homepage first, then propagate it site-wide so branded queries have a consistent entity everywhere.
- Generate schema from the canonical sheet rather than hand-writing it per page, so a fact changes in one place and the markup updates everywhere.
- Include dateModified on anything dated, visible to crawlers and to humans, because fresher answers win and stale stamps get discounted.
- Layer Speakable onto your TL;DR and FAQ blocks so the same answers compose into voice results.
- Keep identifiers stable. A QID or canonical URL that changes breaks every reference pointed at it.
Done well, this move is quietly transformative, because it is write-once, benefit-everywhere. The same Product schema you validate today is still feeding shopping agents and visual search a year from now with no additional work, as long as the underlying fact stays true. That is the compounding the fragmented approach never gets: one correct block of markup paying off across every surface that reads it.
Move three: publish direct answers
Structured facts feed machines. Direct answers feed the surfaces that speak in sentences: AI Overviews, Perplexity, voice assistants, featured snippets, and the synthesis layer of every language model. These systems do not rank your page and hand a user a link. They extract a quotable span and read it out. Your job in this move is to make your answer the easiest clean span to extract in your entire category.
The format is not subtle. Lead with the answer. State it plainly in the first sentence or two under a heading that matches the question. Around a third of Google queries already trigger an AI Overview, and the extraction systems that build those overviews grab the clean answer near the top and skip the preamble. The old SEO habit of burying the payoff under 800 words of throat-clearing is now actively harmful, because you are hiding the exact thing the machine came to take. Every long page in the method gets a TL;DR block up top: two to four sentences that answer the page's core question directly, written to be lifted verbatim.
Then you build FAQs, and you build them from real questions. Open your search console, your sales call notes, your support tickets, and the "People also ask" boxes, and collect the questions people actually type and speak. Phrase your FAQ entries in that exact language, not in marketing language. "Why is our platform the best?" is not a question anyone asks and it will never get cited. "How much does X cost?" and "Is X better than Y for small teams?" are questions, and they map to answers a machine will quote. Keep each answer tight, ideally under 50 words, because short answers fit the snippet slot and long answers get truncated at an unflattering point.
A few patterns get cited far more than their share, and they are worth writing on purpose:
- The numbered list. "The five things a CFO checks first" gives an engine a clean, liftable structure.
- The comparison. "X vs Y vs Z" content owns a huge share of decision-stage answers, because that is the shape of the question.
- The by-something table. Prices by plan, options by use case, specs by model. Answer engines love a table they can read straight into a response.
- The definition. A crisp "X is ..." opener that a machine can quote as the canonical explanation of a term you want to own.
Tone matters more here than anywhere else in the method. Answer engines and the models behind them reward confidence and discard hedging. "It depends" and "many people think" are spans no system wants to read aloud as an answer. If a fact is true, state it as a fact. This is not an invitation to overclaim; it is an instruction to stop qualifying things that do not need qualifying. Your canonical facts from move one are exactly what let you write with this confidence, because you already did the work of deciding what is true and sourcing it. The direct answer is just that truth, phrased for a microphone. Wrap every one of these answers in FAQPage schema so the structure and the content travel together, and you have turned your facts into the raw material every answer surface is built to consume.
Move four: expose your truth to machines
You have agreed on facts, marked them up, and written answers. Now you open the doors machines use to find and fetch them. Most sites are built for human navigation, with menus and links a person clicks. The discovery layer that matters now arrives as crawlers and agents, and those visitors want different doors: files that describe your content, feeds that stream your updates, and endpoints that expose your actions. Exposing is the move where you stop making machines guess and start telling them exactly what you have.
Three small files at your domain root do a lot of work. llms.txt is a clean, machine-readable description of what you do, your primary URL patterns, and the actions available on your site, written for a model rather than a shopper. It is the difference between a language model reverse-engineering your sitemap and a language model reading a plain-language map you handed it. ai.txt states your policy toward training crawlers, opting GPTBot, ClaudeBot, Google-Extended, PerplexityBot, and CCBot in or out per your choice; if you want to be in the training data that future models are built from, this is where you say so explicitly. And a real sitemap.xml, submitted to Search Console and Bing Webmaster Tools, still anchors the whole thing.
Beyond the root files, expose the truth as data streams and endpoints wherever your content justifies it:
- Feeds. An RSS or JSON feed of your articles and updates gives crawlers and aggregators a low-friction way to stay current with your freshest truth.
- Product feeds. A structured product feed keeps marketplaces, shopping surfaces, and AI shopping agents synced to your real prices and availability instead of a scraped snapshot.
- An API or MCP endpoint. When agents will act on your business, not just read about it, an endpoint that exposes your prices, availability, and core actions in a machine-completable shape is how you stay operable in agent-driven commerce.
- Clean, server-rendered HTML. Training and retrieval crawlers parse HTML far more reliably than JavaScript-only pages, so anything you want read should render without a headless browser.
There is a strategic point hiding in this move. The GEO playbook is explicit that heavy JavaScript rendering blocks training crawlers and that no flagship content should sit behind a paywall you want cited, because the crawlers will not run your scripts and will not pay your toll. Exposing is partly about adding files and partly about removing obstacles. Every gate between a crawler and your truth is a place you have chosen to be less findable. If a page matters for discovery, it should be reachable, renderable, and free to read.
Exposing is also where the method quietly future-proofs itself. Today most of these doors are used by crawlers that read. Increasingly they will be used by agents that act, booking, buying, and transacting on a user's behalf. The brands whose prices, availability, specs, and core actions are already structured and machine-completable will be the ones an agent can operate on. The brands whose critical actions are illegible to a machine will be skipped silently, with no error message and no second chance. Building these doors now is cheap. Building them under pressure once agent commerce is the norm will not be.
Move five: distribute the same truth everywhere
Now you syndicate. Distribution is where the single source of truth actually reaches the nineteen surfaces, and the discipline of the whole method lives or dies here. The rule is one sentence: every surface publishes the same values, drawn from the same source, spelled the same way. If your canonical sheet says founded 2019, then your About page, your Wikidata entry, your Crunchbase profile, your LinkedIn page, and your press kit all say 2019. Not approximately. Identically.
Think of distribution as pushing one record outward through concentric rings. The innermost ring is your own site: homepage Organization schema, product pages, the About page, the editorial articles. The next ring is the knowledge graph: your Wikidata entity with its claims set, and, when notability supports it, Wikipedia. Wikidata is permissive and can be filed early; Wikipedia is strict and punishes you for filing without independent press citations, so respect the difference. The next ring is your verified profiles: LinkedIn, Crunchbase, your social accounts, App Store listings if you have an app. The outer ring is the surfaces you influence but do not own: marketplaces, review sites, the communities where your category gets discussed, and the language models that read all of the above.
The reason consistency matters so much is that the machines reading these rings cross-reference them against each other. This is the sameAs principle from move two, playing out across the whole web. When your schema links to your Wikidata entry, and your Wikidata entry links back to your site, and both agree with your Crunchbase profile, a machine resolves all of them to one confident entity and trusts the facts they share. When they disagree, the machine cannot reconcile you and discounts the lot. Consistency is not tidiness for its own sake. It is the signal that tells the structured web your facts are reliable enough to repeat.
Distribution is easy to describe and hard to sustain, so build it to run from the source rather than by hand:
- Generate, do not retype. Where a surface accepts structured input, a feed, an API, a schema block, drive it from the canonical sheet so a fact travels automatically.
- Keep a distribution map. List every place each canonical fact appears, so when the fact changes you know the full blast radius. This map is what makes refresh possible.
- Spell the entity identically everywhere. One primary name, the same aliases, no stray variants that fragment your identity.
- Sequence the rings. Get your own site and schema right first, because the outer rings and the language models cite inward. A clean core makes every outer ring easier.
- Never let a surface freelance its own facts. A profile that someone fills in from memory is how drift re-enters a system you worked to make consistent.
The payoff of doing distribution from a single source is the compounding the fragmented approach never reaches. A press mention that repeats your canonical numbers strengthens the same entity your schema declares and your Wikidata entry confirms. Every consistent reference is a vote, and the votes add up to an entity the machines are confident about. That confidence is what gets you cited, paneled, and recommended, across all nineteen surfaces, from one body of truth you maintained once and syndicated well.
Move six: measure citations and coverage, not just rankings
The old scoreboard was rankings: where you sit in the blue links for a keyword. That number still matters, but it describes one surface out of nineteen, and it says nothing about how an AI answer treats you, how a voice assistant read you, or how a shopping agent could transact on you. The Structured-Truth Method measures a different thing: citation and coverage. Are the engines answering with your truth, and are they answering with it everywhere they should?
Start with the most direct measurement there is, and the one almost nobody does systematically: ask the engines. Take your target questions, the ones your FAQs answer, and put them to ChatGPT, Claude, Gemini, Perplexity, and Google's AI Overview. Record what comes back. Does your brand surface? Are the facts correct? Are they current, or is the model quoting a number you retired last year? Is a competitor cited where you should be? This is manual, it is tedious, and it is the single most honest read on how the discovery layer sees you. Do it on a fixed cadence so you are watching a trend, not a snapshot.
Around that, build a measurement stack that watches the whole funnel from crawl to citation:
- Crawler hits. Your server logs show GPTBot, ClaudeBot, PerplexityBot, and CCBot arriving. If they are not visiting, your expose move is not working, and nothing downstream can.
- Coverage. For each canonical fact, track how many of its intended surfaces actually carry the correct value right now. This is the number that tells you distribution is holding.
- Citation share. For your target questions, what fraction of AI answers cite you versus a competitor versus a stale source. This is the modern equivalent of share of voice.
- Panel and entity health. Does the knowledge panel appear for branded queries, are the Wikidata claims correct, does brand autocomplete resolve cleanly.
- Classic signals. Search Console impressions, clicks, and AI Overview appearances still belong on the dashboard; they are one ring, not the whole picture.
The reason to measure coverage and citation rather than only rankings is that they are the metrics the method can actually move. When citation share drops, you have a specific, fixable cause: a fact drifted, a page fell out of the index, a competitor published fresher data, a model is quoting a stale snapshot. Each of those points back to a move you can rerun. A ranking drop, by contrast, often gives you nothing to act on but anxiety. Coverage and citation are diagnostic. They tell you which fact, on which surface, needs which move.
There is a maturity ladder worth naming. At the bottom, you check rankings and hope. In the middle, you spot-check AI answers occasionally and fix what you notice. At the top, you have a standing list of target questions, a coverage number per canonical fact, and a monthly review where every gap is assigned to an owner and a move. You do not need enterprise tooling to climb it. A spreadsheet of questions, a monthly hour asking the engines, and honest notes will put you ahead of nearly every competitor, because nearly every competitor is still staring at a rankings chart and wondering why the AI does not mention them.
Move seven: refresh before your truth goes stale
Truth has a shelf life. Prices change, headcount grows, a product ships a new version, a certification renews, a founder leaves. The moment a canonical fact changes and your published truth does not, you have reintroduced exactly the drift the whole method exists to prevent. Refresh is the move that keeps the system honest over time, and it is the one teams neglect most, because standing a system up feels like progress and maintaining it feels like chores.
Refresh has two modes, and you need both. The first is event-driven: when a fact changes, you change it at the source and let it propagate. This is where the distribution map from move five earns its keep, because it tells you the full blast radius of a change. Update the price on the canonical sheet, and the map shows you every place that price lives, your product schema, your feed, your marketplace listings, your pricing page, so nothing gets missed. Without the map, an event-driven update is a scavenger hunt that always leaves a straggler behind, and the straggler is what a machine cites six months later.
The second mode is the standing sweep. On a fixed cadence, quarterly is a sane default, you walk the canonical sheet and reverify. Models prefer fresh citations and discount stale ones, so the sweep is not busywork; it is how you stay the current source rather than the outdated one. Watch for these staleness signals in particular:
- A fact whose source is now older than your refresh window and has not been reconfirmed.
- A dateModified stamp that has not moved on a page whose content actually changed.
- An original statistic from a study that is now a year or two old and needs a rerun.
- A profile or knowledge-graph entry that still shows a value you have already changed elsewhere.
- Any fact flagged volatile at Define that has not been checked this cycle.
When you refresh, refresh the whole record, not just the visible copy. Change the number on the page, yes, but also bump the dateModified in the schema so crawlers see the freshness, update the feed, and reconfirm the profiles. Half a refresh is arguably worse than none, because now the page and the schema disagree, and a mismatch between visible content and structured data erodes the trust that gets you cited. The refresh is only done when every ring in your distribution map agrees again.
Refresh is also the move that keeps original data valuable. Your proprietary study or benchmark is a citation magnet the first year and a liability the third, once the numbers are visibly out of date and a competitor has published something newer. Building a rerun into the calendar, an annual survey, a quarterly benchmark, keeps you the fresh source models reach for. Governance, which the next section covers, is what makes refresh reliable instead of heroic: clear owners, a change process, and a cadence, so keeping the truth current is a routine somebody runs, not a fire somebody fights.
How it maps onto the 19 surfaces and the five behaviors
This method is not a competitor to the 19 ranking surfaces. It is the engine that makes optimizing all nineteen tractable. The pillar names the surfaces and the five behaviors that win across every one of them. The Structured-Truth Method is how you actually practice those behaviors from a single source instead of nineteen scattered efforts. Put the two together and the surfaces stop looking like nineteen jobs and start looking like nineteen outputs of one system.
| Surface cluster | Truth it consumes | Moves that feed it |
|---|---|---|
| Classic search (SEO, CWV, E-E-A-T) | Entities, answers, author credibility | Define, mark up, answer |
| AI answers (AEO, GEO, AAO) | Direct answers, schema, exposed content | Answer, mark up, expose |
| Voice and visual (VSO, VxSO) | Speakable answers, ImageObject markup | Answer, mark up |
| Commerce (marketplaces, shopping, ASO) | Product truth as a live feed | Mark up, expose |
| Social and community | Consistent facts and canonical answers | Define, distribute |
| Local, global, identity (KGO, Web3) | Resolved, consistent entity | Define, distribute |
The five winning behaviors from the pillar map cleanly onto the moves in this method. Structured truth through schema is moves one and two, define and mark up. Direct answers is move three. Machine-readable exposure, the doors that let crawlers and agents in, is move four. Consistent entity identity across the web is move five, distribution. And the discipline of freshness and credibility runs through measure and refresh. The method is, in a real sense, the operational version of the behaviors: the same principles, sequenced into moves you can actually run on a Tuesday.
The mapping onto the surfaces themselves is where it gets concrete. Every surface consumes some slice of your structured truth, and the table below shows which slice feeds which cluster. Read it as an argument: notice how few distinct assets are doing all the work. One correct Organization block, one clean Product feed, one set of direct answers, and one Wikidata entity are, between them, feeding almost every surface on the list.
What the table makes obvious is the compounding. Classic search wants your entities, your answers, and your credibility signals, which are moves one through three. AI answer engines want your direct answers, your schema, and your exposed, crawlable, unpaywalled content, which are moves three and four. Voice wants your answers with Speakable layered on, a byproduct of move three. Visual search wants your ImageObject markup and honest alt text, a slice of move two. Commerce surfaces want your Product truth exposed as a feed, moves two and four. Knowledge and identity surfaces want your entity resolved and consistent, which is the entire point of moves one and five. The same small pile of well-made assets satisfies every cluster.
This is why I keep insisting the method is one job, not nineteen. A team that internalizes structured truth is not choosing between SEO and GEO and voice and commerce optimization and rationing its attention across them. It is maintaining a source of truth and letting each surface draw from it. The surface-specific work that remains, a marketplace's particular field, a voice assistant's particular action shortcut, is a thin adapter on top of a shared foundation, not a separate building from the ground up. Get the foundation right and the adapters are cheap. Skip the foundation and every adapter is load-bearing, which is exactly the fragile, uncompounding situation the fragmented approach traps you in.
Governance: keeping the truth consistent across places and people
A single source of truth is a governance problem wearing a technical costume. The schema and the feeds are the easy part; software does what it is told. The hard part is people, because a company is a crowd of humans who each, in good faith, publish facts. Sales updates a deck. Marketing rewrites the homepage. A new hire fills in the LinkedIn page from memory. Every one of them is a potential source of drift, and no amount of clean JSON-LD survives an organization where anyone can publish any fact from anywhere with no reconciliation. Governance is the set of rules that keeps the truth single as the company scales.
The first rule is ownership. Every canonical fact has exactly one owner, a named human who is accountable for its accuracy and the only one authorized to change its value at the source. Ownership is what stops a fact from rotting, because rot happens when everyone assumes someone else is watching. When the price owner changes the price, that change is legitimate and propagates. When a well-meaning contractor changes the price on a landing page, that is not an update, it is drift, and governance is what catches it.
The second rule is a change process. Facts change through the source, not around it. Nobody edits the value on a downstream surface directly; they change it on the canonical sheet and let it flow out through the distribution map. This feels bureaucratic until the first time a fact changes cleanly across eleven surfaces because it went through the process, instead of getting fixed in three places and forgotten in eight. The process is what turns a change from a scavenger hunt into a routine.
Governance also means writing down how you decide and defend the truth, which is where the E-E-A-T discipline becomes concrete:
- An editorial standards page that states how you research, source, and fact-check, and how a reader can flag an error.
- A named author on every article, with a real bio and sameAs links, so credibility resolves to a person, not an anonymous team.
- A corrections policy, so when a fact was wrong, the fix is visible and dated rather than quietly swapped.
- Disclosure of affiliate and sponsor relationships inline, because trust is a ranking input now, not just an ethic.
- One person or small group who owns the canonical sheet itself and adjudicates conflicts when two owners disagree.
Consistency across places is the outcome governance exists to produce, and precision about why it pays is what earns the effort. The machines that decide who gets found are, at bottom, trust engines. They reward the entity whose facts line up across every property and whose claims are owned, sourced, and dated. An organization with clean governance looks, from the outside, like a reliable source, and reliable sources get cited. An organization without it looks like noise, three founding dates and a contradiction, and noise gets skipped no matter how good its product is.
Be honest that governance is where most methods go to die, because it is the part with no dopamine. Defining facts feels like clarity, marking them up feels like building, measuring feels like progress. Enforcing that a contractor changed a number the right way feels like nagging. That asymmetry is exactly why you have to design the enforcement into the system rather than relying on willpower. Make the canonical sheet the path of least resistance: if changing a fact at the source is genuinely easier than editing it on a surface, people change it at the source, and governance stops being a fight. When the right way is also the easy way, consistency maintains itself.
The encouraging part is that governance scales down. A solo founder can run all of this in a single well-structured document and a monthly hour of discipline. A hundred-person company needs named owners, a change process, and a standing review. The shape is the same at both sizes: one source, clear owners, changes through the source, a cadence to catch drift. What you cannot do, at any size, is skip it and hope, because the drift the method prevents does not announce itself. It accumulates quietly, one well-meaning edit at a time, until a machine reads your contradictions back to a customer and cites your competitor instead.
Tooling: what to actually use
You can run the Structured-Truth Method with a spreadsheet and discipline, and for a small team that is genuinely the right answer. Tooling makes the loop faster and harder to break, but no tool substitutes for the decisions in move one. Buy tools to reduce friction on work you have already decided to do, not to decide for you. Here is what earns its place, organized by the job it does rather than by vendor, because the categories outlive the products.
| Job | Tool category | Move it serves |
|---|---|---|
| Hold the source | Spreadsheet, YAML, CMS model, or DB | Define |
| Emit and check markup | Schema generator + validator | Mark up |
| See the crawlers | Server log analysis | Expose, measure |
| Track AI citations | Answer-engine monitors, manual sampling | Measure |
| Watch classic search | Search Console, Bing Webmaster | Measure |
| Keep identity synced | Wikidata, profile management | Distribute |
For the source of truth itself, you want one authoritative store that everything generates from. Early on that is a well-structured spreadsheet or a git-tracked YAML or JSON file, which has the pleasant property of versioning every change to a fact for free. As you grow, a structured content model in your CMS or a small database becomes the home, with fields for value, source, owner, and volatility. The specific technology matters far less than the rule that there is exactly one and everything reconciles against it.
Around that core, a few categories of tooling do the heavy lifting:
- Schema generation and validation. Generate JSON-LD from the canonical store rather than hand-writing it, and validate every type in a rich-results tester until it parses clean. This is the highest-value tooling in the stack because it protects the highest-value move.
- Crawl and log analysis. Something that surfaces GPTBot, ClaudeBot, PerplexityBot, and CCBot hits in your server logs, so you can confirm the machines you are optimizing for are actually visiting.
- Answer-engine monitoring. Tools built to track how AI systems answer and cite are emerging quickly; Profound, AthenaHQ, and Perplexity's own analytics are examples. Until one fits, a spreadsheet of target questions and a monthly manual check beats waiting for the perfect product.
- Classic search consoles. Google Search Console and Bing Webmaster Tools remain the free, authoritative view of crawl status, impressions, and increasingly AI Overview appearances. Submit your sitemap to both.
- Entity and knowledge-graph tools. Wikidata itself, plus whatever you use to keep Crunchbase, LinkedIn, and social profiles consistent, so the identity ring of your distribution stays synced.
The table below sorts these by the move they serve, so you can staff the loop rather than accumulate a drawer of unrelated subscriptions.
One caution about tooling, learned the expensive way. It is easy to buy a monitoring product and mistake its dashboard for progress. A tool that tells you your citation share is falling is useful only if you act on it, and acting means rerunning a move, not admiring the chart. I would rather work with a team that runs the loop by hand and fixes what it finds than a team with a full stack of tools and no habit of responding to them. Start with the free consoles and a canonical spreadsheet, prove you will act on what they tell you, and add paid tooling to remove the specific friction that is slowing you down. Bought in that order, tools compound the method. Bought first, they become another surface you optimize and never read.
The mistakes that quietly sink the method
Most failures of the Structured-Truth Method are not dramatic. Nobody wakes up and decides to publish three founding dates. The method erodes at the edges, through small omissions that each seem harmless and together dismantle the consistency the whole thing depends on. Here are the ones I see most, and why each is more dangerous than it looks.
The first is skipping Define and jumping to schema. It is tempting, because marking up feels productive and agreeing on facts feels like a meeting. But schema built on facts you never actually agreed on just encodes your existing contradictions in a format machines read more reliably. You have made your drift more legible, not less. Define is the boring foundation, and every move above it inherits its quality or its rot.
The second is schema that does not match the page. Structured data claiming a price the visible content does not show reads, to a machine, as a manipulation attempt, and it erodes the trust that makes a machine willing to cite you at all. The fix is to generate schema from the same source as the visible content so they cannot diverge, and to validate that it parses. Broken or mismatched markup is worse than no markup, because no markup is neutral and bad markup is a negative signal.
The rest of the common failures cluster into a short list worth keeping visible:
- Burying the answer. Eight hundred words of preamble before the payoff hides the exact span an answer engine came to lift.
- FAQs written as marketing. "Why are we the best?" is not a question anyone asks and will never get cited; real user phrasing is what gets cited.
- Paywalling or JS-gating content you want cited. Training crawlers will not pay your toll or run your scripts, so gated flagship content is invisible to the surface you built it for.
- Anonymous authorship. A machine cannot confidently cite a faceless team; a named author with a bio and sameAs is a low-effort, high-weight signal you are leaving on the table.
- Stale original data. Last year's proprietary stat, unrefreshed, drifts from citation magnet to liability as competitors publish fresher numbers.
- Fragmented entity naming. "Acme", "Acme Inc", and "acme" as three variants fragments the very identity you are trying to consolidate.
- Refresh that never happens. The system stood up once and left to rot is the fragmented approach with extra steps, because unmaintained truth becomes untruth on a delay.
There is a meta-mistake that sits above all of these: treating the method as a project instead of a practice. A project has an end, and the moment you declare structured truth "done" is the moment it starts decaying, because the facts keep changing with or without your maintenance. The teams that win with this are the ones that internalize it as a standing habit, a source they maintain, a cadence they keep, a set of questions they ask the engines every month. The mistake is not any single omission on the list. It is believing you can stop, when the whole premise of a single source of truth is that keeping it single is the permanent job.
A worked example: one fact from source to citation
Abstraction is easy to nod along to and hard to act on, so let me walk a single fact through the entire loop, end to end, the way it actually happens. Take a mid-size maker of standing desks, and take one fact: the flagship desk's weight capacity. The engineering spec says 355 pounds. That number is going to travel a long way, and every step is a chance to keep it true or let it drift.
At Define, the team writes it down once: entity is the flagship desk, fact is weight capacity, value is 355 lb, source is the load-test report, owner is the head of product, volatility is low but not zero because a redesign could change it. Before this step, the number lived in three places with three values, because the old marketing site rounded it to 350, a sales sheet claimed 400 from a prototype, and the engineering doc said 355. Define kills the ambiguity: 355, sourced, owned. That single decision is the whole method in miniature.
At Mark Up, 355 goes into the Product schema as an additionalProperty, generated from the canonical sheet so it cannot diverge from the visible spec table on the page. It validates clean. Now a machine reading the product page does not infer the capacity from prose; it reads a field. At Answer, the team writes the FAQ the way customers actually ask it: "How much weight can the desk hold?" with a direct answer, "The flagship holds up to 355 lb, tested to load," short enough to lift and confident enough to quote, wrapped in FAQPage schema. At Expose, the product feed carries 355 to the marketplaces, the page is server-rendered and unpaywalled so crawlers read it, and llms.txt points models at the product taxonomy.
At Distribute, 355 propagates outward through the rings:
- The product page and its schema show 355.
- The product feed pushes 355 to every marketplace listing.
- The spec comparison on the category page reads 355.
- The retailer partners pulling the feed show 355.
- The FAQ answer, in schema, says 355.
Every surface now agrees, and the machines cross-referencing them resolve one confident fact instead of discounting a contradiction. A few weeks later, at Measure, someone asks Perplexity and ChatGPT "what is the weight capacity of the flagship standing desk" and both answer 355, citing the product page. That is the payoff: a canonical fact, defined once, is now the answer the discovery layer gives on the team's behalf, in a surface the team does not own and cannot edit directly.
Then the fact changes, and Refresh proves its worth. A hardware revision raises capacity to 400 lb. The head of product updates one cell on the canonical sheet. Because everything downstream generates from that cell, the schema, the feed, the FAQ, and the marketplace listings update together, and the distribution map confirms every ring moved. The dateModified bumps so crawlers see the freshness. At the next Measure, the AI answers have caught up to 400, because the source they trust changed at the source. Compare that to the fragmented world, where 355 would have become 400 on the homepage, stayed 355 in the schema, lingered at 350 on a cached marketplace listing, and left the answer engines quoting whichever stale copy they happened to trust. One fact, one source, one loop, and the machines tell your story correctly, everywhere, without you touching nineteen surfaces by hand.
What comes next for structured truth
A method should be forward-compatible, not just current, so let me end with where this is heading and why building it now is the cheap move. The direction of travel is clear even if the timing is not, and each shift makes structured truth more valuable, not less.
The biggest shift is agents that act instead of answer. Today most AI systems describe your options and hand the decision back to a human. Tomorrow more of them will take the action: comparing and buying the product, booking the appointment, completing the transaction on a user's behalf. When an agent is the one transacting, the brands it can actually operate on will win, and the brands whose prices, availability, and core actions are illegible to a machine will be skipped with no error and no appeal. Structured truth is exactly the readable, machine-completable substrate an agent needs. The teams that expose their real facts and actions now will be operable when agent commerce arrives; the teams that wait will be retrofitting under pressure.
The second shift is discovery going AI-first for a growing share of journeys. More searches will begin, and increasingly end, inside an AI answer rather than on a page of links. That does not delete the web underneath; it changes who gets seen, because the value migrates to being the source the AI trusts and cites. Everything in this method, sourced facts, clean schema, direct answers, consistent identity, is a bet on being that trusted source. A short list of the changes worth preparing for:
- Entity-first ranking. Engines increasingly resolve "what is true about this entity" before ranking documents, which rewards the brands with a clean, resolved entity and punishes the fragmented ones.
- Citation as the new click. Being quoted in an answer, sometimes without a click, becomes a primary outcome, which makes citation share a metric you cannot ignore.
- Provenance and freshness premiums. As models get pickier about what to trust, sourced and recently-updated facts get cited over anonymous, stale ones, so governance and refresh compound in value.
- Structured actions everywhere. The pressure to expose not just facts but completable actions grows as agents mature, extending the method from what is true to what can be done.
The reassuring part is that none of this asks you to chase a new tactic every quarter. The method is deliberately built on things that do not go out of style: know what is true, say it clearly, make it machine-readable, keep it consistent, keep it fresh. Those hold no matter the dominant surface, a search page, an AI answer, a voice assistant, or an agent doing your customer's shopping. Being easy for a machine to find, read, and trust is the durable advantage, and structured truth is how you earn it. Start with Define this week, get one fact clean from source to citation, and let the loop do the rest. If any of this was useful, the pillar it companions, the 19 ranking surfaces, maps the whole terrain this method is built to conquer.
Frequently asked questions
What is the Structured-Truth Method?
It is a way to run modern search optimization from one source instead of many. You maintain a single body of canonical facts, entities, answers, and data, mark it up so machines read it directly, then syndicate the same truth to every surface.
How is this different from regular SEO?
Regular SEO optimizes pages to rank in search results. Structured truth optimizes facts to be resolved and cited across every discovery surface: AI answers, voice, visual, marketplaces, and knowledge panels, not just blue links. SEO is one ring inside the method.
What exactly counts as structured truth?
Five things: entities (your organization, people, products, locations), canonical facts (the agreed value for each attribute), direct answers (clean quotable responses to real questions), schema markup (Schema.org JSON-LD wrapping it all), and original data (proprietary numbers only you can produce).
Do I need a big team or expensive tools to start?
No. A solo founder can run the whole loop with one well-structured spreadsheet, the free Google and Bing consoles, and a monthly hour of discipline. Tools reduce friction on work you have already decided to do; they do not make the decisions for you.
What are the steps of the method?
Seven moves in order: Define your canonical facts and entities, Mark them up in schema, publish direct Answers, Expose your truth to machines via llms.txt, feeds, and clean HTML, Distribute the same values to every surface, Measure citations and coverage, and Refresh before facts go stale.
How does this relate to the 19 ranking surfaces?
The 19 surfaces are the terrain; the Structured-Truth Method is the engine that makes optimizing all of them tractable. The pillar's five winning behaviors map directly onto the method's moves: schema and direct answers, machine exposure, consistent identity, and freshness.
How do I measure if it is working?
Measure citation and coverage, not just rankings. Ask ChatGPT, Claude, Gemini, and Perplexity your target questions on a fixed cadence and record which they answer with and how current the facts are. Track crawler hits in your logs, coverage per canonical fact, citation share versus competitors, and knowledge-panel health.
What is the most common way this fails?
Skipping Define and jumping straight to schema. Marking up facts you never actually agreed on just encodes your existing contradictions in a format machines read more reliably. The other frequent failure is treating the method as a one-time project.
What is llms.txt and do I need one?
llms.txt is a plain-language file at your domain root that describes what you do, your main URL patterns, and the actions available on your site, written for a language model rather than a shopper. It gives models a clean map instead of forcing them to reverse-engineer your sitemap.
How often should I refresh?
Two modes. Event-driven: whenever a canonical fact changes, update it at the source and let it propagate through your distribution map so nothing gets missed. Standing sweep: quarterly is a sane default, walk the canonical sheet, reconfirm sources, bump dateModified stamps, and rerun original studies that have aged.
Does structured truth help with AI shopping agents?
Yes, and it is where the method pays off most as agents mature. An agent that buys or books on a user's behalf can only operate on brands whose prices, availability, specs, and core actions are structured and machine-completable.
Is this only for big brands with a Wikidata entry?
No. The entity and knowledge-graph work scales with you: a small business starts with a canonical name, a consistent sameAs array, and clean Organization schema, and files a Wikidata entry when it can support one.
I'm Frederick Sona, and I've spent most of my career chasing one question: why do some brands break through while others, often the better ones, don't? I've looked for the answer as a marketer, a designer, a technologist, a salesperson, and a founder, and the honest answer is that it takes all of it: being easy to find, easy to trust, and easy to buy from. Search Everywhere Optimization is one piece of how I think about that, but this blog covers the whole picture, from search and technology to brand, design, and the work of turning attention into revenue. If any of this was useful, come say hello at fredericksona.com.