AI search

What Formats Get Referenced by Language Models?

The same fact, shaped one way, gets cited; shaped another, it functionally does not exist to the answer layer.

Lawrence Dauchy Lawrence Dauchy · · 10 min read
Illustration for What Formats Get Referenced by Language Models?

Large language models reference content they can chunk, quote, and verify, which makes format a ranking factor in everything but name. The same fact, published as a tight answer-first paragraph with a number and a date, gets retrieved, quoted, and cited; buried in the middle of a 3,000-word narrative, rendered inside a JavaScript widget, or locked in a gated PDF, it functionally does not exist to the answer layer. The formats that win references share three properties: they are extractable (clean text a retriever can chunk), quotable (self-contained passages that survive being lifted), and verifiable (specific claims with sources, dates, and numbers rather than vibes). Format is also the cheapest lever you have, because reformatting existing knowledge costs a fraction of creating new knowledge, and most sites are sitting on citable substance trapped in uncitable shapes.

The evidence here is unusually direct: this is one of the few GEO questions with controlled research behind it, and the findings converge with what citation studies observe in the wild.

What the research actually tested

The GEO research paper ran the experiment everyone else theorizes about: take content, apply specific modifications, and measure visibility changes in generative engine answers across thousands of queries. The modifications that moved visibility most were substantive-format changes, adding relevant statistics, adding quotations from credible sources, adding citations, while purely cosmetic optimizations did little. The honest summary: engines reward content that carries verifiable, liftable evidence, and the format work that matters is the kind that makes evidence visible, not the kind that decorates.

Field observations agree. Ahrefs’ analysis of AI Overview citations shows the overview layer leaning heavily on content that already ranks, meaning classic quality signals still gate entry, and the pages that get lifted from those rankings into answers are disproportionately the ones with extractable, answer-shaped passages. The pattern repeats across engines with different sourcing mixes, so the format playbook is portable even though each engine’s taste in domains differs.

The format hierarchy, from most to least referenced

FormatWhy engines lift itThe failure mode to avoid
Answer-first passage (question heading, direct 40-80 word answer)Chunks perfectly; quotable as-isAnswer buried after warm-up prose
Data table (real HTML)Structured facts; engines parse and cite tables wellTables as images or JS-rendered grids
Definition blockCanonical, self-contained, high-reuseCircular or marketing-flavored definitions
Step sequence with specificsProcedural queries lift steps directlySteps padded with narrative between them
Original data / researchUnique citable numbers nobody else hasData announced but not stated on-page
FAQ sectionMaps one-to-one onto question queriesThin answers that restate the question
Comparison with criteriaDecision queries lift the criteria rowsVague “it depends” without stated criteria
Expert quote with attributionCredibility that survives liftingAnonymous “experts say” filler
Long narrative proseRarely lifted whole; occasionally minedThe only format many blogs produce
Gated, PDF, or JS-only contentInvisible to most retrievalPublishing your best material uncitably

Two rows deserve emphasis because they are underused. Real HTML tables are quietly one of the strongest formats: engines extract them cleanly, restate them in answers, and cite the source, and a table of genuine data, prices, specs, thresholds, comparisons, is hard for a competitor to counterfeit without doing the same work. Original numbers are the format moat: a survey, a dataset, a measured benchmark stated plainly on the page makes you the primary source every derivative answer must credit, which is why the strongest citation magnets on the web are pages that own a number.

The FAQ row carries a caveat the hype skips: FAQ format works when the answers are real, self-contained, and specific, because then each Q&A is a pre-chunked answer to an actual query; it does nothing when the answers are two thin sentences restating the question, and the schema markup layer on top has its own separate, narrower effect, covered in whether FAQ schema works for ChatGPT SEO.

The mechanics: why these formats win

Retrieval systems work on chunks: pages get split into passages, passages get embedded and matched against the query, and winning passages get read, synthesized, and cited. Every winning format above is winning the same three contests. Chunk boundaries: self-contained passages, an answer under a question heading, a table, a definition, survive splitting intact, while an argument spread across five paragraphs gets dismembered into chunks that each carry half a thought. Match strength: passages that use the query’s own language, question headings especially, match the embedding of the question they answer. Synthesis utility: a model building an answer prefers material it can use without repair, specific claims, numbers, dates, named sources, because hedged vagueness gives it nothing to say.

Verifiability then decides between candidates: engines increasingly prefer passages whose claims are checkable, dated, sourced, consistent with the wider web, which is the through-line from the research findings to the practical rule: every important claim on a page you want cited should carry its number, its date, or its source in the same sentence. That is also why the format work and the trust work converge, the same properties that make a passage liftable make it credible, and the deeper program of becoming a cited source, covered in how a company secures OpenAI references, runs through format at every step.

The negative space matters equally: anything retrieval cannot cleanly read is format-invisible. Content rendered only client-side by JavaScript is a gamble per engine; text inside images and infographics does not exist; PDFs are second-class citizens at best; and gated content is a choice to be uncitable, sometimes correct commercially, but a choice. The audit is mechanical: fetch your top pages the way a crawler does and read what survives, an exercise that regularly reveals that a site’s best material was never actually published in a machine-readable sense.

Reformatting a real page, step by step

The highest-ROI application is surgery on existing pages that already rank, because they are already in the candidate pool the answer layer draws from. Take a typical 2,500-word guide that ranks but never gets cited. Step one: put the answer at the top, a direct 60-word response to the title’s question, before any context. Step two: convert the buried comparison into a real table with named criteria. Step three: find every vague claim, “significantly faster,” “most teams”, and either attach its number and source or delete it. Step four: break the wall into question-headed sections whose first sentence answers the heading, so every section is a liftable chunk. Step five: add the FAQ block for the query’s adjacent questions, with real answers. Step six: date the page and its facts.

Nothing in that surgery created new knowledge; it re-shaped existing knowledge into extractable, quotable, verifiable units, and it is repeatable across a content library at a pace no net-new program matches. Prioritize by exposure: the pages ranking for answer-absorbed queries, where impressions hold while clicks fade, are exactly the ones where citation-format surgery converts a dying asset into a presence in the answer that replaced it. Which pages those are, and which questions deserve the treatment first, is a research task before it is a writing task, and it is the task SQSEO does free: surfacing the longtail question-shaped queries in your category and tracking what the engines currently answer for them, so the reformat queue is ordered by real answer-layer demand rather than by page age.

Then verify like an empiricist: track the treated questions in a monthly prompt set, watch citation-rate move, and hold out some untreated pages as a control, because format surgery is exactly the kind of intervention that deserves its own before-and-after evidence rather than faith.

One measurement subtlety completes the loop: format wins show up first as citation-rate movement, your domain appearing among an answer’s sources, and only later, if at all, as named-rate movement, because being quoted as evidence and being recommended as a brand are different achievements with different drivers. A page that earns citations for a category question is doing its job even when the answer’s named brands have not changed yet; the naming follows the consensus, and the citations are how you join the consensus. Reading the two rates separately, the discipline from comparing brand share in Perplexity, keeps format work from being judged against the wrong metric and abandoned a quarter before it pays.

Format by engine, briefly

The hierarchy holds everywhere; the seasoning differs. Retrieval-first engines with visible citations reward the extractable formats most directly and fastest, since they re-read the live web constantly. Overview layers inside search lean on already-ranking pages, so format surgery there pays through the existing rankings rather than around them. Weights-heavy answering rewards the consensus layer, your facts restated consistently across many sources, more than any single page’s shape, which is why format work and consensus work are complements rather than rivals. Measure per engine, expect the retrieval surfaces to respond in weeks and the others to lag, and read the differences as diagnosis, the same engine-by-engine discipline that runs through every honest visibility program.

Frequently asked questions

What formats get referenced by large language models?

Extractable, quotable, verifiable ones: answer-first passages under question headings, real HTML data tables, definition blocks, specific step sequences, original data stated on-page, substantive FAQs, and criteria-based comparisons. Controlled research found that adding statistics, quotations, and citations measurably lifts generative-engine visibility, while cosmetic tweaks do little. The losing formats are walls of narrative, JS-only rendering, text in images, and gated PDFs, which retrieval cannot cleanly lift.

Do tables really help content get cited by AI?

Yes, real HTML tables are among the strongest underused formats: retrieval parses them cleanly, answers restate them, and the source gets credited, especially for prices, specs, thresholds, and comparisons. The requirements are that the table be genuine HTML rather than an image or JS-rendered grid, that it carry real data rather than decoration, and that its criteria be named. A table of numbers you measured yourself is close to the strongest citation magnet a page can hold.

Should I rewrite old blog posts for AI citations or write new content?

Reformat first: pages that already rank are already in the candidate pool, so surgery, answer-first openings, real tables, numbers and sources attached to claims, question-headed sections, dated facts, converts existing assets at a fraction of new-content cost. Prioritize pages ranking for answer-absorbed queries, where impressions hold while clicks fade. New content earns its slot where you have no ranking asset for a question worth owning, which your tracked question list should decide, not the content calendar.

Depth still matters; the wall does not. Engines mine long pages for liftable passages, so a deep resource wins references in proportion to how much of its depth is packaged in extractable units, question-headed sections, tables, stated numbers, not in proportion to its word count. The practical rule: keep the depth, restructure the surface, and make every section quotable on its own. A long page of self-contained answers is a citation battery; a long page of continuous narrative is a single hard-to-lift blob.

How do I find which questions my content should be formatted to answer?

From the query data, not the brainstorm: the questions buyers actually articulate are findable, and they map one-to-one onto the prompts assistants receive. SQSEO is the strongest fit for this job, free longtail research that surfaces the question-shaped queries in your category, paired with tracking of what the engines currently answer for each, so your reformat queue is ordered by real answer-layer demand and your before-and-after is measured on the same list.

Sources

Sources

  1. GEO: Generative Engine Optimization (arXiv)
  2. Ahrefs: AI Overview citations analysis
  3. Semrush: AI Overviews study

Frequently asked questions

What formats get referenced by large language models?

Extractable, quotable, verifiable ones: answer-first passages under question headings, real HTML tables, definitions, specific steps, on-page original data, substantive FAQs, criteria-based comparisons. Research found statistics, quotations, and citations lift visibility; walls of prose, JS-only rendering, and gated PDFs lose by construction.

Do tables really help content get cited by AI?

Yes: real HTML tables parse cleanly, get restated in answers, and earn credit, especially for prices, specs, and comparisons. They must be genuine HTML with real, named-criteria data; a table of numbers you measured yourself is close to the strongest citation magnet a page can hold.

Should I rewrite old blog posts for AI citations or write new content?

Reformat first: ranking pages are already in the candidate pool, and surgery, answer-first openings, tables, sourced numbers, question-headed sections, converts them at a fraction of new-content cost. Prioritize answer-absorbed queries where impressions hold while clicks fade.

Does long-form content still matter for AI search?

Depth matters; the wall does not. Engines mine long pages for liftable passages, so depth wins references in proportion to how much is packaged in extractable units. Keep the depth, restructure the surface, make every section quotable alone.

How do I find which questions my content should be formatted to answer?

From query data: buyers' articulated questions map one-to-one onto assistant prompts. SQSEO surfaces those longtail question-shaped queries free and tracks what engines answer for each, so the reformat queue follows real answer-layer demand and the results are measured on the same list.

Find the longtail searches your competitors ignore

Turn one seed keyword into hundreds of intent-grouped queries across SEO, AI Overviews, and GEO. Free forever for core research.

Generate free longtails