Skip to content

Agent-Readable Content: Balanced Guidance

Agent-Readable Content: Balanced Canonical Guidance

Section titled “Agent-Readable Content: Balanced Canonical Guidance”
  • Tease: Make content easy for agents to use without turning it into content written for machines.
  • Lede: The durable strategy is original, people-first, crawlable content with clear structure and verifiable evidence. Add crawler controls, Markdown, llms.txt, structured data, or agent protocols only for a real platform or user need; none is a universal ranking recipe.
  • Why it matters:
    • Local practitioner research identifies useful extraction and citation patterns, but Google explicitly warns against AI-specific rewrites, tiny chunks, query-variant page farms, and unsupported generative-search hacks.
    • OpenAI and Anthropic document access controls and citation behavior, but neither publishes a content-format formula that guarantees selection in answers.
    • Over-applying tactics such as question-only headings can make a site repetitive, less useful to people, and vulnerable to the broader search-first patterns Google says it may downrank or exclude.
  • Go deeper: Use the decision table first, then the provider-specific guidance, heading policy, audit order, and ranked source map.

Date: 2026-08-01

Status: Canonical entry point for Agent Ready content recommendations. The linked memos remain evidence for narrower questions, but recommendations should defer to this document when they conflict.

Scope: Public web content intended to be discovered, retrieved, summarized, cited, recommended, or used by Google Search, ChatGPT, Claude, and browser agents. This is not a promise of placement and not a guide to model training optimization.

Do the work that survives platform changes:

  1. Publish one accurate, useful, canonical source for each coherent user need.
  2. Make the source crawlable, indexable where appropriate, accessible, and easy to navigate.
  3. Use descriptive headings chosen for the content: questions for real questions, imperative headings for tasks, and noun or outcome phrases for concepts.
  4. Put the useful answer, definition, comparison, procedure, or evidence near the heading that promises it.
  5. Support claims with original experience, methods, dates, named sources, and accurate citations.
  6. Separate search crawlers, user-directed fetchers, and training crawlers instead of applying one “AI bot” policy.
  7. Treat Markdown, llms.txt, schema, and agent protocols as optional interfaces with explicit consumers, not as universal ranking signals.
  8. Measure eligibility, citations, factual fidelity, qualified visits, and conversions separately.

The core rule is:

Optimize the source for the person it serves, then remove unnecessary friction for the systems that retrieve and use it.

This topic mixes unlike kinds of evidence. Use these labels when turning a source into advice:

LabelWhat it can establishWhat it cannot establish
Official platform guidanceA provider’s stated requirements, controls, policies, and product behaviorA universal rule for other platforms or a placement guarantee
Standard or specificationA documented interface and expected semanticsAdoption, ranking impact, or current client support
Controlled experimentA measured effect under a defined protocolDurable behavior outside the tested models, prompts, corpus, or period
Observed field evidenceWhat real crawlers, logs, or users did in one environmentWhy a ranking system acted or whether another site will see the same result
Practitioner guidanceUseful hypotheses, workflows, and failure patternsProvider policy or causal proof without independent validation
Local conventionA good default for this workspace or productA public-web ranking factor

No scanner, vendor checklist, paper, or anecdote should silently move from one row to another.

“What agents prefer” is too broad. Different systems discover and use content through different pipelines.

StageReal questionPractices that can helpPractices that do not prove success
Access and eligibilityCan the system fetch or index the page?Useful public HTML, correct status, internal links, sitemap, crawler policy, snippet/index controlsA good passage cannot overcome a blocked or ineligible page
SelectionWill the platform retrieve this source for this user and query?Relevance, authority, originality, freshness where needed, platform-access policyllms.txt, schema, Markdown, or question headings do not guarantee selection
Extraction and citationCan the system identify and accurately use the relevant evidence?Clear sections, direct passages, stable terminology, dates, sources, definitions, comparisons, proceduresExtractability does not prove ranking or referral traffic
ActionCan a user-directed agent safely complete a task?Semantic controls, labels, stable state, accessible forms, real APIs/tools, authentication, confirmationsStatic metadata does not make a nonexistent or unsafe capability real

Question headings mainly affect navigation and extraction. They are not a demonstrated shortcut through access or selection.

PracticeDefaultEvidence and counterbalance
Crawlable, canonical, useful HTMLDoRequired for Google eligibility and broadly useful to other retrieval systems.
Original experience, analysis, data, or a genuinely useful synthesisDoGoogle’s current guidance prioritizes unique, non-commodity, people-first value.
Clear paragraphs, sections, and descriptive headingsDoGoogle recommends structure that helps readers; accessibility and agent navigation benefit too.
Direct answers beneath relevant headingsUsually doHelps people and passage extraction. Keep context and nuance; no fixed length is required.
Every heading phrased as a questionDo not defaultNo official provider requires it. Use questions only when they are the clearest label for a real user question.
One page for every query, “People also ask” item, or fan-out variationAvoidGoogle explicitly warns that query-variation page creation for manipulation can violate scaled-content policy.
Tiny “AI-friendly” chunks or an ideal word countAvoid as a ruleGoogle says chunking is not required and there is no ideal page length. Use coherent sections sized for the subject and reader.
Accurate authorship, dates, methodology, and source linksDo when relevantImproves human trust and verification; aligns with Google’s people-first guidance and citation-oriented answer products.
Structured data matching visible contentUse when eligible and usefulNormal Search features may benefit; Google says no special schema is required for generative AI search.
FAQ contentUse for genuine recurring questionsA helpful content type, not a site-wide template. Do not invent questions solely to target queries.
FAQPage structured dataUse only when valid and maintainedGoogle generally limits FAQ rich results to authoritative government and health sites; markup is not an AI-ranking lever.
llms.txtOptionalUseful as a curated map for clients or workflows that intentionally read it. Google says it ignores the file; field adoption remains uneven.
Page-level Markdown or Accept: text/markdownOptional, strongest for docsCan reduce parsing noise and tokens in direct-fetch workflows. Field logs do not show universal crawler use or ranking benefit.
AGENTS.md, MCP, API Catalog, Agent Cards, or WebMCPOnly for a real consumer or capabilityThese are capability or workflow interfaces, not generic SEO. Never publish fake surfaces for a score.
AI-generated pages at scaleAvoid unless each page has real user value and quality controlGoogle treats scaled, low-value production for ranking or generative-response manipulation as spam regardless of how it was generated.
Repeated prompt checks and log reviewDo as measurementUseful evidence of changing behavior; neither reveals the complete retrieval chain nor guarantees causality.

Google now provides a direct counterweight to speculative GEO advice.

  • Continue normal SEO because generative Search features use Google’s Search index, ranking, and quality systems.
  • Create original, non-commodity, useful content based on real knowledge or experience.
  • Organize content for readers with paragraphs, sections, and headings that create a clear navigation structure.
  • Keep pages crawlable, technically sound, and eligible for snippets; provide a good page experience.
  • Use structured data where it supports an existing Search feature and matches visible content.
  • Measure Google’s generative-search visibility in Search Console rather than treating third-party “internal metrics” as official.

Google’s July 2026 generative AI guide says website owners do not need:

  • llms.txt, AI text files, special markup, or Markdown for Google Search visibility;
  • tiny content chunks for AI understanding;
  • content rewritten in a special style for AI systems;
  • every long-tail or fan-out query variation represented verbatim;
  • special Schema.org markup for generative Search;
  • inauthentic mentions or third-party tools promising ranking success.

It also warns against creating separate pages for every possible query or fan-out variation when the purpose is to manipulate rankings or generative answers. Google’s spam policies say sites using scaled content abuse may rank lower or disappear from results.

No official Google Search source found in this pass says a page is downranked merely because its headings use questions. The safer conclusion is narrower:

  • Google wants headings that clearly help readers navigate.
  • Google says not to rewrite content into a special AI-targeted style.
  • Google warns against multiplying content around query and fan-out variations.
  • Google’s developer documentation style guide—a writing guide, not a ranking policy—uses imperative headings for tasks and noun phrases for concepts rather than converting every section into a question.

The likely long-term risk is not the question mark. It is a site-wide pattern that makes the content look search-first, repetitive, mass-produced, or less satisfying to the visitor.

OpenAI’s publisher and crawler documentation establishes access controls, not a copywriting formula.

  • Allow OAI-SearchBot and its published IP ranges when inclusion in ChatGPT search is desired.
  • Treat GPTBot separately; it controls potential foundation-model training use, not ChatGPT Search inclusion.
  • Do not treat ChatGPT-User as the Search crawler; it represents some user-initiated actions and may not follow robots.txt in the same way.
  • Use noindex when a crawlable page should not appear through search products; OpenAI notes the crawler must access the page to read that tag.
  • Track ChatGPT referrals using the utm_source=chatgpt.com parameter.
  • For ChatGPT Agent in Atlas, improve accessibility with accurate roles, labels, and states for interactive elements.
  • Expect no guaranteed top placement; OpenAI says ranking uses multiple factors intended to return reliable, relevant information.

OpenAI does not currently publish an official requirement for:

  • question-only headings;
  • a minimum number of headings;
  • llms.txt or Markdown mirrors for ChatGPT Search ranking;
  • FAQ schema or a special AI schema;
  • a target paragraph, passage, or page length;
  • a formula that guarantees citation or recommendation.

ChatGPT Search may rewrite a user’s request into multiple targeted queries and may use third-party search providers. That supports testing topic coverage and source relevance, but it does not justify building a page for every imagined rewritten query.

Anthropic also documents access lanes and citation behavior without publishing a site-format ranking recipe.

  • Claude-SearchBot supports search indexing and result quality.
  • Claude-User retrieves content for some user-directed requests.
  • ClaudeBot is the model-development crawler.
  • Anthropic says its bots honor robots.txt, respect anti-circumvention controls, and support Crawl-delay where appropriate.
  • Claude web search processes multiple sources and produces a conversational response with source links and citations.
  • Private content should be password-protected; Anthropic also documents noindex, bot controls, and removal requests for excluding content from web-search outputs.

Anthropic does not currently publish an official site-owner rule requiring:

  • question headings or a Q&A template;
  • llms.txt, Markdown, or special Schema.org markup;
  • a preferred passage or page length;
  • a content pattern that guarantees a Claude citation.

Because Claude presents citations, clear evidence and source links can improve human verification and may improve extraction fidelity. Treat the citation-selection benefit as a reasonable hypothesis to measure, not as Anthropic ranking guidance.

The practitioner layer is useful when its scope stays visible.

The local Agent Ready corpus, Cloudflare’s documentation practice, Vercel’s checklist, and the 2024 GEO paper converge on several low-regret practices:

  • make the canonical page easy to fetch and understand;
  • use stable terminology and coherent section boundaries;
  • include concrete facts, definitions, comparisons, procedures, and sources when they help the reader;
  • keep evidence close to the claim it supports;
  • preserve authorship, method, and date context when they affect trust or freshness;
  • measure actual outputs rather than trusting a scanner score.

These are strongest when they improve the human page at the same time.

  • The GEO paper measured visibility effects in a defined generative-engine benchmark. It did not establish Google ranking factors, universal organic discovery, durable referrals, or business conversions.
  • Vercel’s Agent Readability specification is a practical vendor checklist. Its fixed thresholds—such as heading counts or text ratios—are audit heuristics, not official OpenAI, Anthropic, or Google ranking rules.
  • Cloudflare uses llms.txt, Markdown endpoints, and content negotiation for its own documentation workflows, while also saying most content work aligns with normal SEO and content quality.
  • Dries Buytaert’s March 2026 log analysis found no crawler using content negotiation, very limited llms.txt use in the sample, and Markdown discovery through explicit alternate links. This is strong field evidence against calling those features universal discovery channels.
  • Hacker News discussion is consistently skeptical of AEO claims without live citation or log evidence. That skepticism is a useful quality gate, not a substitute for measurement.

The repo’s docs-spec experiment tested 14 documentation variants across token cost, blind human review, and build fidelity. In that bounded context:

  • the bundle of all “best practices” scored below baseline and used 35% more tokens;
  • outcome headings were neutral enough to cut;
  • heading lint was marginal;
  • a Motivation/Explanation page was the one clear improvement.

This was an internal specification eval, not a public Search ranking study. Its relevance is methodological: do not assume that individually plausible writing tactics compound into a better document. Test the bundle and keep the human review axis.

Sometimes. Use a question heading when all of these are true:

  • it is a question the intended reader genuinely asks;
  • the section answers that exact question;
  • the question is the clearest label in the page’s scan path;
  • it does not duplicate another heading or exist only to capture a query variant.

Use a task heading when the section helps someone do something:

  • Configure crawler access
  • Compare the source evidence
  • Measure citation fidelity

Use a noun or outcome phrase when the section explains a concept or result:

  • Crawler policy
  • Citation measurement
  • The safe default

Do not mechanically turn a natural heading such as Crawler policy into What is our crawler policy? merely to look extractable. Mixed, descriptive headings are the default.

After a heading:

  1. Orient or answer in the first paragraph.
  2. Add the evidence, conditions, and exceptions needed to make the answer true.
  3. Keep one main purpose per section, but do not fracture a coherent explanation into tiny pieces.
  4. Link to the primary source or canonical evidence near consequential claims.

There is no required answer length. A definition may need one sentence; a safety or financial claim may need substantial context.

Prefer one strong page when several questions share the same audience, decision, evidence, and maintenance owner. Split a page when the visitor needs a meaningfully different artifact, workflow, or canonical URL—not simply because a query tool produced another phrasing.

A healthy topic cluster usually has:

  • one canonical overview or decision page;
  • separate pages for genuinely distinct procedures, references, products, locations, or evidence sets;
  • stable internal links that explain the relationship;
  • no near-duplicate pages competing to answer the same need.

Avoid:

  • page-per-question generation;
  • city, role, or product variants with only token substitutions;
  • invented FAQs;
  • multiple pages that summarize the same external sources without original value;
  • date changes that imply freshness without a substantive update.

Agent-facing formats are justified when a real consumer benefits and maintenance can stay synchronized.

  • llms.txt as a curated map a known agent, support workflow, or documentation tool reads;
  • Markdown pages for developers or users who intentionally copy or fetch clean text;
  • OpenAPI, MCP, API Catalog, or Agent Cards for working, tested capabilities;
  • accessible names, roles, states, and stable controls for browser agents and people using assistive technology;
  • generated machine views derived from the same canonical source.
  • claiming the file or protocol guarantees Google, ChatGPT, or Claude placement;
  • maintaining a second, drifting set of claims for machines;
  • exposing private, licensed, or sensitive content to improve a score;
  • publishing fake actions, tools, auth flows, or commerce capabilities;
  • serving materially different content to crawlers for ranking manipulation.

The safe architecture is one maintained source of truth with honest derived representations.

Measure each stage separately.

  1. Eligibility: indexed status, snippet eligibility, response status, canonical, robots policy, bot access, and server-log evidence.
  2. Selection: repeated prompt panels by platform, location/account condition, prompt cluster, and date.
  3. Citation fidelity: whether the source is cited, which passage supports the answer, and whether the generated claim matches the page.
  4. Referral quality: Google Search Console generative AI reporting, ChatGPT referral parameters, analytics, qualified visits, and assisted conversions.
  5. Action success: completion rate, error recovery, confirmations, accessibility-tree clarity, and safe handling of destructive or authenticated steps.

Record the prompt, platform, model or product surface when visible, account/location condition, date, cited URLs, answer claim, and outcome. A single successful answer is an example, not a trend.

Use this order so speculative improvements cannot displace fundamentals:

  1. Confirm the page serves a real intended audience and a coherent need.
  2. Verify the facts, provenance, authorship, date, method, and maintenance owner.
  3. Check the canonical URL, crawl/index controls, response, sitemap/internal discovery, and visible main content.
  4. Review the human scan path, heading hierarchy, accessibility, page experience, and mobile rendering.
  5. Test extraction: definitions, comparisons, steps, evidence, caveats, and source links.
  6. Decide search, user-fetch, and training crawler policy separately.
  7. Add only the structured data and agent interfaces that have a real consumer.
  8. Run repeated platform checks and compare changes against the baseline.
  9. Review business and user outcomes before scaling the pattern.

Stop or redesign when a recommendation would require:

  • inventing data, quotes, dates, authors, reviews, or expertise;
  • mass-producing pages for query or fan-out variations;
  • rewriting every heading into a question without a reader benefit;
  • hiding context so a passage is easier to lift but easier to misstate;
  • exposing confidential or private information;
  • adding structured data that does not match visible content;
  • publishing a protocol surface with no working capability behind it;
  • promising a ranking, citation, recommendation, or traffic outcome;
  • treating a crawler hit, scanner score, or one answer as proof of causality.

Use these public pages for the working evidence and implementation guidance behind the balanced rules:

Google:

OpenAI:

Anthropic: