Agent-Readable Content: Balanced Guidance
Agent-Readable Content: Balanced Canonical Guidance
Section titled “Agent-Readable Content: Balanced Canonical Guidance”Tease:Make content easy for agents to use without turning it into content written for machines.Lede:The durable strategy is original, people-first, crawlable content with clear structure and verifiable evidence. Add crawler controls, Markdown,llms.txt, structured data, or agent protocols only for a real platform or user need; none is a universal ranking recipe.Why it matters:- Local practitioner research identifies useful extraction and citation patterns, but Google explicitly warns against AI-specific rewrites, tiny chunks, query-variant page farms, and unsupported generative-search hacks.
- OpenAI and Anthropic document access controls and citation behavior, but neither publishes a content-format formula that guarantees selection in answers.
- Over-applying tactics such as question-only headings can make a site repetitive, less useful to people, and vulnerable to the broader search-first patterns Google says it may downrank or exclude.
Go deeper:Use the decision table first, then the provider-specific guidance, heading policy, audit order, and ranked source map.
Date: 2026-08-01
Status: Canonical entry point for Agent Ready content recommendations. The linked memos remain evidence for narrower questions, but recommendations should defer to this document when they conflict.
Scope: Public web content intended to be discovered, retrieved, summarized, cited, recommended, or used by Google Search, ChatGPT, Claude, and browser agents. This is not a promise of placement and not a guide to model training optimization.
Short answer
Section titled “Short answer”Do the work that survives platform changes:
- Publish one accurate, useful, canonical source for each coherent user need.
- Make the source crawlable, indexable where appropriate, accessible, and easy to navigate.
- Use descriptive headings chosen for the content: questions for real questions, imperative headings for tasks, and noun or outcome phrases for concepts.
- Put the useful answer, definition, comparison, procedure, or evidence near the heading that promises it.
- Support claims with original experience, methods, dates, named sources, and accurate citations.
- Separate search crawlers, user-directed fetchers, and training crawlers instead of applying one “AI bot” policy.
- Treat Markdown,
llms.txt, schema, and agent protocols as optional interfaces with explicit consumers, not as universal ranking signals. - Measure eligibility, citations, factual fidelity, qualified visits, and conversions separately.
The core rule is:
Optimize the source for the person it serves, then remove unnecessary friction for the systems that retrieve and use it.
Evidence labels
Section titled “Evidence labels”This topic mixes unlike kinds of evidence. Use these labels when turning a source into advice:
| Label | What it can establish | What it cannot establish |
|---|---|---|
| Official platform guidance | A provider’s stated requirements, controls, policies, and product behavior | A universal rule for other platforms or a placement guarantee |
| Standard or specification | A documented interface and expected semantics | Adoption, ranking impact, or current client support |
| Controlled experiment | A measured effect under a defined protocol | Durable behavior outside the tested models, prompts, corpus, or period |
| Observed field evidence | What real crawlers, logs, or users did in one environment | Why a ranking system acted or whether another site will see the same result |
| Practitioner guidance | Useful hypotheses, workflows, and failure patterns | Provider policy or causal proof without independent validation |
| Local convention | A good default for this workspace or product | A public-web ranking factor |
No scanner, vendor checklist, paper, or anecdote should silently move from one row to another.
The four stages that tactics affect
Section titled “The four stages that tactics affect”“What agents prefer” is too broad. Different systems discover and use content through different pipelines.
| Stage | Real question | Practices that can help | Practices that do not prove success |
|---|---|---|---|
| Access and eligibility | Can the system fetch or index the page? | Useful public HTML, correct status, internal links, sitemap, crawler policy, snippet/index controls | A good passage cannot overcome a blocked or ineligible page |
| Selection | Will the platform retrieve this source for this user and query? | Relevance, authority, originality, freshness where needed, platform-access policy | llms.txt, schema, Markdown, or question headings do not guarantee selection |
| Extraction and citation | Can the system identify and accurately use the relevant evidence? | Clear sections, direct passages, stable terminology, dates, sources, definitions, comparisons, procedures | Extractability does not prove ranking or referral traffic |
| Action | Can a user-directed agent safely complete a task? | Semantic controls, labels, stable state, accessible forms, real APIs/tools, authentication, confirmations | Static metadata does not make a nonexistent or unsafe capability real |
Question headings mainly affect navigation and extraction. They are not a demonstrated shortcut through access or selection.
Canonical decision table
Section titled “Canonical decision table”| Practice | Default | Evidence and counterbalance |
|---|---|---|
| Crawlable, canonical, useful HTML | Do | Required for Google eligibility and broadly useful to other retrieval systems. |
| Original experience, analysis, data, or a genuinely useful synthesis | Do | Google’s current guidance prioritizes unique, non-commodity, people-first value. |
| Clear paragraphs, sections, and descriptive headings | Do | Google recommends structure that helps readers; accessibility and agent navigation benefit too. |
| Direct answers beneath relevant headings | Usually do | Helps people and passage extraction. Keep context and nuance; no fixed length is required. |
| Every heading phrased as a question | Do not default | No official provider requires it. Use questions only when they are the clearest label for a real user question. |
| One page for every query, “People also ask” item, or fan-out variation | Avoid | Google explicitly warns that query-variation page creation for manipulation can violate scaled-content policy. |
| Tiny “AI-friendly” chunks or an ideal word count | Avoid as a rule | Google says chunking is not required and there is no ideal page length. Use coherent sections sized for the subject and reader. |
| Accurate authorship, dates, methodology, and source links | Do when relevant | Improves human trust and verification; aligns with Google’s people-first guidance and citation-oriented answer products. |
| Structured data matching visible content | Use when eligible and useful | Normal Search features may benefit; Google says no special schema is required for generative AI search. |
| FAQ content | Use for genuine recurring questions | A helpful content type, not a site-wide template. Do not invent questions solely to target queries. |
FAQPage structured data | Use only when valid and maintained | Google generally limits FAQ rich results to authoritative government and health sites; markup is not an AI-ranking lever. |
llms.txt | Optional | Useful as a curated map for clients or workflows that intentionally read it. Google says it ignores the file; field adoption remains uneven. |
Page-level Markdown or Accept: text/markdown | Optional, strongest for docs | Can reduce parsing noise and tokens in direct-fetch workflows. Field logs do not show universal crawler use or ranking benefit. |
AGENTS.md, MCP, API Catalog, Agent Cards, or WebMCP | Only for a real consumer or capability | These are capability or workflow interfaces, not generic SEO. Never publish fake surfaces for a score. |
| AI-generated pages at scale | Avoid unless each page has real user value and quality control | Google treats scaled, low-value production for ranking or generative-response manipulation as spam regardless of how it was generated. |
| Repeated prompt checks and log review | Do as measurement | Useful evidence of changing behavior; neither reveals the complete retrieval chain nor guarantees causality. |
Official Google guidance
Section titled “Official Google guidance”Google now provides a direct counterweight to speculative GEO advice.
What Google recommends
Section titled “What Google recommends”- Continue normal SEO because generative Search features use Google’s Search index, ranking, and quality systems.
- Create original, non-commodity, useful content based on real knowledge or experience.
- Organize content for readers with paragraphs, sections, and headings that create a clear navigation structure.
- Keep pages crawlable, technically sound, and eligible for snippets; provide a good page experience.
- Use structured data where it supports an existing Search feature and matches visible content.
- Measure Google’s generative-search visibility in Search Console rather than treating third-party “internal metrics” as official.
What Google says not to overdo
Section titled “What Google says not to overdo”Google’s July 2026 generative AI guide says website owners do not need:
llms.txt, AI text files, special markup, or Markdown for Google Search visibility;- tiny content chunks for AI understanding;
- content rewritten in a special style for AI systems;
- every long-tail or fan-out query variation represented verbatim;
- special Schema.org markup for generative Search;
- inauthentic mentions or third-party tools promising ranking success.
It also warns against creating separate pages for every possible query or fan-out variation when the purpose is to manipulate rankings or generative answers. Google’s spam policies say sites using scaled content abuse may rank lower or disappear from results.
What this means for the heading question
Section titled “What this means for the heading question”No official Google Search source found in this pass says a page is downranked merely because its headings use questions. The safer conclusion is narrower:
- Google wants headings that clearly help readers navigate.
- Google says not to rewrite content into a special AI-targeted style.
- Google warns against multiplying content around query and fan-out variations.
- Google’s developer documentation style guide—a writing guide, not a ranking policy—uses imperative headings for tasks and noun phrases for concepts rather than converting every section into a question.
The likely long-term risk is not the question mark. It is a site-wide pattern that makes the content look search-first, repetitive, mass-produced, or less satisfying to the visitor.
Official OpenAI guidance
Section titled “Official OpenAI guidance”OpenAI’s publisher and crawler documentation establishes access controls, not a copywriting formula.
Confirmed guidance
Section titled “Confirmed guidance”- Allow
OAI-SearchBotand its published IP ranges when inclusion in ChatGPT search is desired. - Treat
GPTBotseparately; it controls potential foundation-model training use, not ChatGPT Search inclusion. - Do not treat
ChatGPT-Useras the Search crawler; it represents some user-initiated actions and may not followrobots.txtin the same way. - Use
noindexwhen a crawlable page should not appear through search products; OpenAI notes the crawler must access the page to read that tag. - Track ChatGPT referrals using the
utm_source=chatgpt.comparameter. - For ChatGPT Agent in Atlas, improve accessibility with accurate roles, labels, and states for interactive elements.
- Expect no guaranteed top placement; OpenAI says ranking uses multiple factors intended to return reliable, relevant information.
Unsupported extrapolations
Section titled “Unsupported extrapolations”OpenAI does not currently publish an official requirement for:
- question-only headings;
- a minimum number of headings;
llms.txtor Markdown mirrors for ChatGPT Search ranking;- FAQ schema or a special AI schema;
- a target paragraph, passage, or page length;
- a formula that guarantees citation or recommendation.
ChatGPT Search may rewrite a user’s request into multiple targeted queries and may use third-party search providers. That supports testing topic coverage and source relevance, but it does not justify building a page for every imagined rewritten query.
Official Anthropic guidance
Section titled “Official Anthropic guidance”Anthropic also documents access lanes and citation behavior without publishing a site-format ranking recipe.
Confirmed guidance
Section titled “Confirmed guidance”Claude-SearchBotsupports search indexing and result quality.Claude-Userretrieves content for some user-directed requests.ClaudeBotis the model-development crawler.- Anthropic says its bots honor
robots.txt, respect anti-circumvention controls, and supportCrawl-delaywhere appropriate. - Claude web search processes multiple sources and produces a conversational response with source links and citations.
- Private content should be password-protected; Anthropic also documents
noindex, bot controls, and removal requests for excluding content from web-search outputs.
Unsupported extrapolations
Section titled “Unsupported extrapolations”Anthropic does not currently publish an official site-owner rule requiring:
- question headings or a Q&A template;
llms.txt, Markdown, or special Schema.org markup;- a preferred passage or page length;
- a content pattern that guarantees a Claude citation.
Because Claude presents citations, clear evidence and source links can improve human verification and may improve extraction fidelity. Treat the citation-selection benefit as a reasonable hypothesis to measure, not as Anthropic ranking guidance.
Practitioner and research evidence
Section titled “Practitioner and research evidence”The practitioner layer is useful when its scope stays visible.
Durable patterns
Section titled “Durable patterns”The local Agent Ready corpus, Cloudflare’s documentation practice, Vercel’s checklist, and the 2024 GEO paper converge on several low-regret practices:
- make the canonical page easy to fetch and understand;
- use stable terminology and coherent section boundaries;
- include concrete facts, definitions, comparisons, procedures, and sources when they help the reader;
- keep evidence close to the claim it supports;
- preserve authorship, method, and date context when they affect trust or freshness;
- measure actual outputs rather than trusting a scanner score.
These are strongest when they improve the human page at the same time.
Necessary caveats
Section titled “Necessary caveats”- The GEO paper measured visibility effects in a defined generative-engine benchmark. It did not establish Google ranking factors, universal organic discovery, durable referrals, or business conversions.
- Vercel’s Agent Readability specification is a practical vendor checklist. Its fixed thresholds—such as heading counts or text ratios—are audit heuristics, not official OpenAI, Anthropic, or Google ranking rules.
- Cloudflare uses
llms.txt, Markdown endpoints, and content negotiation for its own documentation workflows, while also saying most content work aligns with normal SEO and content quality. - Dries Buytaert’s March 2026 log analysis found no crawler using content negotiation, very limited
llms.txtuse in the sample, and Markdown discovery through explicit alternate links. This is strong field evidence against calling those features universal discovery channels. - Hacker News discussion is consistently skeptical of AEO claims without live citation or log evidence. That skepticism is a useful quality gate, not a substitute for measurement.
Local eval counterbalance
Section titled “Local eval counterbalance”The repo’s docs-spec experiment tested 14 documentation variants across token cost, blind human review, and build fidelity. In that bounded context:
- the bundle of all “best practices” scored below baseline and used 35% more tokens;
- outcome headings were neutral enough to cut;
- heading lint was marginal;
- a Motivation/Explanation page was the one clear improvement.
This was an internal specification eval, not a public Search ranking study. Its relevance is methodological: do not assume that individually plausible writing tactics compound into a better document. Test the bundle and keep the human review axis.
Heading policy
Section titled “Heading policy”Should headings be questions?
Section titled “Should headings be questions?”Sometimes. Use a question heading when all of these are true:
- it is a question the intended reader genuinely asks;
- the section answers that exact question;
- the question is the clearest label in the page’s scan path;
- it does not duplicate another heading or exist only to capture a query variant.
Use a task heading when the section helps someone do something:
Configure crawler accessCompare the source evidenceMeasure citation fidelity
Use a noun or outcome phrase when the section explains a concept or result:
Crawler policyCitation measurementThe safe default
Do not mechanically turn a natural heading such as Crawler policy into What is our crawler policy? merely to look extractable. Mixed, descriptive headings are the default.
Section-writing rule
Section titled “Section-writing rule”After a heading:
- Orient or answer in the first paragraph.
- Add the evidence, conditions, and exceptions needed to make the answer true.
- Keep one main purpose per section, but do not fracture a coherent explanation into tiny pieces.
- Link to the primary source or canonical evidence near consequential claims.
There is no required answer length. A definition may need one sentence; a safety or financial claim may need substantial context.
Content architecture policy
Section titled “Content architecture policy”Prefer one strong page when several questions share the same audience, decision, evidence, and maintenance owner. Split a page when the visitor needs a meaningfully different artifact, workflow, or canonical URL—not simply because a query tool produced another phrasing.
A healthy topic cluster usually has:
- one canonical overview or decision page;
- separate pages for genuinely distinct procedures, references, products, locations, or evidence sets;
- stable internal links that explain the relationship;
- no near-duplicate pages competing to answer the same need.
Avoid:
- page-per-question generation;
- city, role, or product variants with only token substitutions;
- invented FAQs;
- multiple pages that summarize the same external sources without original value;
- date changes that imply freshness without a substantive update.
Agent-only surfaces
Section titled “Agent-only surfaces”Agent-facing formats are justified when a real consumer benefits and maintenance can stay synchronized.
Good uses
Section titled “Good uses”llms.txtas a curated map a known agent, support workflow, or documentation tool reads;- Markdown pages for developers or users who intentionally copy or fetch clean text;
- OpenAPI, MCP, API Catalog, or Agent Cards for working, tested capabilities;
- accessible names, roles, states, and stable controls for browser agents and people using assistive technology;
- generated machine views derived from the same canonical source.
Bad uses
Section titled “Bad uses”- claiming the file or protocol guarantees Google, ChatGPT, or Claude placement;
- maintaining a second, drifting set of claims for machines;
- exposing private, licensed, or sensitive content to improve a score;
- publishing fake actions, tools, auth flows, or commerce capabilities;
- serving materially different content to crawlers for ranking manipulation.
The safe architecture is one maintained source of truth with honest derived representations.
Measurement plan
Section titled “Measurement plan”Measure each stage separately.
- Eligibility: indexed status, snippet eligibility, response status, canonical, robots policy, bot access, and server-log evidence.
- Selection: repeated prompt panels by platform, location/account condition, prompt cluster, and date.
- Citation fidelity: whether the source is cited, which passage supports the answer, and whether the generated claim matches the page.
- Referral quality: Google Search Console generative AI reporting, ChatGPT referral parameters, analytics, qualified visits, and assisted conversions.
- Action success: completion rate, error recovery, confirmations, accessibility-tree clarity, and safe handling of destructive or authenticated steps.
Record the prompt, platform, model or product surface when visible, account/location condition, date, cited URLs, answer claim, and outcome. A single successful answer is an example, not a trend.
Audit order
Section titled “Audit order”Use this order so speculative improvements cannot displace fundamentals:
- Confirm the page serves a real intended audience and a coherent need.
- Verify the facts, provenance, authorship, date, method, and maintenance owner.
- Check the canonical URL, crawl/index controls, response, sitemap/internal discovery, and visible main content.
- Review the human scan path, heading hierarchy, accessibility, page experience, and mobile rendering.
- Test extraction: definitions, comparisons, steps, evidence, caveats, and source links.
- Decide search, user-fetch, and training crawler policy separately.
- Add only the structured data and agent interfaces that have a real consumer.
- Run repeated platform checks and compare changes against the baseline.
- Review business and user outcomes before scaling the pattern.
Stop conditions
Section titled “Stop conditions”Stop or redesign when a recommendation would require:
- inventing data, quotes, dates, authors, reviews, or expertise;
- mass-producing pages for query or fan-out variations;
- rewriting every heading into a question without a reader benefit;
- hiding context so a passage is easier to lift but easier to misstate;
- exposing confidential or private information;
- adding structured data that does not match visible content;
- publishing a protocol surface with no working capability behind it;
- promising a ranking, citation, recommendation, or traffic outcome;
- treating a crawler hit, scanner score, or one answer as proof of causality.
Companion reference trail
Section titled “Companion reference trail”Use these public pages for the working evidence and implementation guidance behind the balanced rules:
- Practical Agent Readiness Audit Priority — the weighted implementation sequence and scanner counterbalance.
- Agentic Empathy — fetch, provenance, citation, and action-boundary guidance.
- WordPress Agent-Ready Tooling — applicability-aware implementation guidance for WordPress sites.
- Docs-Spec evaluation findings — the local evaluation showing that bundling plausible writing tactics can add tokens while reducing human review quality.
- AgentReady resources — the broader official, standards, practitioner, tooling, and safety source map.
Official source links
Section titled “Official source links”Google:
- Optimizing your website for generative AI features on Google Search — current official generative-search guidance and mythbusting; accessed 2026-08-01.
- Creating helpful, reliable, people-first content — content quality, authorship, method, trust, and search-first warning signs.
- Spam policies for Google web search — scaled content abuse, doorway abuse, cloaking, and possible ranking/exclusion consequences.
- Google guidance on generative AI content — accuracy, quality, relevance, disclosure context, and scaled-content limits.
- Google developer documentation heading style — descriptive imperative/noun-phrase heading guidance; a writing reference, not a Search ranking policy.
- Changes to FAQ and How-To rich results — current limits on FAQ rich-result visibility.
- Build agent-friendly websites — DOM, accessibility-tree, semantic control, stable-layout, and action guidance from web.dev.
OpenAI:
- Publishers and Developers FAQ — SearchBot access,
noindex, referral measurement, and Atlas accessibility guidance. - Overview of OpenAI crawlers — independent Search, training, and user-action lanes.
- ChatGPT Search — query rewriting, citations, inclusion controls, and no guaranteed top placement.
Anthropic:
- Anthropic crawler controls —
ClaudeBot,Claude-User,Claude-SearchBot,robots.txt, and crawl policy. - Enabling and using web search — multi-source search, citations, direct fetch, and verification guidance.
- Reporting, blocking, and removing content from Claude — privacy, authentication,
noindex, bot, and removal controls.
Practitioner and research source links
Section titled “Practitioner and research source links”- Dries Buytaert: Markdown, llms.txt and AI crawlers — one month of field logs and a strong adoption counterweight.
- Cloudflare AI consumability — a large documentation operator’s current Markdown and curation practice.
- Vercel Agent Readability specification — implementable checklist; treat fixed thresholds as vendor heuristics.
llms.txtproposal — proposed interface, not a universal standard or ranking signal.- Aggarwal et al., “GEO: Generative Engine Optimization” — KDD 2024 controlled research; interpret within its benchmark scope.
- Hacker News: Answer Engine Optimization — concise practitioner skepticism about advice with no proof of citation impact.