How to Create llms.txt: A Verification-First Procedure

Create and verify llms.txt with a root-level Markdown template, curl checks, link validation, deployment examples, logs, and measured limits on AI consumption claims.

Article highlights

  • Estimated reading time: 8 minutes
  • Published on: October 6, 2026
  • Last updated: October 6, 2026
RA
Published · Updated · 8 min read
Branded cover for the article 'How to Create llms.txt: A Verification-First Procedure': TrustGrowth wordmark and title with callouts: Inventory stable pages, then draft the file, Publish at the root URL and verify retrieval, Test consumption separately from availability.

Article

A published llms.txt proves only that a client can retrieve that URL. It does not prove that Google, an AI crawler, ChatGPT, Claude, or another answer engine fetched or used the file. This procedure creates a reproducible file, verifies its HTTP behavior, and separates measured observations from unknown downstream effects.

Prerequisites

Before starting, you need:

  • A publicly accessible HTTPS website with a configurable root path.
  • Deployment access to your static host, CDN, object-storage bucket, or web server.
  • A terminal with curl and sha256sum or an equivalent checksum utility.
  • Optional Python 3.11+ for link validation.
  • Access to CDN or web-server logs, if available.
  • Basic knowledge of DNS, HTTP status codes, redirects, Markdown, canonical URLs, and version control.
  • A fixed list of public URLs to evaluate and permission to request them.
  • If testing answer surfaces, the same AI provider or model before and after publication, a fixed prompt set, a recorded test date, and a way to save outputs.
  • An understanding that llms.txt is an emerging convention, not a verified control over search engines or LLM behavior.

The current proposal describes a root-level Markdown file containing a short site description and links to useful resources. Treat proposal language as a convention, not a vendor guarantee. The proposal is available at llmstxt.org (accessed 2026-10-05).

1. Define the scope

https://example.com/llms.txt is different from robots.txt, which communicates crawler access preferences; an XML sitemap, which lists URLs for discovery; schema markup, which annotates page content; and Google Search Console, which reports Google Search data. HTTP authentication and firewall rules remain the actual access controls.

Google Search Central's crawling and indexing documentation (accessed 2026-10-05) covers robots.txt, sitemaps, and meta tags; it does not list llms.txt or state that the file controls indexing or rankings. Google Search, Google-Extended, OpenAI, Anthropic, Perplexity, and other systems may ignore, fetch, or interpret the file differently. A successful request is therefore an availability measurement, not a consumption measurement.

Write this scope statement in your issue or repository: “We will publish a concise public resource index at /llms.txt, measure retrieval and linked-page validity, and treat downstream search or answer-engine effects as unknown unless a dated, repeatable experiment supplies evidence.”

Expected output: A written scope statement and a decision to create, test, or defer the file.

2. Inventory stable pages

Create a reviewed inventory of canonical, public URLs. Start with the homepage, product documentation, pricing, API reference, integration guides, comparison pages, and maintained policy pages. Prefer first-party pages with visible authorship, a review date, and a direct answer to a user question.

Exclude login-required pages, duplicate URL variants, tracking parameters, search-result URLs, thin campaign pages, unsupported claims, and obsolete documentation. Each URL should be intended for anonymous public access and should return a successful response.

Use a semantic Markdown list because the inventory is a set of resources. Record one-line purpose, canonical status, HTTP status, and last-reviewed date in version control or a spreadsheet. A canonical URL is the page a site declares as the preferred representative of a document, usually with a rel="canonical" element.

Expected output: A reviewed URL inventory with, for example, https://example.com/docs, its description, status 200, canonical destination, and review date 2025-02-14.

3. Draft the file

Keep the document short enough to maintain. Use absolute HTTPS URLs and descriptions that can be checked against visible page content. Do not add secrets, private URLs, prompt-injection text, instructions to ignore policies, or unsupported E-E-A-T claims. E-E-A-T—experience, expertise, authoritativeness, and trustworthiness—is an evaluation concept, not evidence you can assert without supporting pages.

Save this complete example as llms.txt in Git:

# Example Product

> Example Product helps technical teams monitor public API reliability.

## Start here
- [Product overview](https://example.com/product): Features and intended users.
- [Documentation](https://example.com/docs): Setup and usage instructions.

## Reference
- [API reference](https://example.com/api): Endpoints, authentication, and limits.
- [Pricing](https://example.com/pricing): Current plans and included usage.

## Optional
- [Changelog](https://example.com/changelog): Dated product changes.

Add a “last reviewed” date only when someone owns the review process. Do not duplicate the whole site or make promotional claims that cannot be verified on the linked pages.

Expected output: A version-controlled draft and a review checklist covering accuracy, public access, absolute URLs, stable destinations, secrets, and contradiction with visible content.

4. Publish the root URL

The required target is https://example.com/llms.txt, not /public/llms.txt or a framework source directory. Preserve the lowercase filename. Serve plain Markdown or text where possible, such as text/markdown; charset=utf-8 or text/plain; charset=utf-8.

For Next.js, place the file here:

my-next-app/
├── app/
├── public/
│   └── llms.txt
└── package.json

Next.js serves public/llms.txt at /llms.txt after deployment. For a generic static site, place llms.txt in the directory copied to the domain root:

site-build/
├── index.html
├── docs/
└── llms.txt

For Nginx serving static files, a site-specific location can be:

location = /llms.txt {
    try_files /llms.txt =404;
    default_type text/markdown;
}

These examples are stack-specific. Check your CDN, base path, authentication, bot challenge, and routing rules before release.

Expected output: The deployed file is available at the exact canonical root URL.

5. Verify retrieval and links

Run these commands against your domain. curl 8.4.0 or later is suitable; record your installed version with curl --version.

curl --version
curl -I https://example.com/llms.txt
curl -L --fail --silent --show-error https://example.com/llms.txt
curl -L --fail --silent --show-error https://example.com/llms.txt | sha256sum

Check for status 200, the final URL, redirect count, Content-Type, character encoding, readable Markdown, and a checksum matching the current deployment. The HTTP semantics used here are defined by RFC 9110 (accessed 2026-10-05). Save the UTC date, headers, command output, curl version, and deployment identifier.

You can validate the links with a simple Bash loop. Replace the URLs with the inventory you approved:

set -u
urls=(
  "https://example.com/product"
  "https://example.com/docs"
  "https://example.com/api"
  "https://example.com/pricing"
)
for url in "${urls[@]}"; do
  status=$(curl -L -o /dev/null -sS -w '%{http_code}' --max-time 20 "$url")
  final=$(curl -L -o /dev/null -sS -w '%{url_effective}' --max-time 20 "$url")
  printf '%s\t%s\t%s\n' "$status" "$url" "$final"
done

For a CSV report from a local Markdown file, save this as verify_links.py and run python3 verify_links.py llms.txt:

import csv
import re
import sys
from urllib.error import HTTPError, URLError
from urllib.request import Request, urlopen

path = sys.argv[1] if len(sys.argv) == 2 else "llms.txt"
text = open(path, encoding="utf-8").read()
links = re.findall(r"\[[^\]]+\]\((https://[^)\s]+)\)", text)
with open("llms-link-report.csv", "w", newline="", encoding="utf-8") as output:
    writer = csv.writer(output)
    writer.writerow(["requested_url", "final_url", "status", "error"])
    for url in links:
        try:
            request = Request(url, method="GET", headers={"User-Agent": "llms-link-validator/1.0"})
            with urlopen(request, timeout=20) as response:
                writer.writerow([url, response.geturl(), response.status, ""])
        except HTTPError as error:
            writer.writerow([url, error.geturl(), error.code, "HTTPError"])
        except (URLError, TimeoutError) as error:
            writer.writerow([url, "", "", type(error).__name__])
print(f"Checked {len(links)} links; wrote llms-link-report.csv")

Test from a second network or external HTTP checker if results vary. Record the date, UTC time, headers, final destination, and tool version.

Expected output: A verification log showing status, redirects, content type, checksum, body, and link results—or a precise failure.

6. Test consumption separately

A successful curl request proves only public retrieval. It does not prove that an AI system parsed the body or used it in an answer.

Review CDN or server logs for /llms.txt during a declared interval, such as 2025-02-14T00:00:00Z through 2025-02-21T23:59:59Z. Record request count, status, user-agent, source when available, and whether the request reached the origin. User-agent strings can be spoofed, so a request labeled as an AI crawler is an observation, not proof of ingestion.

For answer-surface testing, save a baseline before publication and a comparison after publication. Use the same provider, model identifier, locale, prompts, URLs, temperature settings where exposed, and test dates. Record citations, quoted claims, and omissions for every response. A practical table has columns for date, model, prompt ID, output file, cited URL, and observed change. Do not call a change causal without a controlled experiment and sufficient sample size.

Expected output: A dated evidence table plus a “what this does not prove” section. Include exact values such as n=20 prompts, model model-name, provider, URL set, and test dates. If those values have not been collected, use TBD, not an invented result.

7. Maintain and reassess

Assign an owner and review cadence, such as monthly or after every documentation, pricing, domain, or information-architecture change. Compare llms.txt with the XML sitemap, canonical tags, visible page content, and access policy.

Remove stale or contradictory links. Keep a change log containing date, editor, reason, checksum, and validation result. A file that conflicts with a visible pricing page is a data-quality defect, not a stronger source of truth.

Expected output: A maintenance record and a repeatable revalidation procedure.

Troubleshooting

  • 404 or wrong path: Confirm the deployment base path and that source public/llms.txt became root /llms.txt.
  • Redirect loop: Inspect CDN, HTTP-to-HTTPS, canonical-domain, and trailing-slash rules.
  • 401, 403, or bot challenge: Check WAF, authentication, robots policy, and edge rules. Do not weaken global security for this file.
  • Wrong content type or encoding: Inspect headers and body; ensure you did not receive an HTML error page.
  • Stale CDN content: Purge or revalidate the object, then compare SHA-256 checksums.
  • Broken or private links: Request every URL without credentials and record status and final destination.
  • Overlong file: Keep maintained, high-value resources; do not copy the site.
  • Misread logs: Treat user-agent and request counts as observations, not proof of AI use.
  • Assumed ranking or citation gains: Use a dated baseline and fixed comparison; otherwise report the effect as unknown.

Verification-first checklist and next steps

You created a version-controlled Markdown file, selected canonical public pages, deployed it at /llms.txt, checked HTTP behavior with curl, calculated a checksum, validated links, and separated retrieval evidence from consumption claims. Re-run the checks after each relevant site change and retain the logs.

If you want to know whether Google can access and understand the underlying site—not merely whether /llms.txt returns 200—run a TrustGrowth Google Search Console read-only audit. TrustGrowth, built by TechWright Labs, reports verified technical, performance, content, and inferred E-E-A-T signals, plus a dated site-specific proof snapshot. That audit can describe the audited site state; it does not prove that llms.txt caused rankings, citations, traffic, or LLM ingestion changes.

technical SEO Google Search Console AI visibility site audits llms.txt HTTP verification
Share:

Know your site's real SEO score

Free GSC-verified audit, E-E-A-T scoring, and AI-powered content strategy.

Get Started Free

Related Articles