Data APIs / Agent-ready check

llms.txt & agents.md Checker API

Probe a domain for the files agents look for — llms.txt, llms-full.txt, agents.md, and the emerging UCP manifest — and get back what exists, its size, and its contents.

From 1 credit per call · cache hits free · failed calls refunded · no card for the free tier

One request, structured data and markdown

Every endpoint accepts format=json, markdown, or both. Use the JSON to filter and store; drop the markdown straight into a prompt.

  • Presence, status, and size of llms.txt, llms-full.txt, agents.md, and UCP files
  • The contents of each file that exists
  • A single agent-readiness verdict per domain
  • Pair with URL-to-markdown to read the pages those files point to
request
curl 'https://scrapewhale.dev/api/v1/agent-ready?domain=anthropic.com&format=json' \
  -H 'Authorization: Bearer <api-key>'
response · json
{
  "extractor": "agent-ready",
  "creditsCharged": 1,
  "data": {
    "domain": "anthropic.com",
    "files": {
      "llms.txt": { "found": true, "status": 200, "bytes": 5210, "content": "# Anthropic\n\n> …" },
      "llms-full.txt": { "found": false, "status": 404 },
      "agents.md": { "found": false, "status": 404 }
    }
  }
}

Abridged.

Endpoints

2 endpoints in this family. Full parameter reference and a try-it console are in the API docs.

get/api/v1/agent-ready1 credit

Agent-readiness probe

Whether a site publishes the agent-facing files AI agents look for: /llms.txt, /agents.md, and the /.well-known/ucp Universal Commerce Protocol merchant profile (with its versions, capabilities, payment handlers, and MCP endpoint when advertised). Follows redirects; HTML soft-404 pages count as absent. Shopify stores publish all three by default.

domain *
The site to probe. Accepts a bare domain or a full URL.
get/api/v1/markdown/url1 credit

URL to markdown

Any public web page as clean, LLM-ready markdown — headings, lists, tables, links, and images preserved; nav, ads, and boilerplate removed. JavaScript-rendered pages included. Plain page reading only: YouTube, X, Facebook, and other structured sources have their own endpoints and are not auto-detected here. Recently read pages are served from cache for free.

url *
The web page to convert.

What teams use it for

Agent-readiness audits

Score a portfolio of client sites on whether agents can discover them.

Agent tooling

Let your agent check for llms.txt before crawling a site the hard way.

Adoption research

Measure how many sites in a category publish llms.txt yet.

Frequently asked questions