Data APIs / URL to markdown

URL to Markdown API

Turn any URL — articles, docs, product pages — or an uploaded PDF, Word, Excel, or image file into clean markdown that drops straight into a prompt or a vector store. JavaScript-rendered pages included.

From 1 credit per call · cache hits free · failed calls refunded · no card for the free tier

One request, structured data and markdown

Every endpoint accepts format=json, markdown, or both. Use the JSON to filter and store; drop the markdown straight into a prompt.

  • Article-quality markdown with headings, lists, tables, links, and images preserved
  • Title and description metadata
  • Files: PDF, Word, Excel, CSV, OpenDocument, Numbers, and images (OCR)
  • Two endpoints: GET for URLs, POST multipart for files — same response shape
request
curl 'https://scrapewhale.dev/api/v1/markdown/url?url=https://example.com/article'   -H 'Authorization: Bearer <api-key>'
response · json
{
  "id": "9djpmHuoC3SGImkcGm7Nl",
  "provider": "jina",
  "title": "Example Domain",
  "markdown": "# Example Domain\n\nThis domain is for use in illustrative examples…",
  "contentChars": 210,
  "creditsCharged": 1
}

Endpoints

2 endpoints in this family. Full parameter reference and a try-it console are in the API docs.

get/api/v1/markdown/url1 credit

URL to markdown

Any public web page as clean, LLM-ready markdown — headings, lists, tables, links, and images preserved; nav, ads, and boilerplate removed. JavaScript-rendered pages included. Plain page reading only: YouTube, X, Facebook, and other structured sources have their own endpoints and are not auto-detected here. Recently read pages are served from cache for free.

url *
The web page to convert.
post/api/v1/markdown/file1 credit

File to markdown

Upload a document and get it back as markdown: PDF, Word, Excel, CSV, OpenDocument, Apple Numbers, and images (scanned PDFs and images are OCR'd). multipart/form-data with a file field, 10 MB cap; HTML files are rejected — use the URL endpoint for web pages. The upload is validated before any credit is charged.

What teams use it for

RAG ingestion

Feed clean pages and PDFs to your embedding pipeline without HTML cleanup.

Agent browsing

Give an agent readable pages at a fraction of the tokens raw HTML costs.

Content monitoring

Diff competitor pages as markdown, not DOM soup.

Frequently asked questions