Solutions · Data extraction
Read pages, tables and files behind your login.
Get a page as Markdown without opening a tab, pull tables and links as JSON, and read downloads (PDF, XLSX, DOCX, CSV) as text or rows.
curl -s -X POST https://api.webpilot.si/v1/read \
-H "Authorization: Bearer $WEBPILOT_TOKEN" -H "content-type: application/json" \
-d '{"url":"https://en.wikipedia.org/wiki/Web_browser","format":"markdown"}'
# → main content as Markdown; "via": "http" or "browser"How it works
01
Fast path first, your browser when needed
browser_read and POST /v1/read fetch the page on the server and return its main content as Markdown or text. YouTube watch pages return their caption transcript, and documents go through the same extractors as downloads.
If the site answers 401, 403 or 429, shows a bot challenge or has no text without JavaScript, the read is done again in your own browser, in a read tab that is closed afterwards. use_browser: true goes there straight away, for pages behind your login.
02
Downloads of any type
Files the browser downloads are kept per user. Read one as extracted text (PDF, XLSX, DOCX, CSV, HTML), as rows, or as the file; stream it over HTTP; or make a short-lived link for a person.
03
Safe by default
Read tabs cannot submit anything. The server-side fetch only connects to public addresses (private and loopback addresses are refused after DNS and on every redirect), and page text reaches the agent fenced as untrusted data.
FAQ
Questions people ask
Does reading use browser time?
The fast path does not start your browser. Only the fallback, or use_browser: true, opens a read tab in it, and browser time counts while the browser runs.
Can it read pages behind my login?
Yes. Log in once (from the vault or by hand in the live view) and the cookies stay in your browser's profile. Use use_browser: true, or open a read tab, to read as yourself.
Which document types can it read?
PDF, DOCX, XLSX and CSV as text or tables, HTML as Markdown, and YouTube transcripts. Other downloads come back as files.
Can I get structured output?
browser_extract returns tables (rows keyed by header), links, meta (title, meta tags, JSON-LD, headings) and form fields as JSON. For a schema of your own, the REST tasks endpoint validates its output against an output_schema you send.
Is there a size limit?
The server-side fetch reads up to 10 MB and gives up after 20 seconds by default; larger files can still be downloaded through the browser.
Will it follow instructions it finds on a page?
The tools return page text fenced and labelled as untrusted data, and our agent instructions say never to follow it. That is a safeguard, not a guarantee; keep sensitive sites read-only in your site rules.
Give your agent a browser that remembers.
Start with the free trial. Connect Claude Code, Cursor, Codex or your own code in a minute.