scraping.fetchHtml
Pages and structure — fetchHtml. Give it url and it answers with html, status, finalUrl, contentType.
What it takes
scraping.fetchHtml takes url. The page to fetch. Required unless html is supplied.
What it gives back
The answer carries html, status, finalUrl, contentType. Each field is declared by the adapter rather than described here, so what this page says it returns and what the call actually returns cannot drift apart.
The same verb everywhere
Running this from the page, posting scraping.fetchHtml to /api/netinfo/run, and giving it to a coding agent over MCP are three doors onto one implementation. There is no separate website version that behaves differently from the documented one, which is the whole reason the verb has the same name in all three places.
- html (string) — The raw markup. A non-2xx status THROWS (status on the error), so this is never an error page body.
- finalUrl (string) — Where the request ended up after redirects — compare with your input to detect one.
- contentType (string) — The response's content-type, or '' when the server sent none.
Questions
Do you store what I look up?
A lookup is answered and not kept against you. A scan is kept — the account, the target, the time and the authorisation statement — because that record is what an abuse report is answered with, and keeping it is the condition of offering the feature at all.
Do you respect robots.txt when scraping?
Yes. The scraper identifies itself honestly, honours robots.txt, holds itself to a couple of requests a second per host, and will not touch anything behind a login. A scraper that ignores those is a scraper that gets our address blocked, which would end the product.