Back to Web research and extraction

Firecrawl Crawl

Firecrawl

Read a bounded set of pages from a public website.

Use this tool with your AI agent

  1. Copy the prompt
  2. Paste it into your agent
  3. Tell it what you need

Read https://scrollport.com/start. Use my existing Scrollport connection, or help me connect if needed. Help me use Firecrawl Crawl (tool ID: firecrawl.site-crawl). Ask me for the details you need, then show me the result.

Already connected
Your agent can use that connection.
New user
Your agent will guide you through setup, sign-in and authorisation.

Provider access is included.

Connect your first agent to receive $1 trial credit.

Current price
$0.0025/page
Availability
Ready to run

Tool details

About this tool

Read up to 50 pages from one public website, with bounded depth and Markdown, HTML or links. Reserve the requested page limit, then charge only for successfully delivered pages and release unused funds. Partial delivery reports completed and failed page counts. The same durable job resumes after interruption. Subdomains, PDF documents, X/Twitter and JSON extraction are excluded. The worked example shows selected fields and at most one item per list; counts retain their recorded values.

Inputs and outputs

Input

formatsarray · optional

Value supplied for formats.

minItems: 1 · maxItems: 4

limitinteger · optional

Value supplied for limit.

minimum: 1 · maximum: 50

maxDiscoveryDepthinteger · optional

Value supplied for max discovery depth.

minimum: 0 · maximum: 5

onlyMainContentboolean · optional

Value supplied for only main content.

urlstring · required

Public HTTPS web page. PDF files and X/Twitter URLs are excluded.

format: uri · maxLength: 2048

Output

attempted_pagesnumber

Returned attempted pages value.

completed_pagesnumber

Returned completed pages value.

failed_pagesnumber

Returned failed pages value.

limitnumber

Returned limit value.

pagesarray

Returned pages value.

partialboolean

Returned partial value.

Example request and output

Inspect returns this reviewed example, its estimated cost and the current input schema. Run calculates the estimate for your actual input.

Request

{
  "tool_id": "firecrawl.site-crawl",
  "input": {
    "url": "https://books.toscrape.com/",
    "limit": 3,
    "formats": [
      "markdown"
    ],
    "onlyMainContent": true,
    "maxDiscoveryDepth": 1
  }
}

Output

{
  "limit": 3,
  "pages": [
    {
      "url": "https://books.toscrape.com/",
      "title": "\n    All products | Books to Scrape - Sandbox\n",
      "status": 200,
      "markdown": "- [Home](https://books.toscrape.com/index.html)\n- All products\n\n# All products\n\n**1000** results - showing **1** to **20**.\n\n\n\n\n\n\n**Warning!** This is a demo website for web scraping purposes. Prices and ratings here were randomly assigned and have no real meaning.\n\n01. [![A Light in the Attic](https://books.toscrape.com/media/cache/2c/da/2cdad67c44b002e7ead0cc35693c0e8b.jpg)](https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html)\n\n\n\n\n\n\n\n    ### [A Light in the ...](https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html \"A Light in the Attic\")\n\n\n\n\n\n    £51.77\n\n\n\n\n\n    In stock\n\n\n\n    Add to basket\n\n02. [![Tipping the Velvet](https://books.toscrape.com/m",
      "truncated": false
    }
  ],
  "partial": false,
  "failed_pages": 0,
  "attempted_pages": 3,
  "completed_pages": 3
}
Errors and limitations
  • invalid_input: The URL, query or limits are outside this tool's supported inputs.example: {"url":"https://example.com/","limit":1,"maxDiscoveryDepth":1,"formats":["markdown"]}
  • upstream_error: The provider could not return usable results; the wallet was not charged.Check the target page and retry with a smaller limit if necessary.