Scrape, search, crawl, and map web data
Integrate fastCRW into the workflow. Scrape pages, search the web, crawl entire sites, and map URL structures. fastCRW is a Firecrawl-compatible web scraper in a single binary — self-host or cloud.
Extract structured content from web pages with comprehensive metadata support. Converts content to markdown or HTML while capturing SEO metadata, Open Graph tags, and page information.
| 参数 | 类型 | 必填 | 描述 |
|---|
url | string | 是 | The URL to scrape content from (e.g., "https://example.com/page"\) |
scrapeOptions | json | 否 | Options for content scraping |
baseUrl | string | 否 | Base URL for self-hosted fastCRW (defaults to https://fastcrw.com/api\) |
apiKey | string | 是 | fastCRW API key |
pricing | per_request | 否 | No description |
rateLimit | string | 否 | No description |
| 参数 | 类型 | 描述 |
|---|
markdown | string | Page content in markdown format |
html | string | Raw HTML content of the page |
metadata | object | Page metadata including SEO and Open Graph information |
↳ title | string | Page title |
↳ description | string | Page meta description |
↳ language | string | Page language code (e.g., "en") |
↳ sourceURL | string | Original source URL that was scraped |
↳ statusCode | number | HTTP status code of the response |
↳ keywords | string | Page meta keywords |
↳ robots | string | Robots meta directive (e.g., "follow, index") |
↳ ogTitle | string | Open Graph title |
↳ ogDescription | string | Open Graph description |
↳ ogUrl | string | Open Graph URL |
↳ ogImage | string | Open Graph image URL |
↳ ogLocaleAlternate | array | Alternate locale versions for Open Graph |
↳ ogSiteName | string | Open Graph site name |
↳ error | string | Error message if scrape failed |
Search for information on the web using fastCRW
| 参数 | 类型 | 必填 | 描述 |
|---|
query | string | 是 | The search query to use |
baseUrl | string | 否 | Base URL for self-hosted fastCRW (defaults to https://fastcrw.com/api\) |
apiKey | string | 是 | fastCRW API key |
pricing | per_request | 否 | No description |
rateLimit | string | 否 | No description |
| 参数 | 类型 | 描述 |
|---|
data | array | Search results data with scraped content and metadata |
↳ title | string | Search result title from search engine |
↳ description | string | Search result description/snippet from search engine |
↳ url | string | URL of the search result |
↳ markdown | string | Page content in markdown (when sources include scraped content) |
↳ metadata | object | Metadata about the search result page |
↳ title | string | Page title |
↳ description | string | Page meta description |
↳ sourceURL | string | Original source URL |
↳ statusCode | number | HTTP status code |
↳ error | string | Error message if scrape failed |
Crawl entire websites and extract structured content from all accessible pages
| 参数 | 类型 | 必填 | 描述 |
|---|
url | string | 是 | The website URL to crawl (e.g., "https://example.com" or "https://docs.example.com/guide"\) |
maxPages | number | 否 | Maximum number of pages to crawl (e.g., 50, 100, 500). Default: 100 |
maxDepth | number | 否 | Maximum depth to crawl from the starting URL (e.g., 1, 2, 3). Controls how many levels deep to follow links |
formats | json | 否 | Output formats for scraped content (e.g., ["markdown"], ["markdown", "html"], ["markdown", "links"]) |
excludePaths | json | 否 | URL paths to exclude from crawling (e.g., ["/blog/", "/admin/", "/*.pdf"]) |
includePaths | json | 否 | URL paths to include in crawling (e.g., ["/docs/", "/api/"]). Only these paths will be crawled |
onlyMainContent | boolean | 否 | Extract only main content from pages |
baseUrl | string | 否 | Base URL for self-hosted fastCRW (defaults to https://fastcrw.com/api\) |
apiKey | string | 是 | fastCRW API key |
pricing | per_request | 否 | No description |
rateLimit | string | 否 | No description |
| 参数 | 类型 | 描述 |
|---|
pages | array | Array of crawled pages with their content and metadata |
↳ markdown | string | Page content in markdown format |
↳ html | string | Processed HTML content of the page |
↳ rawHtml | string | Unprocessed raw HTML content |
↳ links | array | Array of links found on the page |
↳ metadata | object | Page metadata from crawl operation |
↳ title | string | Page title |
↳ description | string | Page meta description |
↳ language | string | Page language code |
↳ sourceURL | string | Original source URL |
↳ statusCode | number | HTTP status code |
↳ ogLocaleAlternate | array | Alternate locale versions |
total | number | Total number of pages found during crawl |
Get a complete list of URLs from any website quickly and reliably. Useful for discovering all pages on a site without crawling them.
| 参数 | 类型 | 必填 | 描述 |
|---|
url | string | 是 | The base URL to map and discover links from (e.g., "https://example.com"\) |
limit | number | 否 | Maximum number of links to return (e.g., 100, 1000, 5000) |
baseUrl | string | 否 | Base URL for self-hosted fastCRW (defaults to https://fastcrw.com/api\) |
apiKey | string | 是 | fastCRW API key |
pricing | per_request | 否 | No description |
rateLimit | string | 否 | No description |
| 参数 | 类型 | 描述 |
|---|
success | boolean | Whether the mapping operation was successful |
links | array | Array of discovered URLs from the website |