AptlyStar

fastCRW

Scrape, search, crawl, and map web data

使用说明

Integrate fastCRW into the workflow. Scrape pages, search the web, crawl entire sites, and map URL structures. fastCRW is a Firecrawl-compatible web scraper in a single binary — self-host or cloud.

工具

crw_scrape

Extract structured content from web pages with comprehensive metadata support. Converts content to markdown or HTML while capturing SEO metadata, Open Graph tags, and page information.

输入

参数类型必填描述
urlstring是The URL to scrape content from (e.g., "https://example.com/page"\)
scrapeOptionsjson否Options for content scraping
baseUrlstring否Base URL for self-hosted fastCRW (defaults to https://fastcrw.com/api\)
apiKeystring是fastCRW API key
pricingper_request否No description
rateLimitstring否No description

输出

参数类型描述
markdownstringPage content in markdown format
htmlstringRaw HTML content of the page
metadataobjectPage metadata including SEO and Open Graph information
↳ titlestringPage title
↳ descriptionstringPage meta description
↳ languagestringPage language code (e.g., "en")
↳ sourceURLstringOriginal source URL that was scraped
↳ statusCodenumberHTTP status code of the response
↳ keywordsstringPage meta keywords
↳ robotsstringRobots meta directive (e.g., "follow, index")
↳ ogTitlestringOpen Graph title
↳ ogDescriptionstringOpen Graph description
↳ ogUrlstringOpen Graph URL
↳ ogImagestringOpen Graph image URL
↳ ogLocaleAlternatearrayAlternate locale versions for Open Graph
↳ ogSiteNamestringOpen Graph site name
↳ errorstringError message if scrape failed

Search for information on the web using fastCRW

输入

参数类型必填描述
querystring是The search query to use
baseUrlstring否Base URL for self-hosted fastCRW (defaults to https://fastcrw.com/api\)
apiKeystring是fastCRW API key
pricingper_request否No description
rateLimitstring否No description

输出

参数类型描述
dataarraySearch results data with scraped content and metadata
↳ titlestringSearch result title from search engine
↳ descriptionstringSearch result description/snippet from search engine
↳ urlstringURL of the search result
↳ markdownstringPage content in markdown (when sources include scraped content)
↳ metadataobjectMetadata about the search result page
↳ titlestringPage title
↳ descriptionstringPage meta description
↳ sourceURLstringOriginal source URL
↳ statusCodenumberHTTP status code
↳ errorstringError message if scrape failed

crw_crawl

Crawl entire websites and extract structured content from all accessible pages

输入

参数类型必填描述
urlstring是The website URL to crawl (e.g., "https://example.com" or "https://docs.example.com/guide"\)
maxPagesnumber否Maximum number of pages to crawl (e.g., 50, 100, 500). Default: 100
maxDepthnumber否Maximum depth to crawl from the starting URL (e.g., 1, 2, 3). Controls how many levels deep to follow links
formatsjson否Output formats for scraped content (e.g., ["markdown"], ["markdown", "html"], ["markdown", "links"])
excludePathsjson否URL paths to exclude from crawling (e.g., ["/blog/", "/admin/", "/*.pdf"])
includePathsjson否URL paths to include in crawling (e.g., ["/docs/", "/api/"]). Only these paths will be crawled
onlyMainContentboolean否Extract only main content from pages
baseUrlstring否Base URL for self-hosted fastCRW (defaults to https://fastcrw.com/api\)
apiKeystring是fastCRW API key
pricingper_request否No description
rateLimitstring否No description

输出

参数类型描述
pagesarrayArray of crawled pages with their content and metadata
↳ markdownstringPage content in markdown format
↳ htmlstringProcessed HTML content of the page
↳ rawHtmlstringUnprocessed raw HTML content
↳ linksarrayArray of links found on the page
↳ metadataobjectPage metadata from crawl operation
↳ titlestringPage title
↳ descriptionstringPage meta description
↳ languagestringPage language code
↳ sourceURLstringOriginal source URL
↳ statusCodenumberHTTP status code
↳ ogLocaleAlternatearrayAlternate locale versions
totalnumberTotal number of pages found during crawl

crw_map

Get a complete list of URLs from any website quickly and reliably. Useful for discovering all pages on a site without crawling them.

输入

参数类型必填描述
urlstring是The base URL to map and discover links from (e.g., "https://example.com"\)
limitnumber否Maximum number of links to return (e.g., 100, 1000, 5000)
baseUrlstring否Base URL for self-hosted fastCRW (defaults to https://fastcrw.com/api\)
apiKeystring是fastCRW API key
pricingper_request否No description
rateLimitstring否No description

输出

参数类型描述
successbooleanWhether the mapping operation was successful
linksarrayArray of discovered URLs from the website

On this page