AptlyStar

Context.dev

Scrape, crawl, search, extract, and enrich web and brand data

使用说明

Integrate Context.dev into the workflow. Scrape pages to markdown or HTML, capture screenshots, list images, crawl entire sites, map sitemaps, search the web, extract structured data and products, pull design systems, classify industries, and retrieve brand assets by domain, name, email, ticker, or transaction ΓÇö all from one API.

工具

context_dev_scrape_markdown

Scrape any URL and return clean, LLM-ready markdown content.

输入

参数类型必填描述
urlstring是The full URL to scrape (must include http:// or https://)
useMainContentOnlyboolean否Return only main content, excluding headers, footers, and navigation
includeLinksboolean否Preserve hyperlinks in the markdown output (default: true)
includeImagesboolean否Include image references in the markdown output (default: false)
includeFramesboolean否Render iframe contents inline (default: false)
maxAgeMsnumber否Cache duration in milliseconds (0-2592000000, default: 86400000)
waitForMsnumber否Browser wait time after page load in milliseconds (0-30000)
timeoutMSnumber否Request timeout in milliseconds (1000-300000)
apiKeystring是Context.dev API key

输出

参数类型描述
markdownstringPage content as clean markdown
urlstringThe scraped URL

context_dev_scrape_html

Scrape any URL and return the raw HTML content of the page.

输入

参数类型必填描述
urlstring是The full URL to scrape (must include http:// or https://)
useMainContentOnlyboolean否Return only main content, excluding headers, footers, and navigation
includeFramesboolean否Render iframe contents inline into the returned HTML (default: false)
maxAgeMsnumber否Cache duration in milliseconds (0-2592000000, default: 86400000)
waitForMsnumber否Browser wait time after page load in milliseconds (0-30000)
timeoutMSnumber否Request timeout in milliseconds (1000-300000)
apiKeystring是Context.dev API key

输出

参数类型描述
htmlstringRaw HTML content of the page
urlstringThe scraped URL
typestringDetected content type (html, xml, json, text, csv, markdown, svg, pdf)

context_dev_scrape_images

Discover every image asset on a page, with optional dimension and type enrichment.

输入

参数类型必填描述
urlstring是The full URL to scrape images from (must include http:// or https://)
maxAgeMsnumber否Cache duration in milliseconds (0-2592000000, default: 86400000)
waitForMsnumber否Browser wait time after page load in milliseconds (0-30000)
timeoutMSnumber否Request timeout in milliseconds (1000-300000)
enrichResolutionboolean否Measure image dimensions (enables 5-credit enrichment)
enrichHostedUrlboolean否Host images on a CDN and return their URL and MIME type (enables enrichment)
enrichClassificationboolean否Classify each image by visual asset type (enables enrichment)
apiKeystring是Context.dev API key

输出

参数类型描述
successbooleanWhether the scrape succeeded
imagesarrayDiscovered image assets with source, element, type, and optional enrichment
↳ srcstringImage source URL or data
↳ elementstringSource element (img, svg, link, source, video, css, object, meta, background)
↳ typestringImage representation (url, html, base64)
↳ altstringAlt text
↳ enrichmentjsonOptional enrichment (width, height, mimetype, url, type) when requested
urlstringThe scraped URL

context_dev_screenshot

Capture a screenshot of any web page and store it as a downloadable image file.

输入

参数类型必填描述
urlstring是The full URL to capture (must include http:// or https://)
fullScreenshotboolean否Capture the full scrollable page instead of just the viewport (default: false)
handleCookiePopupboolean否Attempt to dismiss cookie banners before capturing (default: false)
viewportWidthnumber否Viewport width in pixels (240-7680, default: 1920)
viewportHeightnumber否Viewport height in pixels (240-4320, default: 1080)
maxAgeMsnumber否Cache duration in milliseconds (0-2592000000, default: 86400000)
waitForMsnumber否Post-load delay before capturing in milliseconds (0-30000, default: 3000)
timeoutMSnumber否Request timeout in milliseconds (1000-300000)
apiKeystring是Context.dev API key

输出

参数类型描述
filefileStored screenshot image file
screenshotUrlstringPublic URL of the captured screenshot
screenshotTypestringScreenshot type (viewport or fullPage)
domainstringDomain that was captured
widthnumberScreenshot width in pixels
heightnumberScreenshot height in pixels

context_dev_crawl

Crawl an entire website and return each discovered page as clean markdown.

输入

参数类型必填描述
urlstring是The starting URL to crawl (must include http:// or https://)
maxPagesnumber否Maximum number of pages to crawl (1-500, default: 100)
maxDepthnumber否Maximum link depth from the starting URL (0 = start page only)
urlRegexstring否Regex pattern to filter which URLs are crawled
includeLinksboolean否Preserve hyperlinks in the markdown output (default: true)
includeImagesboolean否Include image references in the markdown output (default: false)
useMainContentOnlyboolean否Strip headers, footers, and sidebars from each page (default: false)
followSubdomainsboolean否Follow links to subdomains of the starting domain (default: false)
maxAgeMsnumber否Cache duration in milliseconds (0-2592000000, default: 86400000)
waitForMsnumber否Browser wait time after page load in milliseconds (0-30000)
stopAfterMsnumber否Soft crawl time budget in milliseconds (10000-110000, default: 80000)
timeoutMSnumber否Request timeout in milliseconds (1000-300000)
apiKeystring是Context.dev API key

输出

参数类型描述
resultsarrayCrawled pages with markdown content and per-page metadata
↳ markdownstringPage content as markdown
↳ metadatajsonPage metadata (url, title, crawlDepth, statusCode)
metadataobjectCrawl summary (numUrls, maxCrawlDepth, numSucceeded, numFailed, numSkipped)

context_dev_map

Build a sitemap of a domain and return every discovered page URL.

输入

参数类型必填描述
domainstring是The domain to build a sitemap for (e.g., "example.com")
maxLinksnumber否Maximum number of URLs to return (1-100000, default: 10000)
urlRegexstring否RE2-compatible regex to filter URLs (max 256 chars)
timeoutMSnumber否Request timeout in milliseconds (1000-300000)
apiKeystring是Context.dev API key

输出

参数类型描述
domainstringThe domain that was mapped
urlsarrayAll page URLs discovered from the sitemap
metaobjectSitemap discovery stats (sitemapsDiscovered, sitemapsFetched, errors)

Search the web with natural language and optionally scrape results to markdown.

输入

参数类型必填描述
querystring是The natural language search query (1-500 characters)
includeDomainsarray否Only return results from these domains
excludeDomainsarray否Exclude results from these domains
freshnessstring否Recency filter (last_24_hours, last_week, last_month, last_year)
queryFanoutboolean否Expand the query into parallel variants for broader coverage
markdownEnabledboolean否Scrape each result page to markdown (default: false)
timeoutMSnumber否Request timeout in milliseconds (1000-300000)
apiKeystring是Context.dev API key

输出

参数类型描述
resultsarraySearch results with url, title, description, relevance, and optional markdown
↳ urlstringResult page URL
↳ titlestringResult page title
↳ descriptionstringResult snippet/description
↳ relevancestringRelevance rating (high, medium, low)
↳ markdownjsonScraped markdown for the result (when markdown scraping is enabled)
querystringThe query that was searched

context_dev_extract

Crawl a website and extract structured data matching a provided JSON schema.

输入

参数类型必填描述
urlstring是The starting website URL (must include http:// or https://)
schemajson是JSON Schema describing the structure of the data to extract
instructionsstring否Optional extraction guidance for link prioritization (max 2000 chars)
factCheckboolean否Require extracted values to be grounded in page facts (default: false)
followSubdomainsboolean否Follow links on subdomains of the starting domain (default: false)
maxPagesnumber否Maximum number of pages to analyze (1-50, default: 5)
maxDepthnumber否Maximum link depth from the starting URL
maxAgeMsnumber否Cache duration in milliseconds (0-2592000000, default: 604800000)
stopAfterMsnumber否Soft crawl time budget in milliseconds (10000-110000, default: 80000)
timeoutMSnumber否Request timeout in milliseconds (1000-300000)
apiKeystring是Context.dev API key

输出

参数类型描述
statusstringExtraction status
urlstringThe starting URL that was crawled
urlsAnalyzedarrayURLs that were analyzed during extraction
datajsonStructured data matching the requested schema
metadataobjectCrawl summary (numUrls, maxCrawlDepth, numSucceeded, numFailed, numSkipped)

context_dev_extract_product

Detect and extract structured product details from a single product page URL.

输入

参数类型必填描述
urlstring是The product page URL (must include http:// or https://)
maxAgeMsnumber否Cache duration in milliseconds (0-2592000000, default: 604800000)
timeoutMSnumber否Request timeout in milliseconds (1000-300000)
apiKeystring是Context.dev API key

输出

参数类型描述
isProductPagebooleanWhether the URL is a product page
platformstringDetected platform (amazon, tiktok_shop, etsy, generic)
productobjectExtracted product details
↳ namestringProduct name
↳ descriptionstringProduct description
↳ pricenumberProduct price
↳ currencystringPrice currency
↳ billing_frequencystringBilling frequency (monthly, yearly, one_time, usage_based)
↳ pricing_modelstringPricing model (per_seat, flat, tiered, freemium, custom)
↳ urlstringProduct URL
↳ categorystringProduct category
↳ featuresjsonProduct features
↳ target_audiencejsonTarget audience
↳ tagsjsonProduct tags
↳ image_urlstringPrimary product image URL
↳ imagesjsonProduct image URLs
↳ skustringProduct SKU

context_dev_extract_products

Extract the product catalog from a brand's website by domain (beta).

输入

参数类型必填描述
domainstring是The domain to extract products from (e.g., "example.com")
maxProductsnumber否Maximum number of products to extract (1-12)
maxAgeMsnumber否Cache duration in milliseconds (0-2592000000, default: 604800000)
timeoutMSnumber否Request timeout in milliseconds (1000-300000)
apiKeystring是Context.dev API key

输出

参数类型描述
productsarrayExtracted products with pricing, features, and metadata
↳ namestringProduct name
↳ descriptionstringProduct description
↳ pricenumberProduct price
↳ currencystringPrice currency
↳ billing_frequencystringBilling frequency (monthly, yearly, one_time, usage_based)
↳ pricing_modelstringPricing model (per_seat, flat, tiered, freemium, custom)
↳ urlstringProduct URL
↳ categorystringProduct category
↳ featuresjsonProduct features
↳ target_audiencejsonTarget audience
↳ tagsjsonProduct tags
↳ image_urlstringPrimary product image URL
↳ imagesjsonProduct image URLs
↳ skustringProduct SKU

context_dev_scrape_fonts

Extract the font families, usage stats, and font files used by a domain.

输入

参数类型必填描述
domainstring是The domain to extract fonts from (e.g., "example.com")
maxAgeMsnumber否Cache max age in milliseconds (86400000-31536000000, default: 7776000000)
timeoutMSnumber否Request timeout in milliseconds (1000-300000)
apiKeystring是Context.dev API key

输出

参数类型描述
statusstringExtraction status
domainstringThe domain that was analyzed
fontsarrayFonts with usage statistics and fallbacks
↳ fontstringFont family name
↳ usesjsonWhere the font is used
↳ fallbacksjsonFallback font families
↳ num_elementsnumberNumber of elements using the font
↳ num_wordsnumberNumber of words rendered in the font
↳ percent_wordsnumberPercent of words using the font
↳ percent_elementsnumberPercent of elements using the font
fontLinksjsonFont family download links keyed by font name (type, files, category)

context_dev_scrape_styleguide

Extract a domain's design system: colors, typography, spacing, shadows, and UI components.

输入

参数类型必填描述
domainstring是The domain to extract the styleguide from (e.g., "example.com")
maxAgeMsnumber否Cache max age in milliseconds (86400000-31536000000, default: 7776000000)
timeoutMSnumber否Request timeout in milliseconds (1000-300000)
apiKeystring是Context.dev API key

输出

参数类型描述
statusstringExtraction status
domainstringThe domain that was analyzed
styleguidejsonDesign system: mode, colors, typography, elementSpacing, shadows, fontLinks, components

context_dev_classify_naics

Classify a brand into NAICS industry codes from its domain or company name.

输入

参数类型必填描述
inputstring是Brand domain or company name to classify (e.g., "stripe.com" or "Stripe")
minResultsnumber否Minimum number of codes to return (1-10, default: 1)
maxResultsnumber否Maximum number of codes to return (1-10, default: 5)
timeoutMSnumber否Request timeout in milliseconds (1000-300000)
apiKeystring是Context.dev API key

输出

参数类型描述
statusstringClassification status
domainstringResolved domain
typestringInput type that was resolved
codesarrayMatched NAICS codes with name and confidence
↳ codestringIndustry code
↳ namestringIndustry name
↳ confidencestringMatch confidence (high, medium, low)

context_dev_classify_sic

Classify a brand into SIC industry codes from its domain or company name.

输入

参数类型必填描述
inputstring是Brand domain or company name to classify (e.g., "stripe.com" or "Stripe")
typestring否SIC taxonomy version: "original_sic" (default) or "latest_sec"
minResultsnumber否Minimum number of codes to return (1-10, default: 1)
maxResultsnumber否Maximum number of codes to return (1-10, default: 5)
timeoutMSnumber否Request timeout in milliseconds (1000-300000)
apiKeystring是Context.dev API key

输出

参数类型描述
statusstringClassification status
domainstringResolved domain
typestringInput type that was resolved
classificationstringSIC taxonomy version used (original_sic or latest_sec)
codesarrayMatched SIC codes with name, confidence, and group metadata
↳ codestringIndustry code
↳ namestringIndustry name
↳ confidencestringMatch confidence (high, medium, low)
↳ majorGroupstringMajor group code (original_sic only)
↳ majorGroupNamestringMajor group name (original_sic only)
↳ officestringSEC office (latest_sec only)

context_dev_get_brand

Retrieve brand data for a domain: logos, colors, backdrops, socials, address, and industry.

输入

参数类型必填描述
domainstring是The domain to retrieve brand data for (e.g., "airbnb.com")
forceLanguagestring否Override the detected language with a supported language code
maxSpeedboolean否Skip time-consuming operations for a faster response (default: false)
maxAgeMsnumber否Cache max age in milliseconds (86400000-31536000000, default: 7776000000)
timeoutMSnumber否Request timeout in milliseconds (1000-300000)
apiKeystring是Context.dev API key

输出

参数类型描述
statusstringRetrieval status
brandobjectBrand data object
↳ domainstringBrand domain
↳ titlestringBrand title
↳ descriptionstringBrand description
↳ sloganstringBrand slogan
↳ colorsjsonBrand colors (hex and name)
↳ logosjsonBrand logos with mode, colors, resolution, and type
↳ backdropsjsonBrand backdrop images
↳ socialsjsonSocial media profiles (type and url)
↳ addressjsonBrand address
↳ stockjsonStock info (ticker and exchange)
↳ is_nsfwbooleanWhether the brand contains adult content
↳ emailstringBrand contact email
↳ phonestringBrand contact phone
↳ industriesjsonIndustry taxonomy (eic industry/subindustry pairs)
↳ linksjsonKey brand links (careers, privacy, terms, blog, pricing)
↳ primary_languagestringPrimary language of the brand site

context_dev_get_brand_by_name

Retrieve brand data by company name: logos, colors, socials, address, and industry.

输入

参数类型必填描述
namestring是Company name to retrieve brand data for (3-30 chars, e.g., "Apple Inc")
countryGlstring否ISO 2-letter country code to prioritize (e.g., "us")
forceLanguagestring否Override the detected language with a supported language code
maxSpeedboolean否Skip time-consuming operations for a faster response (default: false)
maxAgeMsnumber否Cache max age in milliseconds (86400000-31536000000, default: 7776000000)
timeoutMSnumber否Request timeout in milliseconds (1000-300000)
apiKeystring是Context.dev API key

输出

参数类型描述
statusstringRetrieval status
brandobjectBrand data object
↳ domainstringBrand domain
↳ titlestringBrand title
↳ descriptionstringBrand description
↳ sloganstringBrand slogan
↳ colorsjsonBrand colors (hex and name)
↳ logosjsonBrand logos with mode, colors, resolution, and type
↳ backdropsjsonBrand backdrop images
↳ socialsjsonSocial media profiles (type and url)
↳ addressjsonBrand address
↳ stockjsonStock info (ticker and exchange)
↳ is_nsfwbooleanWhether the brand contains adult content
↳ emailstringBrand contact email
↳ phonestringBrand contact phone
↳ industriesjsonIndustry taxonomy (eic industry/subindustry pairs)
↳ linksjsonKey brand links (careers, privacy, terms, blog, pricing)
↳ primary_languagestringPrimary language of the brand site

context_dev_get_brand_by_email

Retrieve brand data from a work email address. Free/disposable emails are rejected (422).

输入

参数类型必填描述
emailstring是Work email address; the domain is extracted (free providers are rejected)
forceLanguagestring否Override the detected language with a supported language code
maxSpeedboolean否Skip time-consuming operations for a faster response (default: false)
maxAgeMsnumber否Cache max age in milliseconds (86400000-31536000000, default: 7776000000)
timeoutMSnumber否Request timeout in milliseconds (1000-300000)
apiKeystring是Context.dev API key

输出

参数类型描述
statusstringRetrieval status
brandobjectBrand data object
↳ domainstringBrand domain
↳ titlestringBrand title
↳ descriptionstringBrand description
↳ sloganstringBrand slogan
↳ colorsjsonBrand colors (hex and name)
↳ logosjsonBrand logos with mode, colors, resolution, and type
↳ backdropsjsonBrand backdrop images
↳ socialsjsonSocial media profiles (type and url)
↳ addressjsonBrand address
↳ stockjsonStock info (ticker and exchange)
↳ is_nsfwbooleanWhether the brand contains adult content
↳ emailstringBrand contact email
↳ phonestringBrand contact phone
↳ industriesjsonIndustry taxonomy (eic industry/subindustry pairs)
↳ linksjsonKey brand links (careers, privacy, terms, blog, pricing)
↳ primary_languagestringPrimary language of the brand site

context_dev_get_brand_by_ticker

Retrieve brand data for a public company by its stock ticker symbol.

输入

参数类型必填描述
tickerstring是Stock ticker symbol (e.g., "AAPL", "GOOGL", "BRK.A")
tickerExchangestring否Exchange code for the ticker (e.g., "NASDAQ", "NYSE", "LSE"). Default: NASDAQ
forceLanguagestring否Override the detected language with a supported language code
maxSpeedboolean否Skip time-consuming operations for a faster response (default: false)
maxAgeMsnumber否Cache max age in milliseconds (86400000-31536000000, default: 7776000000)
timeoutMSnumber否Request timeout in milliseconds (1000-300000)
apiKeystring是Context.dev API key

输出

参数类型描述
statusstringRetrieval status
brandobjectBrand data object
↳ domainstringBrand domain
↳ titlestringBrand title
↳ descriptionstringBrand description
↳ sloganstringBrand slogan
↳ colorsjsonBrand colors (hex and name)
↳ logosjsonBrand logos with mode, colors, resolution, and type
↳ backdropsjsonBrand backdrop images
↳ socialsjsonSocial media profiles (type and url)
↳ addressjsonBrand address
↳ stockjsonStock info (ticker and exchange)
↳ is_nsfwbooleanWhether the brand contains adult content
↳ emailstringBrand contact email
↳ phonestringBrand contact phone
↳ industriesjsonIndustry taxonomy (eic industry/subindustry pairs)
↳ linksjsonKey brand links (careers, privacy, terms, blog, pricing)
↳ primary_languagestringPrimary language of the brand site

context_dev_get_brand_simplified

Retrieve essential brand data for a domain: title, colors, logos, and backdrops.

输入

参数类型必填描述
domainstring是The domain to retrieve simplified brand data for (e.g., "airbnb.com")
maxAgeMsnumber否Cache max age in milliseconds (86400000-31536000000, default: 7776000000)
timeoutMSnumber否Request timeout in milliseconds (1000-300000)
apiKeystring是Context.dev API key

输出

参数类型描述
statusstringRetrieval status
brandobjectSimplified brand data (domain, title, colors, logos, backdrops)
↳ domainstringBrand domain
↳ titlestringBrand title
↳ colorsjsonBrand colors (hex and name)
↳ logosjsonBrand logos with mode, colors, resolution, and type
↳ backdropsjsonBrand backdrop images

context_dev_identify_transaction

Identify the brand behind a raw bank/card transaction descriptor and return its brand data.

输入

参数类型必填描述
transactionInfostring是The raw transaction descriptor or identifier to resolve to a brand
countryGlstring否ISO 2-letter country code from the transaction (e.g., "us", "gb")
citystring否City name to prioritize in the search
mccstring否Merchant Category Code for the business category
phonenumber否Phone number from the transaction for verification
highConfidenceOnlyboolean否Enforce additional verification steps for higher confidence (default: false)
forceLanguagestring否Override the detected language with a supported language code
maxSpeedboolean否Skip time-consuming operations for a faster response (default: false)
timeoutMSnumber否Request timeout in milliseconds (1000-300000)
apiKeystring是Context.dev API key

输出

参数类型描述
statusstringIdentification status
brandobjectBrand data for the identified merchant
↳ domainstringBrand domain
↳ titlestringBrand title
↳ descriptionstringBrand description
↳ sloganstringBrand slogan
↳ colorsjsonBrand colors (hex and name)
↳ logosjsonBrand logos with mode, colors, resolution, and type
↳ backdropsjsonBrand backdrop images
↳ socialsjsonSocial media profiles (type and url)
↳ addressjsonBrand address
↳ stockjsonStock info (ticker and exchange)
↳ is_nsfwbooleanWhether the brand contains adult content
↳ emailstringBrand contact email
↳ phonestringBrand contact phone
↳ industriesjsonIndustry taxonomy (eic industry/subindustry pairs)
↳ linksjsonKey brand links (careers, privacy, terms, blog, pricing)
↳ primary_languagestringPrimary language of the brand site

context_dev_prefetch_domain

Queue a domain for brand-data prefetching to reduce latency on later requests (subscribers; 0 credits).

输入

参数类型必填描述
domainstring是The domain to prefetch brand data for (e.g., "example.com")
timeoutMSnumber否Request timeout in milliseconds (1000-300000)
apiKeystring是Context.dev API key

输出

参数类型描述
statusstringPrefetch status
messagestringHuman-readable prefetch result message
domainstringThe domain queued for prefetching

context_dev_prefetch_by_email

Queue an email's domain for brand-data prefetching to reduce later latency (subscribers; 0 credits). Free/disposable emails are rejected.

输入

参数类型必填描述
emailstring是Work email address whose domain should be prefetched (free providers rejected)
timeoutMSnumber否Request timeout in milliseconds (1000-300000)
apiKeystring是Context.dev API key

输出

参数类型描述
statusstringPrefetch status
messagestringHuman-readable prefetch result message
domainstringThe domain queued for prefetching

On this page