Skip to main content

Firecrawl (Global, Search & Page Scraping)

The platform offers three Firecrawl endpoints: Search for search results, Map for discovering a site's URL structure, and Scrape for extracting the body text of a single page. They are independent — call them on their own, or chain them (Map to find pages → Scrape for content → feed to an LLM).

Official documentation: Firecrawl API Reference ↗

Common conventions​

  • All three endpoints are POST + JSON body and need only one header: Authorization: Bearer $TURING_API_KEY.
  • Each endpoint accepts only the fields listed in the tables below. Any other field returns 422 and the request is never sent. Nested objects are just as strict: an unlisted key inside categories[] or location is rejected too.
  • Field names are best written in camelCase (includeDomains, onlyMainContent, redactPII); the matching snake_case form (include_domains, only_main_content, redact_pii) is also accepted. Capitalization must match one of those two exactly — the URL in ignoreInvalidURLs is three capitals, and ignoreInvalidUrls is rejected as an unlisted field.
  • location is a string on Search and an object on Map / Scrape; the two are not interchangeable.
  • Every url / domain must be a public http(s) address: localhost, *.local, *.internal, private IPs, non-canonical IP forms (e.g. 0x7f.1) and credential-bearing URLs are all rejected. Domain fields take a host — including https:// is rejected.
  • Every timeout is in milliseconds. The platform adds a 5-second buffer on top of your value for its own read timeout.
  • Fields whose value is null are omitted from the response.
  • Endpoint: POST /proxy/firecrawl/search
  • Required parameter: query
ParameterType / defaultDescription
querystring, requiredQuery string, 1–500 characters; surrounding whitespace is trimmed and it must not be empty
limitint, default 10Number of results, 1–100
sourcesstring[]web / images / news, 1–3 items
categoriesobject[]Elements shaped like {"type": "github"}; type is github / research / pdf, up to 3 items. Note this is an array of objects — a bare ["github"] is rejected
includeDomainsstring[]Keep only results from these domains, 1–20 items, host only (no scheme or path), no duplicates
excludeDomainsstring[]Exclude these domains, same rules. Cannot be combined with includeDomains
tbsstringGoogle time-filter string, ≤100 characters
locationstringRetrieval location hint, ≤200 characters
countrystring, default USTwo-letter uppercase country code; lowercase input is uppercased automatically
timeoutint, default 600001000–60000 ms
ignoreInvalidURLsbool, default falseSkip invalid URLs instead of failing the whole request
highlightsbool, default trueWhether to return matched snippets
note

Search returns links and summaries only, not page body text. When you need the content, take the URLs and call Scrape.

curl $TURING_BASE_URL/proxy/firecrawl/search \
-H "Authorization: Bearer $TURING_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"query": "Firecrawl API",
"limit": 5,
"sources": ["web"],
"includeDomains": ["docs.firecrawl.dev"]
}'

Response excerpt:

{
"success": true,
"data": {
"web": [
{
"url": "https://docs.firecrawl.dev",
"position": 1
}
]
},
"creditsUsed": 2
}

The top level returns success and data, and may also carry creditsUsed (credits consumed by this call), warning and id. data is grouped by sources (web / images / news), each group being an array whose items are returned as-is.

Map​

Discovers which URLs a site has without scraping any content. Useful for surveying a site's structure first, then picking pages to Scrape.

  • Endpoint: POST /proxy/firecrawl/map
  • Required parameter: url
ParameterType / defaultDescription
urlstring, requiredSeed URL, ≤2048 characters
searchstringFilter discovered URLs by keyword, 1–200 characters
sitemapdefault includeinclude uses the sitemap plus crawling, only uses the sitemap alone, skip ignores it
includeSubdomainsbool, default falseWhether to include subdomains
ignoreQueryParametersbool, default trueTreat URLs that differ only by query string as the same URL
ignoreCachebool, default falseSkip the cache and rediscover
limitint, default 1000Maximum URLs returned, 1–5000
timeoutint, default 600001000–60000 ms
locationobject{"country": "GB", "languages": ["en-GB"]}; country is a two-letter code and languages takes 1–10 items of ≤35 characters each
curl $TURING_BASE_URL/proxy/firecrawl/map \
-H "Authorization: Bearer $TURING_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://docs.firecrawl.dev",
"search": "api",
"limit": 100
}'

Response excerpt:

{
"success": true,
"links": [
{
"url": "https://docs.firecrawl.dev/api-reference/introduction",
"title": "Introduction",
"description": "Firecrawl API reference"
}
]
}
note

Each entry in links is exactly three fields — url / title / description — of which url is always present and the other two may be absent.

Scrape​

Scrapes the body text of a single URL. To scrape many pages, use Map to get the URL list first, then call Scrape for each one.

  • Endpoint: POST /proxy/firecrawl/scrape
  • Required parameter: url
ParameterType / defaultDescription
urlstring, requiredTarget URL, ≤2048 characters
formatsstring[], default ["markdown"]markdown / summary / html / rawHtml / links, 1–5 items, no duplicates
onlyMainContentbool, default trueKeep only the main content, dropping navigation and footers
onlyCleanContentbool, default falseStrip additional noise content
includeTagsstring[]Keep only these HTML tags / selectors, 1–50 items of ≤128 characters each, no duplicates
excludeTagsstring[]Exclude these tags / selectors, same rules
maxAgeint, default 172800000 (48 hours)Maximum acceptable cache age; 0 forces a fresh scrape
minAgeint, ≥1Minimum acceptable cache age
headersobjectCustom request headers, e.g. {"Cookie": "session=..."}
waitForint, default 0Milliseconds to wait for rendering, 0–5000
mobilebool, default falseScrape with a mobile user agent / viewport
skipTlsVerificationbool, default falseSkip certificate validation; enable only when the target site genuinely has a broken certificate
timeoutint, default 600001000–295000 ms (more generous than Search / Map)
locationobjectSame as Map
blockAdsbool, default trueBlock ads
storeInCachebool, default trueWhether to write this result to the cache
redactPIIbool or object, default falsePass true for the default policy, or an object specifying mode (accurate / aggressive / fast), entities (PERSON / EMAIL / PHONE / LOCATION / FINANCIAL / SECRET) and replaceStyle (tag / mask / remove)

Fields with a fixed value​

These three fields have a fixed value. You may send them explicitly, but only with that value — anything else returns 422:

FieldOnly acceptsEffect
parsers[]No PDF page parsing. A scraped PDF is handled as one whole file, which is also why it bills as 1 credit
removeBase64ImagestrueInline base64 images are always stripped, keeping the response from ballooning
proxy"basic"Uses the basic proxy
curl $TURING_BASE_URL/proxy/firecrawl/scrape \
-H "Authorization: Bearer $TURING_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://docs.firecrawl.dev",
"formats": ["markdown", "links"],
"onlyMainContent": true
}'

Response excerpt:

{
"success": true,
"data": {
"markdown": "# Firecrawl Docs\n...",
"links": ["https://docs.firecrawl.dev/api-reference/introduction"],
"metadata": {
"contentType": "text/html; charset=utf-8"
}
}
}

The keys in data follow the formats you requested — markdown / summary / html / rawHtml / links each map to a key — and metadata (page metadata) may also be present. The contents of data are returned as-is.

Billing​

Billed in Firecrawl credits at $0.00099 per credit. Failed requests — validation errors, upstream failures, invalid responses — are not billed.

EndpointCredit consumption
SearchTaken from creditsUsed in the response; when that field is absent, estimated as ceil(total results / 10) × 2 (total results being the sum of item counts across all source groups)
Map1 credit per successful call, regardless of how many links come back (an empty array still counts as 1)
Scrape1 credit per successful call; an HTML page and a PDF cost the same
Rate limits and errors
  • Rate limits: Search / Map / Scrape are counted separately, at a default of 480 requests per hour per endpoint (enterprise customers are provisioned separately). Quota and rate limits are checked before the request goes out, so exceeding them consumes no Firecrawl quota.
  • 422: The request body failed a field or value check; nothing was sent.
  • 429 (error code 3091): Upstream rate limit hit.
  • 502 (error code 3090): Upstream error, malformed response, or a response body larger than 10 MB. A uniform error message is returned rather than the upstream's raw error content, so include the X-Turing-Trace-Id when reporting an issue.

See Error Codes for error code meanings and Rate Limits for the rate-limiting policy.