Firecrawl (Global, Search & Page Scraping)
The platform offers three Firecrawl endpoints: Search for search results, Map for discovering a site's URL structure, and Scrape for extracting the body text of a single page. They are independent — call them on their own, or chain them (Map to find pages → Scrape for content → feed to an LLM).
Official documentation: Firecrawl API Reference ↗
Common conventions
- All three endpoints are
POST+ JSON body and need only one header:Authorization: Bearer $TURING_API_KEY. - Each endpoint accepts only the fields listed in the tables below. Any other field returns
422and the request is never sent. Nested objects are just as strict: an unlisted key insidecategories[]orlocationis rejected too. - Field names are best written in camelCase (
includeDomains,onlyMainContent,redactPII); the matching snake_case form (include_domains,only_main_content,redact_pii) is also accepted. Capitalization must match one of those two exactly — theURLinignoreInvalidURLsis three capitals, andignoreInvalidUrlsis rejected as an unlisted field. locationis a string on Search and an object on Map / Scrape; the two are not interchangeable.- Every
url/ domain must be a public http(s) address:localhost,*.local,*.internal, private IPs, non-canonical IP forms (e.g.0x7f.1) and credential-bearing URLs are all rejected. Domain fields take a host — includinghttps://is rejected. - Every
timeoutis in milliseconds. The platform adds a 5-second buffer on top of your value for its own read timeout. - Fields whose value is
nullare omitted from the response.
Search
- Endpoint:
POST /proxy/firecrawl/search - Required parameter:
query
| Parameter | Type / default | Description |
|---|---|---|
query | string, required | Query string, 1–500 characters; surrounding whitespace is trimmed and it must not be empty |
limit | int, default 10 | Number of results, 1–100 |
sources | string[] | web / images / news, 1–3 items |
categories | object[] | Elements shaped like {"type": "github"}; type is github / research / pdf, up to 3 items. Note this is an array of objects — a bare ["github"] is rejected |
includeDomains | string[] | Keep only results from these domains, 1–20 items, host only (no scheme or path), no duplicates |
excludeDomains | string[] | Exclude these domains, same rules. Cannot be combined with includeDomains |
tbs | string | Google time-filter string, ≤100 characters |
location | string | Retrieval location hint, ≤200 characters |
country | string, default US | Two-letter uppercase country code; lowercase input is uppercased automatically |
timeout | int, default 60000 | 1000–60000 ms |
ignoreInvalidURLs | bool, default false | Skip invalid URLs instead of failing the whole request |
highlights | bool, default true | Whether to return matched snippets |
Search returns links and summaries only, not page body text. When you need the content, take the URLs and call Scrape.
curl $TURING_BASE_URL/proxy/firecrawl/search \
-H "Authorization: Bearer $TURING_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"query": "Firecrawl API",
"limit": 5,
"sources": ["web"],
"includeDomains": ["docs.firecrawl.dev"]
}'
Response excerpt:
{
"success": true,
"data": {
"web": [
{
"url": "https://docs.firecrawl.dev",
"position": 1
}
]
},
"creditsUsed": 2
}
The top level returns success and data, and may also carry creditsUsed (credits consumed by this call), warning and id. data is grouped by sources (web / images / news), each group being an array whose items are returned as-is.
Map
Discovers which URLs a site has without scraping any content. Useful for surveying a site's structure first, then picking pages to Scrape.
- Endpoint:
POST /proxy/firecrawl/map - Required parameter:
url
| Parameter | Type / default | Description |
|---|---|---|
url | string, required | Seed URL, ≤2048 characters |
search | string | Filter discovered URLs by keyword, 1–200 characters |
sitemap | default include | include uses the sitemap plus crawling, only uses the sitemap alone, skip ignores it |
includeSubdomains | bool, default false | Whether to include subdomains |
ignoreQueryParameters | bool, default true | Treat URLs that differ only by query string as the same URL |
ignoreCache | bool, default false | Skip the cache and rediscover |
limit | int, default 1000 | Maximum URLs returned, 1–5000 |
timeout | int, default 60000 | 1000–60000 ms |
location | object | {"country": "GB", "languages": ["en-GB"]}; country is a two-letter code and languages takes 1–10 items of ≤35 characters each |
curl $TURING_BASE_URL/proxy/firecrawl/map \
-H "Authorization: Bearer $TURING_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://docs.firecrawl.dev",
"search": "api",
"limit": 100
}'
Response excerpt:
{
"success": true,
"links": [
{
"url": "https://docs.firecrawl.dev/api-reference/introduction",
"title": "Introduction",
"description": "Firecrawl API reference"
}
]
}
Each entry in links is exactly three fields — url / title / description — of which url is always present and the other two may be absent.
Scrape
Scrapes the body text of a single URL. To scrape many pages, use Map to get the URL list first, then call Scrape for each one.
- Endpoint:
POST /proxy/firecrawl/scrape - Required parameter:
url
| Parameter | Type / default | Description |
|---|---|---|
url | string, required | Target URL, ≤2048 characters |
formats | string[], default ["markdown"] | markdown / summary / html / rawHtml / links, 1–5 items, no duplicates |
onlyMainContent | bool, default true | Keep only the main content, dropping navigation and footers |
onlyCleanContent | bool, default false | Strip additional noise content |
includeTags | string[] | Keep only these HTML tags / selectors, 1–50 items of ≤128 characters each, no duplicates |
excludeTags | string[] | Exclude these tags / selectors, same rules |
maxAge | int, default 172800000 (48 hours) | Maximum acceptable cache age; 0 forces a fresh scrape |
minAge | int, ≥1 | Minimum acceptable cache age |
headers | object | Custom request headers, e.g. {"Cookie": "session=..."} |
waitFor | int, default 0 | Milliseconds to wait for rendering, 0–5000 |
mobile | bool, default false | Scrape with a mobile user agent / viewport |
skipTlsVerification | bool, default false | Skip certificate validation; enable only when the target site genuinely has a broken certificate |
timeout | int, default 60000 | 1000–295000 ms (more generous than Search / Map) |
location | object | Same as Map |
blockAds | bool, default true | Block ads |
storeInCache | bool, default true | Whether to write this result to the cache |
redactPII | bool or object, default false | Pass true for the default policy, or an object specifying mode (accurate / aggressive / fast), entities (PERSON / EMAIL / PHONE / LOCATION / FINANCIAL / SECRET) and replaceStyle (tag / mask / remove) |
Fields with a fixed value
These three fields have a fixed value. You may send them explicitly, but only with that value — anything else returns 422:
| Field | Only accepts | Effect |
|---|---|---|
parsers | [] | No PDF page parsing. A scraped PDF is handled as one whole file, which is also why it bills as 1 credit |
removeBase64Images | true | Inline base64 images are always stripped, keeping the response from ballooning |
proxy | "basic" | Uses the basic proxy |
curl $TURING_BASE_URL/proxy/firecrawl/scrape \
-H "Authorization: Bearer $TURING_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"url": "https://docs.firecrawl.dev",
"formats": ["markdown", "links"],
"onlyMainContent": true
}'
Response excerpt:
{
"success": true,
"data": {
"markdown": "# Firecrawl Docs\n...",
"links": ["https://docs.firecrawl.dev/api-reference/introduction"],
"metadata": {
"contentType": "text/html; charset=utf-8"
}
}
}
The keys in data follow the formats you requested — markdown / summary / html / rawHtml / links each map to a key — and metadata (page metadata) may also be present. The contents of data are returned as-is.
Billing
Billed in Firecrawl credits at $0.00099 per credit. Failed requests — validation errors, upstream failures, invalid responses — are not billed.
| Endpoint | Credit consumption |
|---|---|
| Search | Taken from creditsUsed in the response; when that field is absent, estimated as ceil(total results / 10) × 2 (total results being the sum of item counts across all source groups) |
| Map | 1 credit per successful call, regardless of how many links come back (an empty array still counts as 1) |
| Scrape | 1 credit per successful call; an HTML page and a PDF cost the same |
- Rate limits: Search / Map / Scrape are counted separately, at a default of 480 requests per hour per endpoint (enterprise customers are provisioned separately). Quota and rate limits are checked before the request goes out, so exceeding them consumes no Firecrawl quota.
422: The request body failed a field or value check; nothing was sent.429(error code 3091): Upstream rate limit hit.502(error code 3090): Upstream error, malformed response, or a response body larger than 10 MB. A uniform error message is returned rather than the upstream's raw error content, so include theX-Turing-Trace-Idwhen reporting an issue.
See Error Codes for error code meanings and Rate Limits for the rate-limiting policy.