Crawl

Create Crawl

POST
/v2/crawl

Authorization

BearerAuth
AuthorizationBearer <token>

In: header

Request Body

application/json

Request body model for the /crawl endpoint

allow_external_links?boolean

Allows the crawler to follow links to external websites.

Defaultfalse
allow_subdomains?boolean

Allows the crawler to follow links to subdomains of the main domain.

Defaultfalse
callback?||

Webhook configuration for receiving crawl results.

crawl_entire_domain?boolean

Allows the crawler to follow internal links to sibling or parent URLs, not just child paths.

Defaultfalse
exclude_paths?array<string>

URL pathname regex patterns that exclude matching URLs from the crawl.

extract_options?
ignore_query_parameters?boolean

Do not re-scrape the same path with different (or none) query parameters.

Defaultfalse
include_paths?array<string>

URL pathname regex patterns that include matching URLs in the crawl.

limit?integer

Maximum number of pages to crawl.

Range1 <= value <= 10000
Default5000
max_discovery_depth?integer

Maximum depth to crawl based on discovery order.

Range1 <= value <= 20
Default5
name?string

Name of the crawl.

sitemap?string

Sitemap and other methods will be used together to find URLs.

Default"include"

Value in

  • "skip"
  • "include"
  • "only"
url*string

Url to crawl.

Response Body

application/json

curl -X POST "https://example.com/v2/crawl" \  -H "Content-Type: application/json" \  -d '{    "url": "string"  }'
{  "account_name": "string",  "completed": 0,  "completed_at": "string",  "crawl_id": "c9eee371-0b50-4a83-baff-95c2f0e36a00",  "crawl_options": {    "allow_external_links": false,    "allow_subdomains": false,    "callback": {      "events": [        "started"      ],      "headers": {        "property1": "string",        "property2": "string"      },      "metadata": {        "property1": null,        "property2": null      },      "url": "http://example.com"    },    "crawl_entire_domain": false,    "exclude_paths": [      "string"    ],    "ignore_query_parameters": false,    "include_paths": [      "string"    ],    "limit": 5000,    "max_discovery_depth": 5,    "sitemap": "skip",    "property1": null,    "property2": null  },  "created_at": "string",  "extract_options": {    "property1": null,    "property2": null  },  "failed": 0,  "name": "string",  "pending": 0,  "status": "queued",  "tasks": [    {      "created_at": "string",      "status": "pending",      "task_id": "string",      "updated_at": "string"    }  ],  "total": 0,  "updated_at": "string",  "url": "http://example.com"}