Apify Website Content Crawler
Apify runs an authorized website content crawler Actor from a URL list and exposes run status and dataset items through its REST API. Actor output can include page text, HTML, Markdown, and metadata depending on the selected crawler input and output schema.
Connect this API to an AI agent Browse all specs
What you can do with Apify Website Content Crawler via MCP
Every endpoint below becomes a governed tool your AI agent (Claude, ChatGPT, or a customer-facing VX agent) can call — with scoped permissions, sandbox testing, and audit logs.
| Method | Path | Operation | Description |
|---|---|---|---|
| POST | /v2/acts/{actor_id}/runs |
postV2ActsByActorIdRuns | Start a crawler Actor run Method: POST Path: /v2/acts/{actor_id}/runs IMPORTANT: This function has 2 REQUIRED parameter(s) and 4 OPTIONAL parameter(s) REQUIRED parameters MUST be provided OPTIONAL parameters can be omitted if not needed Parameters: Path Parameters: REQUIRED: - actor_id: Apify Actor identifier. Use an official or authorized web content crawler Actor. Request Body: REQUIRED: - startUrls: startUrls array OPTIONAL: - maxCrawlDepth: maxCrawlDepth parameter from application/json - maxCrawlPages: maxCrawlPages parameter from application/json - saveMarkdown: saveMarkdown parameter from application/json - proxyConfiguration: proxyConfiguration parameter from application/json |
| GET | /v2/actor-runs/{run_id} |
getV2ActorRunsByRunId | Get Actor run status Method: GET Path: /v2/actor-runs/{run_id} IMPORTANT: This function has 2 REQUIRED parameter(s) and 0 OPTIONAL parameter(s) REQUIRED parameters MUST be provided OPTIONAL parameters can be omitted if not needed Parameters: Path Parameters: REQUIRED: - run_id: Actor run identifier. Query Parameters: REQUIRED: - token: Apify API token. |
| GET | /v2/datasets/{dataset_id}/items |
getV2DatasetsByDatasetIdItems | Get crawler dataset items Method: GET Path: /v2/datasets/{dataset_id}/items IMPORTANT: This function has 2 REQUIRED parameter(s) and 1 OPTIONAL parameter(s) REQUIRED parameters MUST be provided OPTIONAL parameters can be omitted if not needed Parameters: Path Parameters: REQUIRED: - dataset_id: Dataset identifier returned by the run. Query Parameters: REQUIRED: - token: Apify API token. OPTIONAL: - format: Dataset response format. |
Usage guide
## Authentication Create an Apify API token in the Apify Console. Authenticate with the `token` query parameter or the documented `Authorization: Bearer <APIFY_TOKEN>` header. This specification models the query-token form because it is directly supported by Apify API URLs; keep the token server-side. ## Operations - `POST /v2/acts/{actor_id}/runs?token=...` starts an authorized crawler Actor with `startUrls` and crawl limits. - `GET /v2/actor-runs/{run_id}?token=...` checks the run status and obtains the default dataset ID. - `GET /v2/datasets/{dataset_id}/items?token=...` retrieves extracted records, commonly including text/Markdown or HTML fields based on Actor configuration. ## Example ```bash curl -X POST "https://api.apify.com/v2/acts/apify~website-content-crawler/runs?token=$APIFY_TOKEN" \ -H "Content-Type: application/json" \ -d '{"startUrls":[{"url":"https://example.com"}],"maxCrawlPages":10,"maxCrawlDepth":1}' ``` Only crawl websites you are authorized to access and follow the target site's terms and applicable law. Official documentation: https://docs.apify.com/api/v2
How to use this spec
- Create a free VX Agents account (sandbox tools included).
- Open the marketplace and connect “Apify Website Content Crawler” to an agent.
- Expose it as an MCP server for Claude/ChatGPT, or chat with it directly.