VX Agents

OpenAPI Marketplace

Apify Website Content Crawler

v1.0.0 · 3 endpoints · 0 connections · shared by vx.ai

Apify runs an authorized website content crawler Actor from a URL list and exposes run status and dataset items through its REST API. Actor output can include page text, HTML, Markdown, and metadata depending on the selected crawler input and output schema.

Connect this API to an AI agent Browse all specs

What you can do with Apify Website Content Crawler via MCP

Every endpoint below becomes a governed tool your AI agent (Claude, ChatGPT, or a customer-facing VX agent) can call — with scoped permissions, sandbox testing, and audit logs.

MethodPathOperationDescription
POST /v2/acts/{actor_id}/runs postV2ActsByActorIdRuns Start a crawler Actor run Method: POST Path: /v2/acts/{actor_id}/runs IMPORTANT: This function has 2 REQUIRED parameter(s) and 4 OPTIONAL parameter(s) REQUIRED parameters MUST be provided OPTIONAL parameters can be omitted if not needed Parameters: Path Parameters: REQUIRED: - actor_id: Apify Actor identifier. Use an official or authorized web content crawler Actor. Request Body: REQUIRED: - startUrls: startUrls array OPTIONAL: - maxCrawlDepth: maxCrawlDepth parameter from application/json - maxCrawlPages: maxCrawlPages parameter from application/json - saveMarkdown: saveMarkdown parameter from application/json - proxyConfiguration: proxyConfiguration parameter from application/json
GET /v2/actor-runs/{run_id} getV2ActorRunsByRunId Get Actor run status Method: GET Path: /v2/actor-runs/{run_id} IMPORTANT: This function has 2 REQUIRED parameter(s) and 0 OPTIONAL parameter(s) REQUIRED parameters MUST be provided OPTIONAL parameters can be omitted if not needed Parameters: Path Parameters: REQUIRED: - run_id: Actor run identifier. Query Parameters: REQUIRED: - token: Apify API token.
GET /v2/datasets/{dataset_id}/items getV2DatasetsByDatasetIdItems Get crawler dataset items Method: GET Path: /v2/datasets/{dataset_id}/items IMPORTANT: This function has 2 REQUIRED parameter(s) and 1 OPTIONAL parameter(s) REQUIRED parameters MUST be provided OPTIONAL parameters can be omitted if not needed Parameters: Path Parameters: REQUIRED: - dataset_id: Dataset identifier returned by the run. Query Parameters: REQUIRED: - token: Apify API token. OPTIONAL: - format: Dataset response format.

Usage guide

## Authentication Create an Apify API token in the Apify Console. Authenticate with the `token` query parameter or the documented `Authorization: Bearer <APIFY_TOKEN>` header. This specification models the query-token form because it is directly supported by Apify API URLs; keep the token server-side. ## Operations - `POST /v2/acts/{actor_id}/runs?token=...` starts an authorized crawler Actor with `startUrls` and crawl limits. - `GET /v2/actor-runs/{run_id}?token=...` checks the run status and obtains the default dataset ID. - `GET /v2/datasets/{dataset_id}/items?token=...` retrieves extracted records, commonly including text/Markdown or HTML fields based on Actor configuration. ## Example ```bash curl -X POST "https://api.apify.com/v2/acts/apify~website-content-crawler/runs?token=$APIFY_TOKEN" \ -H "Content-Type: application/json" \ -d '{"startUrls":[{"url":"https://example.com"}],"maxCrawlPages":10,"maxCrawlDepth":1}' ``` Only crawl websites you are authorized to access and follow the target site's terms and applicable law. Official documentation: https://docs.apify.com/api/v2

How to use this spec

  1. Create a free VX Agents account (sandbox tools included).
  2. Open the marketplace and connect “Apify Website Content Crawler” to an agent.
  3. Expose it as an MCP server for Claude/ChatGPT, or chat with it directly.