# Botify Agent: HTML Fetch

**ID:** `html_fetch`

Get the HTML of a given URL.


    Fetch the HTML of a URL, from the Botify SiteCrawler archive by default, or
    live through a real-time fetch or the rendering farm (which executes the page
    JavaScript) when `origin` asks for it.

    The HTML is cleaned before being returned: `remove_tags` drops the listed tags,
    `only_body` keeps the body only and `only_text` strips the markup down to text.
    

## Authentication

All endpoints require a **Bearer token** in the `Authorization` header.

```
Authorization: Bearer <your-token>
```

## Calling this agent

This agent supports the following processing modes:

| Mode | Type | Description |
|------|------|-------------|
| **Process** | Synchronous | Single-item processing. Best for real-time requests with immediate response. |
| **Batch Process** | Synchronous | Process multiple items in a single request for efficiency. |
| **Async Process** | Asynchronous | Single-item processing for long-running tasks that exceed timeout limits. |
| **Async Batch Process** | Asynchronous | Large-scale batch jobs with background processing. |

# Synchronous Processing

Synchronous calls block until the result is ready. Use these for quick operations where you need immediate results.

## Single Item Processing

Process a single input and receive the result immediately.

```
POST https://agents.botify.com/{org}/{project}/html_fetch/process
```

### Request body

```json
{
  "item": {
    "url": "<url>"
  }
}
```

### Response

Returns a single processed item (HTTP 200).

### cURL example

```bash
curl -X POST "https://agents.botify.com/{org}/{project}/html_fetch/process" \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
  "item": {
    "url": "<url>"
  }
}'
```

## Batch Processing

Process multiple inputs in a single synchronous request. More efficient than making individual calls. Items are processed concurrently on the server.

```
POST https://agents.botify.com/{org}/{project}/html_fetch/batch_process
```

### Request body

```json
{
  "items": [
    {
      "url": "<url>"
    }
  ]
}
```

### Response

Returns an array of results. Each element is either a successful processed item (`"status": "success"`) or an error (`"status": "error"`).

### cURL example

```bash
curl -X POST "https://agents.botify.com/{org}/{project}/html_fetch/batch_process" \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
  "items": [
    {
      "url": "<url>"
    }
  ]
}'
```

### When to use synchronous calls

- Processing completes within timeout limits (typically 30 seconds)
- You need immediate results for user-facing features
- Small to medium batch sizes (up to ~100 items)

# Asynchronous Processing

Asynchronous calls return immediately with a batch ID. You then poll for results using the `async_batches/` endpoints.

### Async processing flow

1. **Submit job** -- `POST` to `async_process` or `async_batch_process`
2. **Receive batch ID** -- response contains `batch_id`
3. **Poll status** -- `HEAD` request to check readiness (lightweight, no body)
4. **Get results** -- `GET` request to retrieve processed results

## Async Single Item

Submit a single item for background processing. Use the `/single` endpoint to retrieve the result.

```
POST https://agents.botify.com/{org}/{project}/html_fetch/async_process
```

### Request body

```json
{
  "item": {
    "url": "<url>"
  }
}
```

### cURL example

```bash
# Step 1: Submit the job
curl -X POST "https://agents.botify.com/{org}/{project}/html_fetch/async_process" \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
  "item": {
    "url": "<url>"
  }
}'

# Response: { "batch_id": "abc123" }

# Step 2: Check if batch is ready (HEAD request -- lightweight check)
curl -I -X HEAD "https://agents.botify.com/{org}/{project}/html_fetch/async_batches/abc123" \
  -H "Authorization: Bearer $TOKEN"
# Returns 200 if ready, 204 if still processing

# Step 3: Get the result (for single-item async)
curl -X GET "https://agents.botify.com/{org}/{project}/html_fetch/async_batches/abc123/single" \
  -H "Authorization: Bearer $TOKEN"
# Returns 200 with result, 202 if still processing, 404 if batch not found
```

## Async Batch Processing

Submit multiple items for background processing. The request body uses **JSONL** (newline-delimited JSON) format.

```
POST https://agents.botify.com/{org}/{project}/html_fetch/async_batch_process
```

### Request body (JSONL, `Content-Type: text/plain`)

The first line contains `config` and `batch_config`. Each subsequent line is an item with a unique `id`.

```
{"config": {}, "batch_config": {}}
{"id": "item_1", "item": {"url": "<url>"}}
```

### Response

Returns a JSON object with `batch_id`, `status`, and a `Location` header.

### cURL example

```bash
# Step 1: Submit the batch job
curl -X POST "https://agents.botify.com/{org}/{project}/html_fetch/async_batch_process" \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: text/plain" \
  -d '{"config": {}, "batch_config": {}}\n{"id": "item_1", "item": {"url": "<url>"}}'

# Response: { "batch_id": "xyz789" }

# Step 2: Check if batch is ready (HEAD request)
curl -I -X HEAD "https://agents.botify.com/{org}/{project}/html_fetch/async_batches/xyz789" \
  -H "Authorization: Bearer $TOKEN"
# Returns 200 if ready, 204 if still processing

# Step 3: Get all results
curl -X GET "https://agents.botify.com/{org}/{project}/html_fetch/async_batches/xyz789" \
  -H "Authorization: Bearer $TOKEN"
```

## Checking Batch Status

Use a lightweight `HEAD` request to check if your batch is ready without transferring data.

```
HEAD https://agents.botify.com/{org}/{project}/html_fetch/async_batches/{batch_id}
```

```bash
curl -I -X HEAD "https://agents.botify.com/{org}/{project}/html_fetch/async_batches/{batch_id}" \
  -H "Authorization: Bearer $TOKEN"

# Response headers indicate status:
# HTTP/1.1 200 OK        -> Batch is complete, results ready
# HTTP/1.1 204 No Content -> Still processing
# HTTP/1.1 404 Not Found -> Batch does not exist
```

## Retrieving Results

Once the batch is ready, fetch results using the appropriate endpoint.

```bash
# For async_process (single item) -- use /single endpoint
curl -X GET "https://agents.botify.com/{org}/{project}/html_fetch/async_batches/{batch_id}/single" \
  -H "Authorization: Bearer $TOKEN"

# For async_batch_process (multiple items) -- use base endpoint
curl -X GET "https://agents.botify.com/{org}/{project}/html_fetch/async_batches/{batch_id}" \
  -H "Authorization: Bearer $TOKEN"
# Returns streaming JSONL where each line is a processed item or error
```

### Async response codes

| Code | Endpoint | Meaning | Action |
|------|----------|---------|--------|
| 200 | `HEAD` / `GET` | Batch complete, results ready | Read results from response body |
| 204 | `HEAD` | Still processing | Continue polling |
| 202 | `GET /single` | Still processing | Continue polling |
| 400 | `GET /single` | All items failed (client error) | Check error in response body |
| 404 | `HEAD` / `GET` | Batch does not exist | Verify `batch_id` is correct |

> **HEAD vs GET:** Use `HEAD` requests for lightweight status checks (no response body). Use `GET` only when ready to retrieve results to minimize bandwidth.

### When to use asynchronous calls

- Processing takes longer than 30 seconds
- Large batch sizes (100+ items)
- Background processing where immediate results aren't required
- Integration with job queues or workflow systems

## Python Example

A complete Python example showing both synchronous and asynchronous patterns.

```python
import requests
import time

API_TOKEN = "your_api_token"
BASE_URL = "https://agents.botify.com/{organization}/{project}"
AGENT = "html_fetch"

headers = {
    "Authorization": f"Bearer {API_TOKEN}",
    "Content-Type": "application/json"
}

single_payload = {
  "item": {
    "url": "<url>"
  }
}


def process_sync(payload: dict) -> dict:
    """Synchronous single-item processing."""
    response = requests.post(
        f"{BASE_URL}/{AGENT}/process",
        headers=headers,
        json=payload
    )
    response.raise_for_status()
    return response.json()


def process_async(
    payload: dict,
    poll_interval: int = 5,
    max_retries: int = 60
) -> dict:
    """Asynchronous single-item processing with polling."""
    # Submit job
    response = requests.post(
        f"{BASE_URL}/{AGENT}/async_process",
        headers=headers,
        json=payload
    )
    batch_id = response.json()["batch_id"]

    # Poll for results using HEAD (lightweight check)
    for _ in range(max_retries):
        status = requests.head(
            f"{BASE_URL}/{AGENT}/async_batches/{batch_id}",
            headers=headers
        )

        if status.status_code == 200:
            result = requests.get(
                f"{BASE_URL}/{AGENT}/async_batches/{batch_id}/single",
                headers=headers
            )
            return result.json()
        elif status.status_code == 204:
            time.sleep(poll_interval)
        else:
            raise Exception(f"Unexpected status: {status.status_code}")

    raise Exception("Max retries exceeded")


def process_async_batch(
    items: list,
    config: dict | None = None,
    poll_interval: int = 5
) -> dict:
    """Asynchronous batch processing with polling."""
    import json as _json

    # Build JSONL payload: first line is config, subsequent lines are items
    lines = [_json.dumps({"config": config or {}, "batch_config": {}})]
    for i, item in enumerate(items):
        lines.append(_json.dumps({"id": f"item_{i}", "item": item}))
    body = "\n".join(lines)

    # Submit batch (JSONL, text/plain)
    response = requests.post(
        f"{BASE_URL}/{AGENT}/async_batch_process",
        headers={**headers, "Content-Type": "text/plain"},
        data=body
    )
    batch_id = response.json()["batch_id"]

    # Poll until ready
    while True:
        status = requests.head(
            f"{BASE_URL}/{AGENT}/async_batches/{batch_id}",
            headers=headers
        )

        if status.status_code == 200:
            results = requests.get(
                f"{BASE_URL}/{AGENT}/async_batches/{batch_id}",
                headers=headers
            )
            return results.json()

        time.sleep(poll_interval)


# Example usage
# result = process_sync(single_payload)
# result = process_async(single_payload)
# results = process_async_batch([{"input": "item1"}, {"input": "item2"}])
```

## Billing

Fixed cost per URL requested. A page already in your SiteCrawler crawl is read from it and adds nothing. A live page fetch is billed only when you ask for one, or when the page is missing from the crawl and you allowed the fallback. Pages that refuse an ordinary fetch are retried on harder routes and every attempt is billed, so a protected page costs a live page fetch plus a protected-page fetch, and a hard-blocked one the whole sequence. Botify remembers for a day which route worked for a domain, so repeat fetches of the same site usually stay on the first.

These usage SKUs can be charged on a call.

| SKU | Credits | Description |
| --- | ------- | ----------- |
| Tool call | 1 per request | Charged once per successful item, on top of any usage below. |
| Live page fetch | 495 per 1,000 page | Fetches a page from the live web as an ordinary visitor would. |
| Live page fetch, hard-blocked page | 7,425 per 1,000 page | A page that refused every gentler attempt. Charged on top of them, so a hard-blocked page costs the whole sequence. |
| Live page fetch, protected page | 2,475 per 1,000 page | A page that refused an ordinary fetch and had to be retried. Charged on top of the ordinary attempt, not instead of it. |


## Schemas

### Item

```json
{
  "type": "object",
  "properties": {
    "url": {
      "description": "The URL to fetch HTML for. Temporary uploaded URLs on https://app.botify.com/:organization/:project/o/storage/tmp are allowed.",
      "title": "URL",
      "type": "string"
    },
    "origin": {
      "$ref": "#/$defs/OriginEnum",
      "default": "site_crawler",
      "description": "The origin of the URL",
      "enumNames": [
        "Botify SiteCrawler",
        "Real Time Fetch URL",
        "Rendering Farm",
        "Temporary upload"
      ],
      "title": "Origin"
    },
    "country_code": {
      "default": "us",
      "description": "The country code to use for the real-time fetch. 2 letters lowercase",
      "title": "Country Code",
      "type": "string"
    },
    "fallback_real_time_fetch": {
      "default": false,
      "description": "Fallback to real_time_fetch_url if the URL is not found in SiteCrawler",
      "title": "Fallback To Real Time Fetch",
      "type": "boolean"
    },
    "only_text": {
      "default": false,
      "description": "Only return the text of the page",
      "title": "Only Text",
      "type": "boolean"
    },
    "only_body": {
      "default": false,
      "description": "Only return the body of the page",
      "title": "Only Body",
      "type": "boolean"
    },
    "remove_tags": {
      "default": [
        "script",
        "noscript",
        "style",
        "header",
        "footer",
        "nav",
        "aside",
        "form",
        "svg",
        "iframe"
      ],
      "description": "Remove all tags from the HTML",
      "items": {
        "description": "The tag to remove",
        "title": "Tag",
        "type": "string"
      },
      "title": "Remove Tags",
      "type": "array"
    },
    "save_to_tmp_storage": {
      "default": false,
      "description": "Save the fetched HTML to Botify temporary storage and return the temporary URL instead of the HTML. Use this to save output tokens. The returned URL can be passed to other HTML-processing tools.",
      "title": "Save To Temporary Storage",
      "type": "boolean"
    }
  },
  "required": [
    "url"
  ],
  "$defs": {
    "OriginEnum": {
      "enum": [
        "site_crawler",
        "real_time_fetch_url",
        "rendering_farm",
        "temporary_upload"
      ],
      "title": "OriginEnum",
      "type": "string"
    }
  }
}
```

### Config

This agent uses the default config (no custom fields).

### Response

```json
{
  "$defs": {
    "OriginEnum": {
      "enum": [
        "site_crawler",
        "real_time_fetch_url",
        "rendering_farm",
        "temporary_upload"
      ],
      "title": "OriginEnum",
      "type": "string"
    }
  },
  "properties": {
    "origin": {
      "anyOf": [
        {
          "$ref": "#/$defs/OriginEnum"
        },
        {
          "type": "null"
        }
      ],
      "default": null,
      "description": "The origin of the HTML (site_crawler, real_time_fetch_url, rendering_farm, or temporary_upload)"
    },
    "crawl": {
      "anyOf": [
        {
          "type": "string"
        },
        {
          "type": "null"
        }
      ],
      "default": null,
      "description": "Crawl Slug from SiteCrawler",
      "title": "Crawl"
    },
    "html": {
      "anyOf": [
        {
          "type": "string"
        },
        {
          "type": "null"
        }
      ],
      "default": null,
      "description": "HTML content of the page. Omitted when `save_to_tmp_storage` is True; see `html_tmp_storage_url` instead.",
      "title": "Html"
    },
    "html_tmp_storage_url": {
      "anyOf": [
        {
          "type": "string"
        },
        {
          "type": "null"
        }
      ],
      "default": null,
      "description": "Temporary storage URL of the HTML page. Returned instead of `html` when `save_to_tmp_storage` is True. Can be passed to other HTML-processing tools.",
      "title": "Html Tmp Storage Url"
    },
    "status_code": {
      "anyOf": [
        {
          "type": "integer"
        },
        {
          "type": "null"
        }
      ],
      "default": null,
      "description": "HTTP status code of the response",
      "title": "Status Code"
    },
    "headers": {
      "anyOf": [
        {
          "additionalProperties": true,
          "type": "object"
        },
        {
          "type": "null"
        }
      ],
      "default": null,
      "description": "HTTP headers of the response",
      "title": "Headers"
    },
    "iframe_allowed": {
      "anyOf": [
        {
          "type": "boolean"
        },
        {
          "type": "null"
        }
      ],
      "default": null,
      "description": "Whether iframes are allowed in the response",
      "title": "Iframe Allowed"
    },
    "reason": {
      "anyOf": [
        {
          "type": "string"
        },
        {
          "type": "null"
        }
      ],
      "default": null,
      "description": "Reason for failure, if any",
      "title": "Reason"
    }
  },
  "title": "HTMLFetchProcessedItem",
  "type": "object"
}
```

## Best Practices

**Use exponential backoff for polling.** Start with short intervals (1-2 seconds) and increase the delay between polls to reduce API load.

**Set reasonable timeouts.** For synchronous calls, configure your HTTP client with appropriate timeout values (30-60 seconds).

**Handle rate limits gracefully.** Implement retry logic with backoff when you receive 429 (Too Many Requests) responses.

**Batch when possible.** Use batch endpoints to reduce the number of API calls and improve throughput.
