CrawlCheck

Quickstart

Ask before your agent acts

Before an agent reads, cites, connects to or transacts with a domain, CrawlCheck answers what the domain claims, what it actually serves, what is safe to do, and what evidence proves it, as one signed decision. Pick your stack; each block runs as written.

Decisions: allow proceed · warn proceed only if your policy accepts the stated doubt · require_confirmation a human approves first · block do not act · unsupported nothing to act on. A policy you send can make a decision stricter, never looser. Templates: safe_citation, safe_data_retrieval, safe_api_connect, safe_mcp_tool, safe_commerce, safe_oauth, support_escalation.

Jump to: Claude Code, Claude Desktop or any MCP client · Any language: one HTTP call · Python · OpenAI Agents SDK · Claude API tool use · LangChain / LangGraph · Cloudflare Worker or any fetch() runtime · TypeScript SDK

Claude Code, Claude Desktop or any MCP client

Add the server once. The tools preflight (a signed decision for an action), resolve_domain (the full signed answer) and mcp_servers (what a domain's MCP endpoints actually answered) are then available to the model.

claude mcp add --transport http crawlcheck https://crawlcheck.io/mcp

# or in an MCP client config file:
{ "mcpServers": { "crawlcheck": { "type": "http", "url": "https://crawlcheck.io/mcp" } } }

Any language: one HTTP call

Ask before the action. The answer is signed, so keep it with the action it authorised.

curl -s -X POST https://crawlcheck.io/api/v1/preflight \
  -H 'content-type: application/json' \
  -d '{"domain":"example.com","action":"cite","agent":"GPTBot","template":"safe_citation"}'

Python

No dependencies. guard() asks, verifies the receipt locally against the published key, and only returns proceed=True when the decision allows it.

pip install crawlcheck

from crawlcheck import CrawlCheck
cc = CrawlCheck()
proceed, receipt = cc.guard("example.com", action="read", agent="MyAgentBot")
if not proceed:
    print(receipt["decision"], [s["why"] for s in receipt["steps"]])

OpenAI Agents SDK (Python)

Expose preflight as a tool and tell the agent to call it before it reads, cites or acts on a site.

from agents import Agent, Runner, function_tool
from crawlcheck import CrawlCheck

cc = CrawlCheck()

@function_tool
def preflight(domain: str, action: str) -> str:
    """Ask CrawlCheck whether this action on this domain is allowed. action: read, cite, connect or transact."""
    r = cc.preflight(domain, action=action, agent="MyAgentBot")
    return r["decision"] + ": " + "; ".join(s["why"] for s in r["steps"])

agent = Agent(name="researcher", tools=[preflight],
              instructions="Before you read, cite or act on any website, call preflight. Do not proceed on block; ask the user on require_confirmation.")
print(Runner.run_sync(agent, "Is example.com a safe source to cite?").final_output)

Claude API tool use (Python)

The same tool, declared for the Messages API. Run the call when Claude asks for it and return the result.

import anthropic, json
from crawlcheck import CrawlCheck

cc, client = CrawlCheck(), anthropic.Anthropic()
tools = [{"name": "preflight", "description": "Ask CrawlCheck whether an action on a domain is allowed (read, cite, connect, transact).",
          "input_schema": {"type": "object", "properties": {"domain": {"type": "string"}, "action": {"type": "string"}}, "required": ["domain", "action"]}}]

def run_tool(block):
    r = cc.preflight(block.input["domain"], action=block.input["action"])
    return {"type": "tool_result", "tool_use_id": block.id, "content": json.dumps({"decision": r["decision"], "steps": r["steps"]})}

LangChain / LangGraph

A plain tool works in any LangChain agent or LangGraph node.

from langchain_core.tools import tool
from crawlcheck import CrawlCheck

cc = CrawlCheck()

@tool
def preflight(domain: str, action: str = "read") -> dict:
    """Ask CrawlCheck whether an action on a domain is allowed before taking it."""
    r = cc.preflight(domain, action=action)
    return {"decision": r["decision"], "why": [s["why"] for s in r["steps"]]}

Cloudflare Worker or any fetch() runtime

Gate an outbound fetch: ask first, fetch only on allow.

async function guardedFetch(url, agent = "MyAgentBot") {
  const domain = new URL(url).hostname;
  const r = await fetch("https://crawlcheck.io/api/v1/preflight", { method: "POST", headers: { "content-type": "application/json" },
    body: JSON.stringify({ domain, action: "read", agent }) }).then((x) => x.json());
  if (r.decision !== "allow") throw new Error("CrawlCheck: " + r.decision + " - " + r.steps.map((s) => s.why).join("; "));
  return fetch(url, { headers: { "user-agent": agent } });
}

TypeScript SDK

The typed client and the offline verifier for evidence bundles and receipts.

npm i @crawlcheck/sdk

What the answer rests on

Every decision names the resolve answer it was made from by its SHA-256, and both are signed with the key published at /.well-known/http-message-signatures-directory. Verify either one offline with crawlcheck-verify.mjs or python -m crawlcheck verify receipt.json. Reference: POST /api/v1/preflight · all endpoints · SDKs.