Integrations

Where the check happens decides what it can catch. Every integration below states its check point, whether robots.txt is enforced, and what to gate on.

Check point What CrawlSign sees Enforces robots.txt?
Fetch time The request, the raw HTML and the real response headers Yes, before the page is requested
Post-fetch raw HTML Raw HTML and headers that another tool fetched No
Post-extraction text Plain text from a loader No; hidden links, meta tags and headers are already gone

Integrations marked "example" ship as source files in the examples bundle that comes with your package: copy the file into your project and adapt it.

requests: SafeFetcher

The reference integration. One shared fetcher checks robots.txt before every page, caps the response size, records redirects and owns the detector.

from crawlsign import CrawlSignConfig, SafeFetcher

kept, skipped = [], []
with SafeFetcher.with_pooled_session(CrawlSignConfig()) as fetcher:
    for url in urls:
        fetched = fetcher.fetch(url)
        if fetched["safe"]:
            kept.append(fetched["content"])
        else:
            skipped.append((url, fetched["result"]["action"], fetched["result"]["reasons"]))
  • Check point: fetch time. Domain escalation: yes; the fetcher owns one detector.
  • Gate: fetched["safe"]. Pass CrawlSignConfig.from_file("crawlsign.yaml") for your own configuration.

httpx and other HTTP clients (example)

Your client does the request, so ask a shared SafeFetcher about robots.txt first and reuse its detector for every page:

def guarded_get(client, url, fetcher):
    robots_stop = fetcher.check_robots(url)
    if robots_stop is not None:
        return {"safe": False, "result": robots_stop.to_dict(), "html": None}
    response = client.get(url)
    result = fetcher.detector.analyze(url=str(response.url), html=response.text, headers=dict(response.headers))
    safe = result.action in {"continue", "log"}
    return {"safe": safe, "result": result.to_dict(), "html": response.text if safe else None}

The httpx_crawler.py example adds the response size cap.

Playwright (example)

Analysis runs on the rendered DOM, so a noai tag added by JavaScript is caught, and the main response's headers are passed so X-Robots-Tag works. robots.txt is checked before the browser navigates.

from playwright.async_api import async_playwright

from crawlsign import SafeFetcher
from playwright_guard import guarded_goto  # from the examples bundle


async def crawl(urls):
    fetcher = SafeFetcher()  # one per crawl: robots cache and domain escalation
    pages = []
    async with async_playwright() as playwright:
        browser = await playwright.chromium.launch()
        page = await browser.new_page()
        for url in urls:
            outcome = await guarded_goto(page, url, fetcher)
            if outcome["safe"]:
                pages.append(outcome["html"])
        await browser.close()
    return pages

Scrapy (example)

The downloader middleware sees responses, so keep Scrapy's own ROBOTSTXT_OBEY on; CrawlSign's refusal handling (X-Robots-Tag, meta robots, noai) is added on top. A quarantine or stop response raises IgnoreRequest, so the spider never sees it.

# settings.py
ROBOTSTXT_OBEY = True
DOWNLOADER_MIDDLEWARES = {
    "myproject.middlewares.CrawlSignDownloaderMiddleware": 543,  # scrapy_middleware.py example
}

Per-action and per-reason counts appear in Scrapy's stats as crawlsign/action/<action> and crawlsign/reason/<code>.

Crawl4AI (example)

The guard checks robots.txt before Crawl4AI navigates, analyzes the rendered HTML with the real response headers, and returns markdown only for pages CrawlSign allows. CrawlSign does not use Crawl4AI's stealth options: a refusal is a reason to stop, not a block to get around.

from crawl4ai import AsyncWebCrawler

from crawlsign import SafeFetcher
from crawl4ai_guard import guarded_arun  # from the examples bundle


async def crawl(urls):
    fetcher = SafeFetcher()
    pages = []
    async with AsyncWebCrawler() as crawler:
        for url in urls:
            outcome = await guarded_arun(crawler, url, fetcher)
            if outcome["safe"]:
                pages.append(outcome["markdown"])
    return pages

Firecrawl (example)

Firecrawl fetches and renders the page, so CrawlSign analyzes the rawHtml it returns before you use the markdown. Two limits: Firecrawl does the request, so the optional robots.txt pre-check uses your user agent rather than Firecrawl's, and Firecrawl returns no response headers, so a refusal sent only in X-Robots-Tag is invisible. Meta robots and noai tags in the HTML are still checked.

from firecrawl import Firecrawl

from crawlsign import SafeFetcher
from firecrawl_guard import guarded_scrape  # from the examples bundle

firecrawl = Firecrawl(api_key="fc-...")
fetcher = SafeFetcher()
outcome = guarded_scrape(firecrawl, "https://example.com/", fetcher.detector, fetcher=fetcher)
if outcome["safe"]:
    page_html = outcome["html"]  # None for quarantine and stop

LangChain

CrawlSignWebLoader fetches each URL through one shared SafeFetcher, analyzes the raw HTML, and only then extracts text. Pages that are not safe never become documents; they are listed, without content, in loader.skipped. Install the langchain extra. Experimental.

from crawlsign.integrations.langchain import CrawlSignWebLoader

loader = CrawlSignWebLoader(urls)
docs = loader.load()  # metadata: crawlsign_action, crawlsign_score, crawlsign_reasons, ...
audit = loader.skipped  # url, status_code, result, quarantine_path

Pass config= or a shared fetcher= to change thresholds or share robots.txt caching across loaders, and extractor= to replace text extraction (it only ever sees approved HTML).

LlamaIndex

CrawlSignWebReader is the equivalent reader (the llamaindex extra). The crawlsign_* metadata keys are excluded from embeddings and prompts. Experimental.

from crawlsign.integrations.llamaindex import CrawlSignWebReader

reader = CrawlSignWebReader()
documents = reader.load_data(urls)
audit = reader.skipped

Local HTTP API and CrawlSignClient

For other languages, other processes, or one service shared by several crawlers. Start crawlsign serve (the api extra) on your own infrastructure; it listens on 127.0.0.1:8765 by default. Point clients at wherever you run it. See the HTTP API reference.

from crawlsign import CrawlSignClient

with CrawlSignClient("http://127.0.0.1:8765") as client:
    # HTML you already fetched: send the response headers too. robots.txt is not checked here.
    checked = client.analyze_html(url="https://example.com/page", html=html, headers=headers)
    # The server fetches (robots.txt enforced, private addresses blocked) and analyzes.
    live = client.analyze_url("https://example.com/")
    safe = live.result.action in {"continue", "log"}

Node and Crawlee (example)

A Node crawler calls the local API. The node_crawlee_client.mjs example is a dependency-free client (Node 18+) with checkUrl, checkHtml and isSafe:

import { checkHtml, isSafe } from "./node_crawlee_client.mjs";

// In a Crawlee requestHandler, after the response arrives:
const result = await checkHtml(request.loadedUrl, body, new Headers(response.headers));
if (!isSafe(result)) {
  return; // quarantine or stop: drop the page, keep result.reasons, do not retry
}

/v1/analyze/html cannot check robots.txt, so keep Crawlee's respectRobotsTxtFile on, or let the server fetch with /v1/analyze/url.

MCP server for agents

crawlsign mcp runs an MCP server on stdio (the mcp extra) so an agent can check a page before it uses it. Experimental.

Tool What it does
check_url(url) Fetches through SafeFetcher and returns the decision: safe, action, score, reasons, stop_scope. Never returns page content.
read_page(url) The same check, plus the page text only when the page is safe.
analyze_html(html, url, headers?) Offline analysis of HTML the agent already has.
{
  "mcpServers": {
    "crawlsign": { "command": "crawlsign", "args": ["mcp"] }
  }
}

No tool returns the content of a page CrawlSign said to stop on, and the server instructs the agent not to retry a refused page. Private-network targets are refused unless you pass --allow-private-targets.

WARC files and Common Crawl batches

crawlsign analyze-warc runs every stored HTTP response in a WARC file (plain or .gz) through one detector. A WARC keeps response headers, so X-Robots-Tag and noai checks work offline; robots.txt cannot be enforced because the pages are already fetched. Needs the warc extra.

crawlsign analyze-warc crawl.warc.gz --output report.jsonl
crawlsign analyze-warc crawl.warc.gz --config crawlsign.yaml --limit 1000

Each JSON Lines row has url, status_code, content_type, action, score, should_stop, reasons, stop_scope and ignored_refusal_signals, and a summary of action and reason counts is printed to stderr. From Python:

from crawlsign import PoisonDetector
from crawlsign.warc import analyze_warc

detector = PoisonDetector()  # one detector per file: domain escalation applies
stopped = [r["url"] for r in analyze_warc("crawl.warc.gz", detector) if r["result"]["should_stop"]]