Getting started

CrawlSign is a crawler-safety engine that runs on your own infrastructure. It checks every page before it enters a dataset, an index or an agent's context, and returns one decision: continue, log, quarantine or stop, with stable reason codes that explain why.

The engine is licensed to design partners. The same detector is available three ways: a CLI, a Python SDK, and a local HTTP API for any other language. Pages, headers and results stay on your machines.

Install

Your pilot agreement comes with the CrawlSign package for your platform. CrawlSign needs Python 3.10 or newer. Install the core package, plus the extras for the interfaces you use:

pip install "./crawlsign-0.1.0-py3-none-any.whl"          # CLI, SDK, SafeFetcher, PoisonDetector
pip install "./crawlsign-0.1.0-py3-none-any.whl[api]"     # local HTTP API (crawlsign serve)
pip install "./crawlsign-0.1.0-py3-none-any.whl[warc]"    # WARC and Common Crawl batches

Other extras: mcp (agent tool server), scrapy, playwright, crawl4ai, firecrawl, langchain and llamaindex. If you received a private package index URL instead of a file, pass it with --index-url and install crawlsign by name.

First scan

crawlsign scan https://example.com
crawlsign analyze-file page.html --url https://example.com/page --format markdown

scan checks robots.txt first, then fetches with size, time and redirect limits, then analyzes. analyze-file analyzes HTML you already have, with no network access. Both print the result as JSON, and exit with code 2 when the decision is to stop.

{
  "url": "https://example.com/noai",
  "should_stop": true,
  "score": 80,
  "action": "stop",
  "reasons": ["ai_refusal_signal_detected"],
  "threshold": 70,
  "metadata": {"stop_scope": "url_or_path", "signature_list_version": 2}
}

From Python

from crawlsign import CrawlSignConfig, SafeFetcher

with SafeFetcher.with_pooled_session(CrawlSignConfig()) as fetcher:
    for url in urls:
        fetched = fetcher.fetch(url)
        if fetched["safe"]:
            ingest(fetched["content"])
        else:
            audit(url, fetched["result"]["action"], fetched["result"]["reasons"])

fetched["content"] is None unless the page is safe, so a quarantined or stopped page cannot be ingested by accident.

Four rules every integration follows

  1. Gate on action. Ingest only continue and log. A quarantine page has should_stop set to false (you may keep crawling the host) but must not be ingested. should_stop is the crawl decision; metadata.stop_scope says whether it covers the URL or the whole domain.
  2. Check at fetch time when you can. Only SafeFetcher (and the API's /v1/analyze/url) can enforce robots.txt, because they run before the request. Analyzing HTML someone else fetched still catches X-Robots-Tag, meta robots and noai, but not robots.txt.
  3. Use one fetcher or detector per crawl. The robots.txt cache, rate limits, budgets and domain-wide stop escalation live on the instance.
  4. A refusal is never routed around. There is no retry, user-agent rotation or bypass option. A refusal ends in stop, with an audit trail.

Try it without blocking anything

Start a pilot in monitor mode: CrawlSign still honors refusals, but pages flagged only by its risk detectors are logged instead of withheld, together with what enforcement would have done. crawlsign monitor-report turns a run into a report you can review before switching enforcement on.

Next