Getting started
CrawlSign is a crawler-safety engine that runs on your own infrastructure. It checks every page
before it enters a dataset, an index or an agent's context, and returns one decision: continue,
log, quarantine or stop, with stable reason codes that explain why.
The engine is licensed to design partners. The same detector is available three ways: a CLI, a Python SDK, and a local HTTP API for any other language. Pages, headers and results stay on your machines.
Install
Your pilot agreement comes with the CrawlSign package for your platform. CrawlSign needs Python 3.10 or newer. Install the core package, plus the extras for the interfaces you use:
pip install "./crawlsign-0.1.0-py3-none-any.whl" # CLI, SDK, SafeFetcher, PoisonDetector
pip install "./crawlsign-0.1.0-py3-none-any.whl[api]" # local HTTP API (crawlsign serve)
pip install "./crawlsign-0.1.0-py3-none-any.whl[warc]" # WARC and Common Crawl batches
Other extras: mcp (agent tool server), scrapy, playwright, crawl4ai, firecrawl,
langchain and llamaindex. If you received a private package index URL instead of a file,
pass it with --index-url and install crawlsign by name.
First scan
crawlsign scan https://example.com
crawlsign analyze-file page.html --url https://example.com/page --format markdown
scan checks robots.txt first, then fetches with size, time and redirect limits, then
analyzes. analyze-file analyzes HTML you already have, with no network access. Both print the
result as JSON, and exit with code 2 when the decision is to stop.
{
"url": "https://example.com/noai",
"should_stop": true,
"score": 80,
"action": "stop",
"reasons": ["ai_refusal_signal_detected"],
"threshold": 70,
"metadata": {"stop_scope": "url_or_path", "signature_list_version": 2}
}
From Python
from crawlsign import CrawlSignConfig, SafeFetcher
with SafeFetcher.with_pooled_session(CrawlSignConfig()) as fetcher:
for url in urls:
fetched = fetcher.fetch(url)
if fetched["safe"]:
ingest(fetched["content"])
else:
audit(url, fetched["result"]["action"], fetched["result"]["reasons"])
fetched["content"] is None unless the page is safe, so a quarantined or stopped page cannot
be ingested by accident.
Four rules every integration follows
- Gate on
action. Ingest onlycontinueandlog. Aquarantinepage hasshould_stopset tofalse(you may keep crawling the host) but must not be ingested.should_stopis the crawl decision;metadata.stop_scopesays whether it covers the URL or the whole domain. - Check at fetch time when you can. Only
SafeFetcher(and the API's/v1/analyze/url) can enforcerobots.txt, because they run before the request. Analyzing HTML someone else fetched still catchesX-Robots-Tag, meta robots andnoai, but notrobots.txt. - Use one fetcher or detector per crawl. The
robots.txtcache, rate limits, budgets and domain-wide stop escalation live on the instance. - A refusal is never routed around. There is no retry, user-agent rotation or bypass option.
A refusal ends in
stop, with an audit trail.
Try it without blocking anything
Start a pilot in monitor mode: CrawlSign still honors refusals, but
pages flagged only by its risk detectors are logged instead of withheld, together with what
enforcement would have done. crawlsign monitor-report turns a run into a report you can review
before switching enforcement on.
Next
- Integrations: Scrapy, Playwright, Crawl4AI, Firecrawl, LangChain, LlamaIndex, Node, MCP and WARC.
- Results and monitor mode: every field of a result.
- Configuration: thresholds, refusal handling and limits.
- HTTP API and reason codes.