Integrations
Where the check happens decides what it can catch. Every integration below states its check
point, whether robots.txt is enforced, and what to gate on.
| Check point | What CrawlSign sees | Enforces robots.txt? |
|---|---|---|
| Fetch time | The request, the raw HTML and the real response headers | Yes, before the page is requested |
| Post-fetch raw HTML | Raw HTML and headers that another tool fetched | No |
| Post-extraction text | Plain text from a loader | No; hidden links, meta tags and headers are already gone |
Integrations marked "example" ship as source files in the examples bundle that comes with your package: copy the file into your project and adapt it.
requests: SafeFetcher
The reference integration. One shared fetcher checks robots.txt before every page, caps the
response size, records redirects and owns the detector.
from crawlsign import CrawlSignConfig, SafeFetcher
kept, skipped = [], []
with SafeFetcher.with_pooled_session(CrawlSignConfig()) as fetcher:
for url in urls:
fetched = fetcher.fetch(url)
if fetched["safe"]:
kept.append(fetched["content"])
else:
skipped.append((url, fetched["result"]["action"], fetched["result"]["reasons"]))
- Check point: fetch time. Domain escalation: yes; the fetcher owns one detector.
- Gate:
fetched["safe"]. PassCrawlSignConfig.from_file("crawlsign.yaml")for your own configuration.
httpx and other HTTP clients (example)
Your client does the request, so ask a shared SafeFetcher about robots.txt first and reuse
its detector for every page:
def guarded_get(client, url, fetcher):
robots_stop = fetcher.check_robots(url)
if robots_stop is not None:
return {"safe": False, "result": robots_stop.to_dict(), "html": None}
response = client.get(url)
result = fetcher.detector.analyze(url=str(response.url), html=response.text, headers=dict(response.headers))
safe = result.action in {"continue", "log"}
return {"safe": safe, "result": result.to_dict(), "html": response.text if safe else None}
The httpx_crawler.py example adds the response size cap.
Playwright (example)
Analysis runs on the rendered DOM, so a noai tag added by JavaScript is caught, and the main
response's headers are passed so X-Robots-Tag works. robots.txt is checked before the browser
navigates.
from playwright.async_api import async_playwright
from crawlsign import SafeFetcher
from playwright_guard import guarded_goto # from the examples bundle
async def crawl(urls):
fetcher = SafeFetcher() # one per crawl: robots cache and domain escalation
pages = []
async with async_playwright() as playwright:
browser = await playwright.chromium.launch()
page = await browser.new_page()
for url in urls:
outcome = await guarded_goto(page, url, fetcher)
if outcome["safe"]:
pages.append(outcome["html"])
await browser.close()
return pages
Scrapy (example)
The downloader middleware sees responses, so keep Scrapy's own ROBOTSTXT_OBEY on; CrawlSign's
refusal handling (X-Robots-Tag, meta robots, noai) is added on top. A quarantine or
stop response raises IgnoreRequest, so the spider never sees it.
# settings.py
ROBOTSTXT_OBEY = True
DOWNLOADER_MIDDLEWARES = {
"myproject.middlewares.CrawlSignDownloaderMiddleware": 543, # scrapy_middleware.py example
}
Per-action and per-reason counts appear in Scrapy's stats as crawlsign/action/<action> and
crawlsign/reason/<code>.
Crawl4AI (example)
The guard checks robots.txt before Crawl4AI navigates, analyzes the rendered HTML with the real
response headers, and returns markdown only for pages CrawlSign allows. CrawlSign does not use
Crawl4AI's stealth options: a refusal is a reason to stop, not a block to get around.
from crawl4ai import AsyncWebCrawler
from crawlsign import SafeFetcher
from crawl4ai_guard import guarded_arun # from the examples bundle
async def crawl(urls):
fetcher = SafeFetcher()
pages = []
async with AsyncWebCrawler() as crawler:
for url in urls:
outcome = await guarded_arun(crawler, url, fetcher)
if outcome["safe"]:
pages.append(outcome["markdown"])
return pages
Firecrawl (example)
Firecrawl fetches and renders the page, so CrawlSign analyzes the rawHtml it returns before you
use the markdown. Two limits: Firecrawl does the request, so the optional robots.txt pre-check
uses your user agent rather than Firecrawl's, and Firecrawl returns no response headers, so a
refusal sent only in X-Robots-Tag is invisible. Meta robots and noai tags in the HTML are
still checked.
from firecrawl import Firecrawl
from crawlsign import SafeFetcher
from firecrawl_guard import guarded_scrape # from the examples bundle
firecrawl = Firecrawl(api_key="fc-...")
fetcher = SafeFetcher()
outcome = guarded_scrape(firecrawl, "https://example.com/", fetcher.detector, fetcher=fetcher)
if outcome["safe"]:
page_html = outcome["html"] # None for quarantine and stop
LangChain
CrawlSignWebLoader fetches each URL through one shared SafeFetcher, analyzes the raw HTML, and
only then extracts text. Pages that are not safe never become documents; they are listed, without
content, in loader.skipped. Install the langchain extra. Experimental.
from crawlsign.integrations.langchain import CrawlSignWebLoader
loader = CrawlSignWebLoader(urls)
docs = loader.load() # metadata: crawlsign_action, crawlsign_score, crawlsign_reasons, ...
audit = loader.skipped # url, status_code, result, quarantine_path
Pass config= or a shared fetcher= to change thresholds or share robots.txt caching across
loaders, and extractor= to replace text extraction (it only ever sees approved HTML).
LlamaIndex
CrawlSignWebReader is the equivalent reader (the llamaindex extra). The crawlsign_*
metadata keys are excluded from embeddings and prompts. Experimental.
from crawlsign.integrations.llamaindex import CrawlSignWebReader
reader = CrawlSignWebReader()
documents = reader.load_data(urls)
audit = reader.skipped
Local HTTP API and CrawlSignClient
For other languages, other processes, or one service shared by several crawlers. Start
crawlsign serve (the api extra) on your own infrastructure; it listens on 127.0.0.1:8765
by default. Point clients at wherever you run it. See the HTTP API reference.
from crawlsign import CrawlSignClient
with CrawlSignClient("http://127.0.0.1:8765") as client:
# HTML you already fetched: send the response headers too. robots.txt is not checked here.
checked = client.analyze_html(url="https://example.com/page", html=html, headers=headers)
# The server fetches (robots.txt enforced, private addresses blocked) and analyzes.
live = client.analyze_url("https://example.com/")
safe = live.result.action in {"continue", "log"}
Node and Crawlee (example)
A Node crawler calls the local API. The node_crawlee_client.mjs example is a dependency-free
client (Node 18+) with checkUrl, checkHtml and isSafe:
import { checkHtml, isSafe } from "./node_crawlee_client.mjs";
// In a Crawlee requestHandler, after the response arrives:
const result = await checkHtml(request.loadedUrl, body, new Headers(response.headers));
if (!isSafe(result)) {
return; // quarantine or stop: drop the page, keep result.reasons, do not retry
}
/v1/analyze/html cannot check robots.txt, so keep Crawlee's respectRobotsTxtFile on, or
let the server fetch with /v1/analyze/url.
MCP server for agents
crawlsign mcp runs an MCP server on stdio (the mcp extra) so an agent can check a page before
it uses it. Experimental.
| Tool | What it does |
|---|---|
check_url(url) |
Fetches through SafeFetcher and returns the decision: safe, action, score, reasons, stop_scope. Never returns page content. |
read_page(url) |
The same check, plus the page text only when the page is safe. |
analyze_html(html, url, headers?) |
Offline analysis of HTML the agent already has. |
{
"mcpServers": {
"crawlsign": { "command": "crawlsign", "args": ["mcp"] }
}
}
No tool returns the content of a page CrawlSign said to stop on, and the server instructs the
agent not to retry a refused page. Private-network targets are refused unless you pass
--allow-private-targets.
WARC files and Common Crawl batches
crawlsign analyze-warc runs every stored HTTP response in a WARC file (plain or .gz) through
one detector. A WARC keeps response headers, so X-Robots-Tag and noai checks work offline;
robots.txt cannot be enforced because the pages are already fetched. Needs the warc extra.
crawlsign analyze-warc crawl.warc.gz --output report.jsonl
crawlsign analyze-warc crawl.warc.gz --config crawlsign.yaml --limit 1000
Each JSON Lines row has url, status_code, content_type, action, score, should_stop,
reasons, stop_scope and ignored_refusal_signals, and a summary of action and reason counts
is printed to stderr. From Python:
from crawlsign import PoisonDetector
from crawlsign.warc import analyze_warc
detector = PoisonDetector() # one detector per file: domain escalation applies
stopped = [r["url"] for r in analyze_warc("crawl.warc.gz", detector) if r["result"]["should_stop"]]