Configuration

Every interface takes the same YAML file: --config on the CLI, CrawlSignConfig.from_file(path) in Python, and config_path on the HTTP API. Unknown keys are rejected, so a typo fails loudly instead of silently falling back to a default. Check a file with:

crawlsign validate-config crawlsign.yaml

Bundled profiles

Profile Use it when
default Standard crawling with the default thresholds.
audit-only Monitor mode: see what CrawlSign would do before enforcing it.
strict-enterprise Lower thresholds; quarantine on minimal hidden-link evidence.

On the HTTP API, config_path: "configs/default.yaml" (and the other two names) resolves to the bundled copy when the server has no configs/ directory of its own.

Mode

mode: monitor   # or enforce (the default)

enforce acts on every detector. monitor still enforces refusals and crawler-safety stops, but pages flagged only by risk detectors are logged with the would-be action in metadata.monitor_action. See monitor mode.

Thresholds

The summed score is bucketed into an action: at or above stop_url_score is stop, at or above quarantine_score is quarantine, at or above log_score is log, anything lower is continue. Escalating a stop to the whole domain needs stop_domain_score and at least two such URLs on the host. The values must satisfy log_score <= quarantine_score <= stop_url_score <= stop_domain_score.

Refusal handling

respect turns each refusal signal on or off: robots_txt, noai, noimageai, noindex and nofollow. All are on by default. A signal seen while its setting is off is still recorded, in metadata.ignored_refusal_signals, so the decision to ignore it is auditable.

Refusal directives are honored even when they name a specific crawler (googlebot: noindex, <meta name="gptbot" content="noai">): CrawlSign does not assume its own user agent is exempt.

Fetch limits, rate limits and budgets

crawler sets the user agent, connect, read and total deadlines, a minimum-throughput floor against slow-drip tarpits, the response size cap and retries. rate_limit adds a per-host delay and honors Retry-After. budgets caps pages, bytes and redirect hops per host; an exhausted budget returns stop with crawl_budget_exhausted without making the request. Rate limits and budgets live on the fetcher, so reuse one fetcher for the whole crawl.

Detectors and signatures

detectors.link_maze and detectors.content_anomaly are experimental and off by default. signatures.file points to your own signature list in YAML; when set, it replaces the bundled list. Ask us before writing one: signatures that match ordinary prose are the most common cause of false positives.

Quarantine

quarantine.enabled and quarantine.destination control where withheld pages and their evidence are written on the fetch path.

The default profile

This is the bundled default profile, which also documents every key:

crawler:
  user_agent: "CrawlSignCrawler/0.1 (+responsible crawler; respects robots.txt)"
  timeout_seconds: 10          # per-read socket timeout (seconds)
  connect_timeout_seconds: 10  # TCP/TLS handshake timeout (seconds)
  total_deadline_seconds: 30   # whole-response slow-stream deadline; 0 disables
  min_bytes_per_second: 0      # minimum sustained throughput; 0 disables
  connection_pool_size: 10     # bounded pool size for SafeFetcher.with_pooled_session
  max_retries: 2               # GET retries for 429/5xx + connection errors (pooled session)
  max_response_bytes: 5242880

respect:
  robots_txt: true
  noai: true
  noimageai: true
  noindex: true
  nofollow: true

thresholds:
  # Score >= log_score is recorded as a "log" action (below quarantine_score).
  log_score: 30
  quarantine_score: 50
  stop_url_score: 70
  stop_domain_score: 90

limits:
  max_links_per_page: 250
  max_hidden_links_per_page: 10
  max_words_per_page_before_anomaly_check: 2500

rate_limit:
  per_host_delay_seconds: 0.0          # minimum delay between page fetches to a host; 0 disables
  respect_retry_after: true            # honor Retry-After for the next request to a host
  max_concurrent_requests_per_host: 0  # cap in-flight requests per host for threaded callers; 0 = unbounded

budgets:
  per_host_pages: 0      # max page fetches per host; 0 = unlimited
  per_host_bytes: 0      # max cumulative response bytes per host; 0 = unlimited
  per_host_redirects: 0  # max cumulative redirect hops per host; 0 = unlimited

detectors:
  link_maze: false
  content_anomaly: false

quarantine:
  enabled: true
  destination: "./quarantine"