Configuration
Every interface takes the same YAML file: --config on the CLI,
CrawlSignConfig.from_file(path) in Python, and config_path on the HTTP API. Unknown keys are
rejected, so a typo fails loudly instead of silently falling back to a default. Check a file with:
crawlsign validate-config crawlsign.yaml
Bundled profiles
| Profile | Use it when |
|---|---|
default |
Standard crawling with the default thresholds. |
audit-only |
Monitor mode: see what CrawlSign would do before enforcing it. |
strict-enterprise |
Lower thresholds; quarantine on minimal hidden-link evidence. |
On the HTTP API, config_path: "configs/default.yaml" (and the other two names) resolves to the
bundled copy when the server has no configs/ directory of its own.
Mode
mode: monitor # or enforce (the default)
enforce acts on every detector. monitor still enforces refusals and crawler-safety stops, but
pages flagged only by risk detectors are logged with the would-be action in
metadata.monitor_action. See monitor mode.
Thresholds
The summed score is bucketed into an action: at or above stop_url_score is stop, at or above
quarantine_score is quarantine, at or above log_score is log, anything lower is
continue. Escalating a stop to the whole domain needs stop_domain_score and at least two such
URLs on the host. The values must satisfy
log_score <= quarantine_score <= stop_url_score <= stop_domain_score.
Refusal handling
respect turns each refusal signal on or off: robots_txt, noai, noimageai, noindex and
nofollow. All are on by default. A signal seen while its setting is off is still recorded, in
metadata.ignored_refusal_signals, so the decision to ignore it is auditable.
Refusal directives are honored even when they name a specific crawler (googlebot: noindex,
<meta name="gptbot" content="noai">): CrawlSign does not assume its own user agent is exempt.
Fetch limits, rate limits and budgets
crawler sets the user agent, connect, read and total deadlines, a minimum-throughput floor
against slow-drip tarpits, the response size cap and retries. rate_limit adds a per-host delay
and honors Retry-After. budgets caps pages, bytes and redirect hops per host; an exhausted
budget returns stop with crawl_budget_exhausted without making the request. Rate limits and
budgets live on the fetcher, so reuse one fetcher for the whole crawl.
Detectors and signatures
detectors.link_maze and detectors.content_anomaly are experimental and off by default.
signatures.file points to your own signature list in YAML; when set, it replaces the bundled
list. Ask us before writing one: signatures that match ordinary prose are the most common cause of
false positives.
Quarantine
quarantine.enabled and quarantine.destination control where withheld pages and their evidence
are written on the fetch path.
The default profile
This is the bundled default profile, which also documents every key:
crawler:
user_agent: "CrawlSignCrawler/0.1 (+responsible crawler; respects robots.txt)"
timeout_seconds: 10 # per-read socket timeout (seconds)
connect_timeout_seconds: 10 # TCP/TLS handshake timeout (seconds)
total_deadline_seconds: 30 # whole-response slow-stream deadline; 0 disables
min_bytes_per_second: 0 # minimum sustained throughput; 0 disables
connection_pool_size: 10 # bounded pool size for SafeFetcher.with_pooled_session
max_retries: 2 # GET retries for 429/5xx + connection errors (pooled session)
max_response_bytes: 5242880
respect:
robots_txt: true
noai: true
noimageai: true
noindex: true
nofollow: true
thresholds:
# Score >= log_score is recorded as a "log" action (below quarantine_score).
log_score: 30
quarantine_score: 50
stop_url_score: 70
stop_domain_score: 90
limits:
max_links_per_page: 250
max_hidden_links_per_page: 10
max_words_per_page_before_anomaly_check: 2500
rate_limit:
per_host_delay_seconds: 0.0 # minimum delay between page fetches to a host; 0 disables
respect_retry_after: true # honor Retry-After for the next request to a host
max_concurrent_requests_per_host: 0 # cap in-flight requests per host for threaded callers; 0 = unbounded
budgets:
per_host_pages: 0 # max page fetches per host; 0 = unlimited
per_host_bytes: 0 # max cumulative response bytes per host; 0 = unlimited
per_host_redirects: 0 # max cumulative redirect hops per host; 0 = unlimited
detectors:
link_maze: false
content_anomaly: false
quarantine:
enabled: true
destination: "./quarantine"