HTTP API

crawlsign serve (the api extra) runs the engine as a local HTTP service, so crawlers in any language, other processes, or several workers can share one detector. It runs on your infrastructure; there is no hosted CrawlSign API.

crawlsign serve                       # http://127.0.0.1:8765
curl -s http://127.0.0.1:8765/v1/analyze/url \
  -H 'content-type: application/json' \
  -d '{"url": "https://example.com/"}'

127.0.0.1:8765 is the server you run, not an Apotropic endpoint: page content and results never leave your infrastructure. If the server runs on another host in your network, use its address in place of 127.0.0.1:8765 throughout this page.

The running server also serves its OpenAPI schema at /openapi.json and interactive docs at /docs; start it with --no-docs to turn both off.

Choosing an endpoint

  • POST /v1/analyze/url when the server can do the fetch. robots.txt is enforced before the request, and private, loopback, link-local and internal targets are refused.
  • POST /v1/analyze/html when your crawler already fetched the page. Send the response headers too, or X-Robots-Tag refusals are invisible. robots.txt is not checked here, so check it in your crawler.

Gate on result.action: ingest only continue and log.

Binding and authentication

The server binds to loopback by default and needs no token there. To expose it on another interface, pass --host with --allow-public and set CRAWLSIGN_API_TOKEN; the server refuses to start publicly without one. With a token set, every endpoint except /health requires Authorization: Bearer <token>, and a wrong token returns 401.

Shared state

One server process keeps one detector and fetcher per config file for its lifetime. Domain-stop escalation therefore works across requests and across both analyze endpoints, and robots.txt is cached per host. Editing a config file starts fresh state for it; restarting the server clears everything. Per-host rate limits and budgets in a config accumulate for the server's lifetime.

The fixture endpoints analyze the demo pages from the examples bundle and return an empty list when the server runs without it.

Endpoints

Method Path Purpose
POST /v1/analyze/url Fetch and analyze a URL
POST /v1/analyze/html Analyze submitted HTML
GET /health Check local API health
GET /v1/fixtures List demo fixtures
POST /v1/analyze/fixture Analyze a checked-in fixture

POST /v1/analyze/url

Fetch and analyze a URL. Fetches through SafeFetcher, preserves robots behavior, and validates URL targets.

Request body: UrlAnalyzeRequest

Status Body Meaning
200 AnalyzeResponse Successful Response
400 - Invalid, unsupported, or blocked URL/config input.
422 HTTPValidationError Validation Error

POST /v1/analyze/html

Analyze submitted HTML. Analyzes caller-supplied HTML without fetching network resources.

Request body: HtmlAnalyzeRequest

Status Body Meaning
200 AnalyzeResponse Successful Response
400 - Invalid config path or malformed config/signature file.
413 - Submitted HTML exceeds the selected config byte limit.
422 HTTPValidationError Validation Error

GET /health

Check local API health. Returns a small deterministic health payload for local service checks.

Status Body Meaning
200 HealthResponse Successful Response

GET /v1/fixtures

List demo fixtures. Returns checked-in deterministic MVP fixtures. Arbitrary fixture paths are not accepted.

Status Body Meaning
200 list of FixtureCatalogEntry Successful Response

POST /v1/analyze/fixture

Analyze a checked-in fixture. Analyzes one fixture from /v1/fixtures using the existing CrawlSign engine.

Request body: FixtureAnalyzeRequest

Status Body Meaning
200 AnalyzeResponse Successful Response
400 - Unknown or unsafe fixture ID.
422 HTTPValidationError Validation Error

Schemas

AnalyzeResponse

Field Type Required Description
fetch object or null no Optional URL-fetch summary.
fixture object or null no Optional fixture metadata.
quarantine_path string or null no Optional quarantine directory path.
result DetectionResultResponse yes Canonical CrawlSign result payload.
source one of fixture, html, url yes Input surface used for this analysis.
url string or null no Input or fetched URL.
view ResultView or null no Optional UI-ready summary, evidence, JSON, and Markdown fields.

DetectionResultResponse

Field Type Required Description
action one of continue, log, quarantine, stop yes MVP action band derived from score and evidence.
metadata object no Detector evidence metadata.
reasons list of string no Stable reason codes.
score integer yes CrawlSign risk score.
should_stop boolean yes Whether ingestion should stop for this result.
threshold integer yes Configured stop threshold.
url string yes Analyzed URL.

EvidenceItem

Field Type Required Description
label string yes Evidence item label.
value string yes Evidence item value.

EvidenceSection

Field Type Required Description
id string yes Stable section identifier.
items list of EvidenceItem yes Evidence items in this section.
title string yes Human-readable section title.

FixtureAnalyzeRequest

Field Type Required Description
fixture_id string yes Stable fixture ID from /v1/fixtures.
include_view boolean no Include UI-ready summary/evidence fields. Default true.

FixtureCatalogEntry

Field Type Required Description
config_path string yes Repository-relative config path for the fixture.
expected_action one of continue, log, quarantine, stop yes Expected MVP action for the fixture.
fixture_path string yes Repository-relative fixture path.
id string yes Stable fixture identifier accepted by /v1/analyze/fixture.
label string yes Human-readable fixture label.
url string yes URL used when analyzing the fixture.

HTTPValidationError

Field Type Required Description
detail list of ValidationError no

HealthResponse

Field Type Required Description
mvp boolean yes Whether this service exposes the MVP surface.
service string yes Stable service identifier.
status string yes Service health status.
version string or null no Installed package version when available.

HtmlAnalyzeRequest

Field Type Required Description
config_path string no Repository-relative config path under configs/. Default "configs/default.yaml".
headers object no Optional response headers used for X-Robots-Tag detection.
html string yes HTML content to analyze without network fetching.
include_view boolean no Include UI-ready summary/evidence fields. Default true.
url string yes Source URL associated with the submitted HTML.

ResultView

Field Type Required Description
evidence_sections list of EvidenceSection yes Evidence grouped by detector type.
markdown_report string yes Markdown-formatted detection report.
raw_json string yes JSON-formatted detection result.
raw_result object yes Raw detection result dict.
summary ViewSummary yes Summary fields for UI display.

UrlAnalyzeRequest

Field Type Required Description
config_path string no Repository-relative config path under configs/. Default "configs/default.yaml".
include_view boolean no Include UI-ready summary/evidence fields. Default true.
url string yes Absolute http or https URL to fetch and analyze.

ValidationError

Field Type Required Description
ctx object no
input any no
loc list of string or integer yes
msg string yes
type string yes

ViewSummary

Field Type Required Description
action one of continue, log, quarantine, stop yes MVP action band.
action_label string yes Human-readable action label.
recommended_next_step string yes Recommended next action.
score integer yes CrawlSign risk score.
severity string yes Severity identifier.
severity_label string yes Human-readable severity label.
should_stop boolean yes Whether ingestion should stop.
top_reason_codes list of string yes Top reason codes, up to 3.
url string yes Analyzed URL.