Reason codes

Every result lists the reasons behind its action as stable, machine-readable codes. They are a public contract: a code is never renamed or reused for a different meaning, so you can key alerts, dashboards and audit queries on them. Codes marked reserved are part of the contract but are not emitted yet.

Refusal signals

The site said no. Honored by default; see the respect settings in the configuration.

Code Meaning
robots_txt_disallow robots.txt disallowed crawling.
x_robots_tag_disallow X-Robots-Tag requested no crawling/indexing.
meta_robots_disallow HTML meta robots requested no crawling/indexing.
ai_refusal_signal_detected noai/noimageai or equivalent AI refusal signal detected.

Known signatures

The URL, path or page matches a curated poison or tarpit signature.

Code Meaning
known_poison_source_detected Known poison source or endpoint reference detected.
known_tarpit_reference_detected Known tarpit keyword/reference detected.

Links a human cannot see but a crawler would follow.

Code Meaning
hidden_link_detected One or more hidden links detected.
hidden_internal_link_cluster_detected Cluster of hidden internal links detected.

Enabled with detectors.link_maze.

Code Meaning
excessive_internal_links Page contains unusually many internal links.
randomized_link_maze_detected Page contains many random/UUID/hash-like internal paths.
crawler_trap_link_pattern_detected Reserved. Link pattern appears crawler-trap-like.
infinite_crawl_pattern_detected Reserved. Crawl graph suggests endless expansion.

Content (experimental, off by default)

Enabled with detectors.content_anomaly.

Code Meaning
bulk_generated_text_detected Large body of generated-looking text detected.
low_coherence_content_detected Text features suggest low coherence.
repeated_template_content_detected Repeated text templates detected.
same_url_high_content_variance Reserved. Same URL changed drastically across fetches.
crawler_specific_response_anomaly Reserved. Crawler UA receives suspiciously different response.

Fetch safety

Protections for the crawler itself, enforced in every mode.

Code Meaning
oversized_response_detected Response exceeds configured size limits.
streaming_tarpit_detected Slow or never-ending streaming behavior detected.
redirect_loop_detected Redirect loop detected.
crawl_budget_exhausted A per-host page/byte/redirect budget was exhausted; no further fetch was made.
fetch_failed The fetch failed (network error or non-retryable HTTP failure); no content was analyzed.
non_html_response_skipped Response declared a non-HTML/binary content type; body not read or analyzed.

Crawl control

Scope of a stop decision.

Code Meaning
crawl_stopped_url URL crawling stopped.
crawl_stopped_path_prefix Path-prefix crawling stopped.
crawl_stopped_domain Domain crawling stopped.
content_quarantined Reserved. Content stored outside the main index.

Refusal signals

Refusal directives (noindex, nofollow, none, noai, noimageai) are honored even when they are scoped to a specific crawler or AI bot rather than to all crawlers.

  • X-Robots-Tag: both global (noindex) and bot-scoped (googlebot: noindex) forms are refusals.
  • <meta> robots: name="robots" and the names of well-known crawlers and AI bots (for example googlebot, bingbot, gptbot, claudebot, ccbot, google-extended) are read. Other metadata such as description is never mined for directive-like words.
  • noindex and nofollow express one indexing-refusal intent and count once even when they appear in both the header and the meta tag; both x_robots_tag_disallow and meta_robots_disallow are still listed so the audit trail shows where each one was found.