Reason codes
Every result lists the reasons behind its action as stable, machine-readable codes. They are a public contract: a code is never renamed or reused for a different meaning, so you can key alerts, dashboards and audit queries on them. Codes marked reserved are part of the contract but are not emitted yet.
Refusal signals
The site said no. Honored by default; see the respect settings in the configuration.
| Code | Meaning |
|---|---|
robots_txt_disallow |
robots.txt disallowed crawling. |
x_robots_tag_disallow |
X-Robots-Tag requested no crawling/indexing. |
meta_robots_disallow |
HTML meta robots requested no crawling/indexing. |
ai_refusal_signal_detected |
noai/noimageai or equivalent AI refusal signal detected. |
Known signatures
The URL, path or page matches a curated poison or tarpit signature.
| Code | Meaning |
|---|---|
known_poison_source_detected |
Known poison source or endpoint reference detected. |
known_tarpit_reference_detected |
Known tarpit keyword/reference detected. |
Hidden links
Links a human cannot see but a crawler would follow.
| Code | Meaning |
|---|---|
hidden_link_detected |
One or more hidden links detected. |
hidden_internal_link_cluster_detected |
Cluster of hidden internal links detected. |
Link structure (experimental, off by default)
Enabled with detectors.link_maze.
| Code | Meaning |
|---|---|
excessive_internal_links |
Page contains unusually many internal links. |
randomized_link_maze_detected |
Page contains many random/UUID/hash-like internal paths. |
crawler_trap_link_pattern_detected |
Reserved. Link pattern appears crawler-trap-like. |
infinite_crawl_pattern_detected |
Reserved. Crawl graph suggests endless expansion. |
Content (experimental, off by default)
Enabled with detectors.content_anomaly.
| Code | Meaning |
|---|---|
bulk_generated_text_detected |
Large body of generated-looking text detected. |
low_coherence_content_detected |
Text features suggest low coherence. |
repeated_template_content_detected |
Repeated text templates detected. |
same_url_high_content_variance |
Reserved. Same URL changed drastically across fetches. |
crawler_specific_response_anomaly |
Reserved. Crawler UA receives suspiciously different response. |
Fetch safety
Protections for the crawler itself, enforced in every mode.
| Code | Meaning |
|---|---|
oversized_response_detected |
Response exceeds configured size limits. |
streaming_tarpit_detected |
Slow or never-ending streaming behavior detected. |
redirect_loop_detected |
Redirect loop detected. |
crawl_budget_exhausted |
A per-host page/byte/redirect budget was exhausted; no further fetch was made. |
fetch_failed |
The fetch failed (network error or non-retryable HTTP failure); no content was analyzed. |
non_html_response_skipped |
Response declared a non-HTML/binary content type; body not read or analyzed. |
Crawl control
Scope of a stop decision.
| Code | Meaning |
|---|---|
crawl_stopped_url |
URL crawling stopped. |
crawl_stopped_path_prefix |
Path-prefix crawling stopped. |
crawl_stopped_domain |
Domain crawling stopped. |
content_quarantined |
Reserved. Content stored outside the main index. |
Refusal signals
Refusal directives (noindex, nofollow, none, noai, noimageai) are honored even when they
are scoped to a specific crawler or AI bot rather than to all crawlers.
X-Robots-Tag: both global (noindex) and bot-scoped (googlebot: noindex) forms are refusals.<meta>robots:name="robots"and the names of well-known crawlers and AI bots (for examplegooglebot,bingbot,gptbot,claudebot,ccbot,google-extended) are read. Other metadata such asdescriptionis never mined for directive-like words.noindexandnofollowexpress one indexing-refusal intent and count once even when they appear in both the header and the meta tag; bothx_robots_tag_disallowandmeta_robots_disalloware still listed so the audit trail shows where each one was found.