CrawlSign / Crawler safety for AI ingestion

Stop poisoned pages before ingestion.

CrawlSign runs inside your crawler and checks every page before it enters a dataset, an index, or an agent's context. It honors opt-outs, catches tarpits, poison sources, and hidden-link traps, and returns one decision with an audit trail. Engine and data stay on your infrastructure.

  • Proprietary, by Apotropic
  • v0.1.0, pre-release for design partners
  • Runs on your infrastructure, no telemetry
stop score 80
{
  "url": "https://example.com/noai",
  "should_stop": true,
  "score": 80,
  "action": "stop",
  "reasons": [
    "ai_refusal_signal_detected"
  ],
  "threshold": 70
}
One decision per page. An offline demo with the same result contract the CLI, SDK, and API return.

Why check before ingestion

Three risks

  1. One ingestion window

    Poisoned content enters at index time. Re-indexing is the only remedy once it is in a dataset.

  2. Real legal exposure

    noai and X-Robots-Tag are cited in opt-out frameworks. Ignoring them carries risk beyond a policy question.

  3. Invisible crawl cost

    Tarpit pages inflate crawl depth and quota with no useful signal. Hidden-link clusters compound the problem quietly.

Run it where your crawler already works

Three transports

The same detector core, three transports. Start with the CLI or SDK; run the local HTTP API when crawlers in other languages or processes need the same decisions.

CLI

Scan live URLs, URL lists, or saved HTML from a terminal. Output is JSON or Markdown, stop decisions exit with code 2 for CI, and reason codes are stable across releases.

crawlsign scan https://example.com
crawlsign scan-file urls.txt --compact
crawlsign analyze-file page.html --url https://example.com/page --format markdown
crawlsign analyze-warc crawl.warc.gz --output report.jsonl

Drop into your ingestion stack

Framework adapters

Built-in loaders and drop-in guards for the stacks AI and RAG teams already use. Each one gates on the same action and never hands a withheld page to your pipeline.

LangChain

CrawlSignWebLoader fetches through one shared SafeFetcher, so robots.txt is checked first, and only safe pages become Document objects. Withheld pages are listed, without content, for the audit trail.

Integration guide →
from crawlsign.integrations.langchain import CrawlSignWebLoader

loader = CrawlSignWebLoader(urls)
docs = loader.load()     # metadata: crawlsign_action, crawlsign_reasons, ...
audit = loader.skipped   # what was withheld, and why

What gets caught

Five patterns

Five patterns the detector flags, drawn from the bundled fixture library and the mocked gzip relay in the offline demo.

Page markup (strict-enterprise config)

<div style="display:none">
  <a href="/poison/aaa...">hidden</a>
  <a href="/poison/bbb...">hidden</a>
  <!-- 8 more hidden anchors -->
</div>
quarantine score 40
  • hidden_link_detected
  • hidden_internal_link_cluster_detected

Page markup

<meta name="robots"
  content="noai,noimageai">
stop score 80
  • ai_refusal_signal_detected

Endpoint path

GET /nepenthes
User-Agent: ResearchCrawler/1.0
quarantine score 50
  • known_tarpit_reference_detected

Page markup

<h1>Hello</h1>
<p>A normal article.</p>
<a href="/about">About</a>
continue score 0
  • No signals detected

Poison Fountain-style gzip relay (live fetch)

GET /hidden/relay
Content-Encoding: gzip
85 bytes on the wire, 71 bytes decoded:
This page relays rnsaffn.com/poison2 content.
stop score 70
  • known_poison_source_detected
  • Decoded before analysis; content withheld from the caller

Designed to stop safely, not bypass controls

Decision path

  1. 1.0

    Fetch politely

    Live URL scans check robots.txt first, then fetch with size, time, and redirect limits and decode gzip before analysis.

  2. 2.0

    Inspect signals

    The detector checks refusal tags, known poison markers, tarpit endpoints, and hidden links. Signature matches take the highest score, not the sum.

  3. 3.0

    Return the action

    Every run emits the same schema: score, action, reasons, threshold, and metadata, bucketed into four default bands.

  4. 4.0

    Quarantine risk

    On the live fetch path, stopped and quarantined pages are written to a local quarantine folder with their evidence, and their content is never returned to the caller.

Default score bands Every edge is set in config. Monitor mode logs what it would withhold without blocking, strict-enterprise quarantines on a single hidden same-origin link, and repeated stops on one host escalate to a domain-wide stop.

  • continue0-29
  • log30-49
  • quarantine50-69
  • stop70+

Evidence a crawler can act on

Detection surface

Detector Signals Typical action
Robots and refusal robots.txt, X-Robots-Tag, noai, noimageai stop
Poison signatures Curated signatures for known poison sources and tarpit tools, matched by URL, endpoint path, and page marker, plus your own YAML list quarantine or stop
Hidden links Anchors hidden by inline CSS, the hidden or aria-hidden attributes, zero size, visually-hidden classes, or local <style> rules, including hiding inherited from parent elements. Same-origin targets are listed so the crawler can drop them from its frontier. log to quarantine
Safe fetching HTTP/HTTPS only; connect, read, and total deadlines plus a minimum-throughput floor against slow-drip tarpits; size caps; redirect validation and loop detection; non-HTML skipped; per-host rate limits and page, byte, and redirect budgets. The local API also rejects loopback, private, link-local, and .internal targets reject, stop, or analyze
Compressed relays Content-Encoding: gzip decoded before signature and marker analysis; compressed and decompressed sizes both capped by max_response_bytes stop, content withheld

Start in monitor mode

Design-partner pilot

See what enforcement would change, before it changes anything.

Pilots run on your infrastructure. In monitor mode CrawlSign still honors refusals, but only logs what it would withhold. The monitor report shows exactly what enforcement would change, by host and reason code, with a column for marking false positives. Enforce when the report matches what you expect.

  • Refusals such as robots.txt and noai are honored throughout
  • Risk evidence is logged, not blocked, until you switch to enforce
  • Same engine and result contract in both modes
# 1. Run your crawl, or an existing WARC, in monitor mode
crawlsign analyze-warc crawl.warc.gz \
  --config configs/audit-only.yaml --output run.jsonl

# 2. Review what enforcement would have withheld
crawlsign monitor-report run.jsonl \
  --format html --output monitor-report.html

# 3. Enforce: same engine, default config
crawlsign serve
Monitor, report, then enforce.

What you actually need to know

Common questions

How is this different from urllib.robotparser or reppy?

Standard robots parsers only read robots.txt. CrawlSign also inspects X-Robots-Tag headers, <meta name="robots"> with noai and noimageai tokens, known poison signatures, and hidden-link traps in page markup. All signals are scored and bucketed into a single action you can act on without writing decision logic yourself.

What does a Scrapy, LangChain, LlamaIndex, or Playwright integration look like?

LangChain and LlamaIndex get built-in loaders; Scrapy, Playwright, Crawl4AI, Firecrawl, httpx, and Node crawlers get drop-in guards. The integration guide states what each one checks, including where robots.txt can and cannot be enforced. The offline path (PoisonDetector.analyze) works on any bytes you already have, with no network I/O added to the critical path.

How are signatures updated? How do I avoid false positives?

Apotropic curates the signature list and ships updates with each release. Every signature is verified against the tool or source it targets and regression-tested against ordinary pages that share its words, so a page about a plant called Nepenthes is not treated as a tarpit. Every result records the signature list version it was checked against. You can add your own list in YAML.

Is there a hosted API?

No. The HTTP API runs on your infrastructure and binds to loopback (127.0.0.1:8765) by default, so pages and results never leave your machines. Binding to another interface requires --allow-public and a CRAWLSIGN_API_TOKEN. Every endpoint except /health then requires that bearer token.

Does CrawlSign bypass website controls?

No. The product is designed for crawler self-protection and responsible ingestion. It stops crawlers from ingesting harmful content and respects refusal signals by default.

Can it analyze private or internal URLs?

The live URL endpoint blocks localhost, private IP ranges, link-local hosts, multicast, and internal schemes by default. The offline analyzer accepts any bytes you supply.

Where are the API docs?

The HTTP API reference covers every endpoint and schema. A running server also serves its OpenAPI schema at /openapi.json and interactive docs at /docs.

How do we get CrawlSign?

CrawlSign is licensed to design partners during the pre-release. Request access and tell us what you crawl and where the pages end up; we will set up a pilot that starts in monitor mode.

Can we try it without blocking our crawl?

Yes. In monitor mode nothing is withheld on risk evidence alone: those pages are logged with what enforcement would have done. Refusals such as robots.txt and noai are still honored, because ignoring them is never part of a trial.

Crawling for a dataset, an index, or an agent?

CrawlSign is licensed to a small group of design partners. Tell us what you crawl and where the pages end up, and we will set up a pilot that starts in monitor mode.

adi@apotropic.com