How is this different from urllib.robotparser or reppy?
Standard robots parsers only read robots.txt. CrawlSign also inspects
X-Robots-Tag headers, <meta name="robots"> with
noai and noimageai tokens, known poison signatures, and
hidden-link traps in page markup. All signals are scored and bucketed into a single
action you can act on without writing decision logic yourself.
What does a Scrapy, LangChain, LlamaIndex, or Playwright integration look like?
LangChain and LlamaIndex get built-in loaders; Scrapy, Playwright, Crawl4AI,
Firecrawl, httpx, and Node crawlers get drop-in guards. The
integration guide states what each one checks,
including where robots.txt can and cannot be enforced. The offline path
(PoisonDetector.analyze) works on any bytes you already have, with no
network I/O added to the critical path.
How are signatures updated? How do I avoid false positives?
Apotropic curates the signature list and ships updates with each release. Every
signature is verified against the tool or source it targets and regression-tested
against ordinary pages that share its words, so a page about a plant called
Nepenthes is not treated as a tarpit. Every result records the signature list
version it was checked against. You can add your own list in YAML.
Is there a hosted API?
No. The HTTP API runs on your infrastructure and binds to loopback (127.0.0.1:8765) by default, so pages and results never leave your machines. Binding to another interface requires --allow-public and a CRAWLSIGN_API_TOKEN. Every endpoint except /health then requires that bearer token.
Does CrawlSign bypass website controls?
No. The product is designed for crawler self-protection and responsible ingestion. It stops crawlers from ingesting harmful content and respects refusal signals by default.
Can it analyze private or internal URLs?
The live URL endpoint blocks localhost, private IP ranges, link-local hosts, multicast, and internal schemes by default. The offline analyzer accepts any bytes you supply.
Where are the API docs?
The HTTP API reference covers every endpoint and schema. A running server also serves its OpenAPI schema at /openapi.json and interactive docs at /docs.
How do we get CrawlSign?
CrawlSign is licensed to design partners during the pre-release. Request access and tell us what you crawl and where the pages end up; we will set up a pilot that starts in monitor mode.
Can we try it without blocking our crawl?
Yes. In monitor mode nothing is withheld on risk evidence alone: those pages are logged with what enforcement would have done. Refusals such as robots.txt and noai are still honored, because ignoring them is never part of a trial.