Crawler policy

Draft — not yet reviewed by counsel

This is a beta draft, written to be complete enough for legal review, not yet reviewed by counsel. Do not treat it as final or binding until that review is complete and this frontmatter's status changes.

Identity

suo's crawler identifies itself as SuoBot in its User-Agent header on every request.

Verified domains only

SuoBot only crawls a domain after its owner completes domain verification — a DNS TXT record or a meta tag. It will not crawl a domain on anyone's behalf without that proof of control, so it cannot be pointed at a site you do not own.

What it obeys

  • robots.txt, checked before any page on a domain is fetched, including a User-agent: SuoBot block scoped specifically to it.
  • crawl-delay, if your robots.txt specifies one for SuoBot or for *.
  • Rate limits: independent of robots.txt, SuoBot self-limits its request rate per domain to avoid meaningfully adding to your traffic.

What it never does

  • Runs JavaScript. It reads server-rendered HTML only — content that only appears after client-side rendering is invisible to it.
  • Sends credentials, cookies, or authentication headers of any kind.
  • Submits a form, clicks a button, or otherwise interacts with a page beyond following <a href> links and reading a sitemap.
  • Crawls a domain that has not been verified.

Blocking SuoBot

Add a rule to your robots.txt:

User-agent: SuoBot
Disallow: /

Or scope it to specific paths:

User-agent: SuoBot
Disallow: /internal/
Disallow: /drafts/

Changes to robots.txt are picked up before the next crawl; there is no need to contact us to have a disallow rule respected.

Rate limits

SuoBot limits itself to a conservative number of requests per second per domain, backing off further if your server responds slowly or with error statuses, so a crawl should never be mistaken for load pressure on your site.

Contact

Questions about SuoBot's behavior, a request to have it re-crawl sooner, or a security report: crawler@usesuo.com (placeholder — to be confirmed before this page leaves draft status).


Placeholders for legal review: an explicit statement of what happens to already-indexed content if verification is later revoked, and whether a domain-wide Disallow: / in robots.txt before verification is itself sufficient to block a crawl attempt, are intentionally left for confirmation with engineering and counsel together.