TextWarden

About TextWardenBot

TextWardenBot is the crawler behind TextWarden. It reads public pages so the people who run those sites can find spelling, grammar, and style mistakes in their own copy. This page explains exactly what it does and how to send it away for good.

How to recognise it

Every request carries this User-Agent header and comes from our own servers. Its link points to this page.

User-Agent: TextWardenBot/0.1 (+https://textwarden.com/bot)

The User-Agent names textwarden.com on all our domains: it is the exact string our servers send.

What it does

  • Reads robots.txt first, then the sitemap. When there is no usable sitemap, it follows links from the homepage instead.
  • Fetches a page's HTML and extracts the text a visitor sees. It ignores images, scripts, and styles.
  • Loads a page in a headless browser only when its text appears after JavaScript runs, so a JavaScript-only site can be read too.
  • Checks that text for spelling, grammar, and style problems, and reports them to the person who asked us to watch that site.
  • Re-reads a page later to see what changed, using conditional requests so an unchanged page costs you almost nothing.

What it does not do

  • It does not submit forms, sign in, click buttons, or take any action on your site. It only reads.
  • It does not collect personal data, email addresses, or prices, and it does not resell anything it reads.
  • It does not use your content to train AI models.
  • It does not republish your pages. Extracted text is only shown back to the account watching that site.

How it behaves

  • It honours the Disallow rules in robots.txt. It reads Crawl-delay but caps it, so a scan finishes in minutes rather than hours.
  • It waits between requests to the same host and reads a limited number of pages per visit.
  • It sends If-None-Match / If-Modified-Since, so an unchanged page is answered with a cheap 304.
  • It respects noindex in a robots meta tag or an X-Robots-Tag header on the sites that ask us to.

Blocking it yourself

To handle it in your own configuration, add this to robots.txt. TextWardenBot then stays away from the paths you name.

User-agent: TextWardenBot
Disallow: /

Or opt out here

Prefer a switch you don't have to maintain? Tell us the domain and TextWarden stops fetching it everywhere: scheduled scans, one-off scans, and the public single-page checker. The opt-out covers every subdomain too.

We email a confirmation link to check that the request comes from the domain itself, so use an address at the domain you are opting out.

Protected by Cloudflare Turnstile. Privacy policy.
← Back to home