Your blog is training data: lessons from the USA Today suit
On Thursday 8 October 2026, USA Today Co., Inc. and 13 affiliated entities sued OpenAI in the U.S. District Court for the Southern District of New York, case 1:26-cv-08892. They allege OpenAI copied hundreds of thousands of articles from their 19 publications, including USA TODAY, the Detroit Free Press and The Arizona Republic, to train and operate its GPT models. They are seeking damages "in excess of $250 million". OpenAI spokespeople did not immediately respond to a request for comment, and other AI companies have argued in court that training on copyrighted material is fair use.
I run a blog, a wiki and a few local models on my own hardware, so the courtroom is not the part that interests me. The mechanism is. The complaint points at datasets built from ordinary public web crawls, and those crawls visit your blog too.
What it really takes to keep your content out of AI training
The plaintiffs allege their publications make up more than 160,000 entries in WebText, the corpus OpenAI built to train GPT-2, including 83,266 entries from usatoday.com and 12,994 from freep.com. They also allege their content accounts for more than 122 million tokens in C4, a filtered English-language subset of a 2019 snapshot of the Common Crawl web archive, with roughly 23 million of those tokens from usatoday.com. They say these figures come from OpenAI's own published descriptions of its training data.
Common Crawl is a public archive of the web that anyone can download. Nobody had to target a newspaper for its pages to land in a slice of it. If your blog was public and reachable when a crawl ran, it had the same exposure.
You can check your own domain. Common Crawl runs a public index at index.commoncrawl.org, with one index per crawl. This lists every page it captured from your site in one crawl:
curl "https://index.commoncrawl.org/CC-MAIN-2019-18-index?url=yourblog.com/*&output=json"Swap in other crawl IDs from the list on that page to see other years.
So here is the honest picture. Whatever was public in past crawls already sits in datasets that have been copied and trained on, and no setting you change today reaches back into them. Opt-outs and blocklists govern the future, not the past, and even there it helps to separate what you control from what you can only request.
- You control: whether a page is public at all, whether a login sits in front of it, which machine your files live on, and which AI tools ever see your private material.
- You can only request: that a crawler skip your site, that a company not train on your pages, and that a bot identify itself honestly so you can block it.
The levers you have, and why none of them is a wall
The first lever is robots.txt. Here is a block worth adding to your own:
User-agent: GPTBot
Disallow: /
User-agent: CCBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: ClaudeBot
Disallow: /Those names cover OpenAI's crawler, Common Crawl's crawler and Anthropic's crawler. Google-Extended is a control token rather than a separate bot; it tells Google not to use your pages for its AI models. Check each company's documentation for new names now and then.
robots.txt is a note on the door. A crawler reads it and decides whether to obey. A scraper can ignore it, or arrive posing as a normal browser.
Publisher opt-out signals, such as meta tags or HTTP headers that declare "no AI training", work the same way. They state your policy; the reader still decides whether to honour it.
Bot blocking at the CDN layer is a step stronger. Cloudflare and other CDNs can refuse requests from known AI crawlers instead of asking politely, but only the ones they recognise by user agent, IP range or behaviour. A scraper that looks enough like a person gets through.
The complaint itself shows the limit. The plaintiffs allege OpenAI scraped copyrighted material regardless of paywalls or other access restrictions, and used tools to strip away the copyright management information that told readers the work was protected. Those claims have not been tested in court. But a paywall is a much bigger obstacle than a line in a text file, and the plaintiffs' claim is that it did not stop the collection.
That does not make opt-outs pointless. Here is what they buy you:
- They bind the crawlers that honour a declared policy, and the large AI companies say theirs do.
- They stop crawlers run by companies that want to play by the rules.
- They put your intent on record.
- They cost five minutes and nothing else.
What they do not do is remove a single document from a corpus that was built years ago. Treat the block above as a stated preference, not as protection.
The architectural answer: control what is exposed
The only control that does not depend on someone else's behaviour is not exposing the page. A crawler cannot copy what it cannot reach. That changes the question from "how do I block bots" to "what should be public at all".
Hosted platforms are where people slip. A draft in a hosted editor, an unlisted page, a document shared with "anyone with the link": all of it lives on someone else's server under someone else's terms, and an unlisted URL is one forwarded email away from public. I am not claiming any platform trains on drafts. You are relying on their policy, and policies change. It is the same point I made about rented AI in why a cloud AI agent is a dependency you do not control.
For anything genuinely private, this is the setup I use and recommend:
- A self-hosted wiki or notes app. BookStack, Wiki.js, Outline, Trilium, or a plain folder of Markdown files synced between your machines with Syncthing. Pick the one you will actually use.
- Auth in front of it. Best is keeping it off the public internet and reaching it over WireGuard or Tailscale. If it has to be reachable from anywhere, put it behind a reverse proxy with a login: Authelia or Authentik for single sign-on and two-factor, basic auth at minimum.
- Backups you hold. restic or borg to a disk in your house, plus one encrypted copy offsite.
- A deliberate public and private split. Decide per section, not per page. If I would not be happy to see it in a dataset, it stays private.
If you go the reverse proxy route, a minimal Caddy config looks like this:
wiki.example.com {
basic_auth {
hitesh $2a$14$REPLACE_WITH_HASH
}
reverse_proxy 127.0.0.1:3000
}Generate the hash with caddy hash-password. Port 3000 is the Wiki.js default; change it to whatever your app listens on. Then test from a machine outside your network with curl -I https://wiki.example.com. You want a 401 or no answer at all, never a 200.
Using AI on your own archive without feeding anyone else's model
Once your notes live on your own box, you may want AI to search and summarise them. Pasting them into a hosted chat sends them to someone else's server under someone else's retention policy, which undoes the work above.
You need three pieces:
- A local model runner. Ollama is the easiest start; llama.cpp gives you more control. A general model and a small embedding model are enough:
ollama pull llama3.1:8bandollama pull nomic-embed-text. - Retrieval instead of pasting. A retrieval (RAG) setup splits your documents into chunks, turns each chunk into an embedding with the local embedding model, and stores them in an index. A question pulls only the relevant chunks into the model. Open WebUI and AnythingLLM both do this with local models. If you only want search, a self-hosted index like Meilisearch may be enough.
- The index on your own disk. Know where your tool stores its vector database, back it up, and keep it out of any folder synced to a cloud drive.
The model never trains on your files; it reads the relevant chunks at question time. Delete a note, rebuild the index, and it is gone.
Check that nothing is listening publicly. On a native install, Ollama binds to 127.0.0.1:11434 by default, and ss -tlnp | grep 11434 should show that address, not 0.0.0.0. If you changed OLLAMA_HOST to reach it from another machine, put it behind the same VPN as your wiki.
Be honest about quality. An 8B model on a home GPU trails the frontier hosted models: it misses nuance and long documents strain it. For "find the note where I wrote down my borg retention policy" it works well. The trade is control, not better answers. I wrote more about that trade in running your own AI agent.
Where I land
I am not going to guess how this case ends, and the lesson does not depend on the verdict. Public means crawlable, crawlable means it can end up in a training set, and opt-outs only bind the crawlers that choose to respect them. Publish what you are comfortable seeing in a dataset, keep the rest on a machine you own, and run your own AI over it so the index never leaves your box.
If you would rather have someone set this up and harden it for you, from a private wiki behind auth to a local model with retrieval over your own files, that is the work I do. Details are on the services page. Fiverr: hiteshsaini459 · Upwork: hiteshsaini25