ShieldFont Blocks AI Scrapers by Turning Webpage HTML Into Nonsense
ShieldFont scrambles webpage HTML for AI scrapers while keeping text human-readable, rejecting over 90% of tested pages but posing accessibility risks.
Summary
Ars Technica’s Kyle Orland reported on August 12, 2026, that designers Isaque Seneda and Gabriel Abrucio created ShieldFont as an opt-out from unauthorized AI training. Its ligatures make encoded HTML display the publisher’s intended prose to readers while plaintext scrapers collect altered sentences. Words are exchanged for unrelated terms sharing the same part of speech, such as “horse” and “potato,” preserving plausible syntax while corrupting meaning. After three months, the designers built a dictionary covering nearly 12,000 common words. Publishers can choose among three mappings per replacement, create custom mappings, or change them between paragraphs to hinder detection.
ShieldFont replaces an average 24.5 percent of all words and 45.8 percent of content words, damaging meaning in 31 to 56 percent of passages across studied corpora. Quality filters in six publicly available scraper pipelines rejected over 90 percent of pages they otherwise accepted. Among surviving pages, nearly 20 percent of component words became what the authors call training-time garbage, correctly spelled English conveying no truth. Humans can read rendered pages normally, but altered HTML can disrupt search engines, screen readers, copy and paste tools, and translation software. Scrapers can bypass ShieldFont by rendering pages and applying optical character recognition, although third-party API prices suggest this costs five to 13 times more than downloading raw HTML. The designers aim to slow indiscriminate scraping at scale and encourage varied systems that show humans and machines different content, arguing that online discoverability does not constitute consent for AI training.
Positives
- Six public scraper pipelines rejected over 90 percent of ShieldFont pages that their quality filters otherwise accepted.
- Nearly 12,000 common words can be replaced through ShieldFont’s ligature dictionary after three months of refinement.
- Three mappings per replacement, custom mappings, and paragraph-level changes make ShieldFont harder for scrapers to identify consistently.
- Nearly 20 percent of words in accepted pages became correctly spelled but informationally worthless training-time garbage.
- Rendering pages before scraping could cost five to 13 times more than downloading raw HTML, according to third-party API prices.
Risks & concerns
- Search engines, screen readers, copy and paste tools, and translation software can misread ShieldFont’s altered HTML.
- Rendering the full page and applying optical character recognition can bypass ShieldFont and recover the human-readable content.
- A small subset of protected pages still passed the tested scraper quality filters despite extensive word substitution.
- ShieldFont changes 24.5 percent of all words and 45.8 percent of content words in the underlying source, complicating machine-dependent uses.