October 2026 edition
State of agent readiness
We scanned the home pages of 100 well-known websites the way an AI agent reads them. 27 could not be read at all. The other 73 scored a median of 71 out of 100.
27 of 100
home pages an automated agent could not read
71/100
median score of the 73 pages that could be read
77%
of readable sites make one robots.txt decision for AI search and AI training
41%
of readable home pages carry no JSON-LD structured data
More than a quarter of the sites turned the agent away
27 of the 100 home pages returned nothing an agent could use. 18 of those answered with a bot-protection challenge: 8 from Cloudflare, 4 from DataDome, 3 from HUMAN (PerimeterX), and 3 from Akamai. The rest returned an HTTP error, closed the connection, did not answer within eight seconds, or sent a page larger than the scanner’s 5 MB limit.
News was the hardest category to read: 7 of 10 home pages were unreadable. A site that challenges every unfamiliar automated visitor is making a choice, and it may be the right one. The cost is that an agent sent by a customer to compare prices or check a policy gets a captcha page, and reports back about a competitor instead.
One caution that matters here. The scanner announces itself honestly under its own name, which no site has reason to recognise. A site that challenged it may still admit the crawlers it knows, such as GPTBot or ClaudeBot. This number measures how a site treats an unfamiliar agent, not every agent.
What the readable pages get wrong
Of the 73 home pages that could be read, 4 scored Excellent (85 or more), 36 Good, 21 Fair, and 12 Poor (under 50). The highest score was 100 and the lowest 18. These are the checks that failed most often.
Search and training crawlers decided separately
56 of 73 failed (77%)
Being cited and being trained on are one decision instead of two, so the site either gives away more than intended or blocks itself out of AI answers by accident.
JSON-LD in the server HTML
30 of 73 failed (41%)
Every fact about the page has to be inferred from prose. An agent guesses at who wrote it and when, and guesses wrong on pages that do not follow a familiar shape.
Structured data says what the page is
30 of 73 failed (41%)
The markup describes the site and the organisation but never this page, so nothing carries its headline, its dates or its author.
Open Graph tags with an image that resolves
23 of 73 failed (32%)
A link to this page becomes a bare URL wherever it is shared or previewed, including inside chat interfaces.
One H1 per page
18 of 73 failed (25%)
Once markup is stripped, nothing says what the page is called. An agent has to guess the title from the body text, and often guesses the navigation.
Sitemap lastmod dates look real
16 of 73 failed (22%)
Crawlers cannot tell what changed, so they either recrawl everything or trust none of it. Stamped-looking dates train them to ignore the field.
Organization and contact information
15 of 73 failed (21%)
Who publishes this and how to reach them is not stated, which is part of what anyone weighs before relying on a page.
Semantic landmarks mark the content
13 of 73 failed (18%)
Nothing separates the content from the navigation and footer, so an extractor keeps the menu and the boilerplate alongside the text.
By category
| Category | Readable | Not readable | Median score of readable |
|---|---|---|---|
| Ecommerce | 5 of 10 | 5 | 52 |
| Software | 9 of 10 | 1 | 84 |
| Developer docs | 10 of 10 | 0 | 62 |
| News | 3 of 10 | 7 | 78 |
| Finance | 10 of 10 | 0 | 84 |
| Travel | 5 of 10 | 5 | 30 |
| Health | 6 of 10 | 4 | 69 |
| Education | 9 of 10 | 1 | 73 |
| Government | 9 of 10 | 1 | 76 |
| AI companies | 7 of 10 | 3 | 71 |
A median over three or five sites is a description of those sites, not of an industry. Read the category rows with the “readable” column in view.
All 100 sites
Each score is for the home page as it was served to the scanner on 2026-10-11. Scores change when pages change. Anyone can run the same scan on any of these.
| Site | Category | Home page result |
|---|---|---|
| Zoom | Software | 100 Excellent |
| WHO | Health | 96 Excellent |
| Coursera | Education | 95 Excellent |
| Vercel Docs | Developer docs | 86 Excellent |
| Atlassian | Software | 84 Good |
| Bank of America | Finance | 84 Good |
| Capital One | Finance | 84 Good |
| Cloudflare Docs | Developer docs | 84 Good |
| Dropbox | Software | 84 Good |
| Fidelity | Finance | 84 Good |
| Google DeepMind | AI companies | 84 Good |
| Harvard | Education | 84 Good |
| HubSpot | Software | 84 Good |
| Kubernetes Docs | Developer docs | 84 Good |
| Mistral AI | AI companies | 84 Good |
| Nike | Ecommerce | 84 Good |
| PayPal | Finance | 84 Good |
| Salesforce | Software | 84 Good |
| Stripe | Finance | 84 Good |
| USA.gov | Government | 84 Good |
| Wells Fargo | Finance | 84 Good |
| Shopify | Software | 82 Good |
| BBC | News | 81 Good |
| edX | Education | 81 Good |
| European Union | Government | 81 Good |
| Next.js Docs | Developer docs | 78 Good |
| The Verge | News | 78 Good |
| Cursor | AI companies | 77 Good |
| Notion | Software | 77 Good |
| The White House | Government | 77 Good |
| Wikipedia | Education | 77 Good |
| GOV.UK | Government | 76 Good |
| NASA | Government | 76 Good |
| KAYAK | Travel | 75 Good |
| Stanford | Education | 73 Good |
| Vanguard | Finance | 72 Good |
| Anthropic | AI companies | 71 Good |
| Slack | Software | 71 Good |
| UC Berkeley | Education | 71 Good |
| Healthline | Health | 70 Good |
| Airbnb | Travel | 69 Fair |
| Cleveland Clinic | Health | 69 Fair |
| Cohere | AI companies | 69 Fair |
| USPS | Government | 69 Fair |
| Walgreens | Health | 69 Fair |
| IKEA | Ecommerce | 67 Fair |
| Library of Congress | Government | 66 Fair |
| MDN Web Docs | Developer docs | 66 Fair |
| IRS | Government | 65 Fair |
| Hugging Face | AI companies | 64 Fair |
| WebMD | Health | 63 Fair |
| MIT | Education | 62 Fair |
| Replicate | AI companies | 60 Fair |
| Canada.ca | Government | 59 Fair |
| GitHub Docs | Developer docs | 58 Fair |
| The Guardian | News | 58 Fair |
| Asana | Software | 57 Fair |
| CDC | Health | 53 Fair |
| American Express | Finance | 52 Fair |
| Target | Ecommerce | 52 Fair |
| Stripe Docs | Developer docs | 50 Fair |
| The Home Depot | Ecommerce | 49 Poor |
| AWS Docs | Developer docs | 48 Poor |
| React | Developer docs | 48 Poor |
| Python Docs | Developer docs | 46 Poor |
| Chase | Finance | 44 Poor |
| Amazon | Ecommerce | 33 Poor |
| Southwest | Travel | 30 Poor |
| Duolingo | Education | 27 Poor |
| Delta | Travel | 24 Poor |
| Booking.com | Travel | 21 Poor |
| Khan Academy | Education | 20 Poor |
| Charles Schwab | Finance | 18 Poor |
| AP News | News | Not readable. Challenged by Cloudflare |
| Best Buy | Ecommerce | Not readable. No response within 8 seconds |
| Bloomberg | News | Not readable. Challenged by HUMAN (PerimeterX) |
| Canva | Software | Not readable. Challenged by Cloudflare |
| CNN | News | Not readable. Page larger than 5 MB |
| CVS | Health | Not readable. Page larger than 5 MB |
| eBay | Ecommerce | Not readable. Challenged by Akamai |
| Etsy | Ecommerce | Not readable. Challenged by DataDome |
| Expedia | Travel | Not readable. HTTP 429 |
| Hilton | Travel | Not readable. HTTP 403 |
| Johns Hopkins Medicine | Health | Not readable. Challenged by Cloudflare |
| Marriott | Travel | Not readable. Challenged by Akamai |
| Mayo Clinic | Health | Not readable. Challenged by Akamai |
| NIH | Health | Not readable. Challenged by Cloudflare |
| NPR | News | Not readable. Connection closed by the server |
| OpenAI | AI companies | Not readable. Challenged by Cloudflare |
| Oxford | Education | Not readable. Challenged by Cloudflare |
| Perplexity | AI companies | Not readable. Challenged by Cloudflare |
| Reuters | News | Not readable. Challenged by DataDome |
| Social Security | Government | Not readable. HTTP 403 |
| The New York Times | News | Not readable. Challenged by DataDome |
| The Washington Post | News | Not readable. Connection closed by the server |
| Tripadvisor | Travel | Not readable. Challenged by DataDome |
| United | Travel | Not readable. Connection closed by the server |
| Walmart | Ecommerce | Not readable. Challenged by HUMAN (PerimeterX) |
| Wayfair | Ecommerce | Not readable. Challenged by HUMAN (PerimeterX) |
| xAI | AI companies | Not readable. Challenged by Cloudflare |
How this was measured, and what it cannot tell you
- The sites. We picked 100, ten in each of ten categories, because people know them. It is not a random sample and not a ranking of the web.
- One page each. Only the home page was scanned, once, on 2026-10-11, from one location. Sites that failed for a transport reason were tried a second time.
- The scanner. The Agent Analyzer, check version 2026.10.1. It fetches a page without running JavaScript and identifies itself as AgentExperiencesScanner. It obeys robots.txt.
- The score. A weighted total of automated checks on access, structure, metadata, structured data, and trust signals. It measures how easily an agent can read and attribute a page. It does not measure whether AI systems cite the site, and a high score does not predict that they will.
- Size limit. Two home pages were larger than the scanner’s 5 MB limit and count as not readable. That is our limit, though an agent with a fixed context budget faces the same problem.