Agent Experiences

October 2026 edition

State of agent readiness

We scanned the home pages of 100 well-known websites the way an AI agent reads them. 27 could not be read at all. The other 73 scored a median of 71 out of 100.

27 of 100

home pages an automated agent could not read

71/100

median score of the 73 pages that could be read

77%

of readable sites make one robots.txt decision for AI search and AI training

41%

of readable home pages carry no JSON-LD structured data

More than a quarter of the sites turned the agent away

27 of the 100 home pages returned nothing an agent could use. 18 of those answered with a bot-protection challenge: 8 from Cloudflare, 4 from DataDome, 3 from HUMAN (PerimeterX), and 3 from Akamai. The rest returned an HTTP error, closed the connection, did not answer within eight seconds, or sent a page larger than the scanner’s 5 MB limit.

News was the hardest category to read: 7 of 10 home pages were unreadable. A site that challenges every unfamiliar automated visitor is making a choice, and it may be the right one. The cost is that an agent sent by a customer to compare prices or check a policy gets a captcha page, and reports back about a competitor instead.

One caution that matters here. The scanner announces itself honestly under its own name, which no site has reason to recognise. A site that challenged it may still admit the crawlers it knows, such as GPTBot or ClaudeBot. This number measures how a site treats an unfamiliar agent, not every agent.

What the readable pages get wrong

Of the 73 home pages that could be read, 4 scored Excellent (85 or more), 36 Good, 21 Fair, and 12 Poor (under 50). The highest score was 100 and the lowest 18. These are the checks that failed most often.

  1. Being cited and being trained on are one decision instead of two, so the site either gives away more than intended or blocks itself out of AI answers by accident.

  2. JSON-LD in the server HTML

    30 of 73 failed (41%)

    Every fact about the page has to be inferred from prose. An agent guesses at who wrote it and when, and guesses wrong on pages that do not follow a familiar shape.

  3. The markup describes the site and the organisation but never this page, so nothing carries its headline, its dates or its author.

  4. A link to this page becomes a bare URL wherever it is shared or previewed, including inside chat interfaces.

  5. One H1 per page

    18 of 73 failed (25%)

    Once markup is stripped, nothing says what the page is called. An agent has to guess the title from the body text, and often guesses the navigation.

  6. Sitemap lastmod dates look real

    16 of 73 failed (22%)

    Crawlers cannot tell what changed, so they either recrawl everything or trust none of it. Stamped-looking dates train them to ignore the field.

  7. Who publishes this and how to reach them is not stated, which is part of what anyone weighs before relying on a page.

  8. Nothing separates the content from the navigation and footer, so an extractor keeps the menu and the boilerplate alongside the text.

By category

CategoryReadableNot readableMedian score of readable
Ecommerce5 of 10552
Software9 of 10184
Developer docs10 of 10062
News3 of 10778
Finance10 of 10084
Travel5 of 10530
Health6 of 10469
Education9 of 10173
Government9 of 10176
AI companies7 of 10371

A median over three or five sites is a description of those sites, not of an industry. Read the category rows with the “readable” column in view.

All 100 sites

Each score is for the home page as it was served to the scanner on 2026-10-11. Scores change when pages change. Anyone can run the same scan on any of these.

SiteCategoryHome page result
ZoomSoftware100 Excellent
WHOHealth96 Excellent
CourseraEducation95 Excellent
Vercel DocsDeveloper docs86 Excellent
AtlassianSoftware84 Good
Bank of AmericaFinance84 Good
Capital OneFinance84 Good
Cloudflare DocsDeveloper docs84 Good
DropboxSoftware84 Good
FidelityFinance84 Good
Google DeepMindAI companies84 Good
HarvardEducation84 Good
HubSpotSoftware84 Good
Kubernetes DocsDeveloper docs84 Good
Mistral AIAI companies84 Good
NikeEcommerce84 Good
PayPalFinance84 Good
SalesforceSoftware84 Good
StripeFinance84 Good
USA.govGovernment84 Good
Wells FargoFinance84 Good
ShopifySoftware82 Good
BBCNews81 Good
edXEducation81 Good
European UnionGovernment81 Good
Next.js DocsDeveloper docs78 Good
The VergeNews78 Good
CursorAI companies77 Good
NotionSoftware77 Good
The White HouseGovernment77 Good
WikipediaEducation77 Good
GOV.UKGovernment76 Good
NASAGovernment76 Good
KAYAKTravel75 Good
StanfordEducation73 Good
VanguardFinance72 Good
AnthropicAI companies71 Good
SlackSoftware71 Good
UC BerkeleyEducation71 Good
HealthlineHealth70 Good
AirbnbTravel69 Fair
Cleveland ClinicHealth69 Fair
CohereAI companies69 Fair
USPSGovernment69 Fair
WalgreensHealth69 Fair
IKEAEcommerce67 Fair
Library of CongressGovernment66 Fair
MDN Web DocsDeveloper docs66 Fair
IRSGovernment65 Fair
Hugging FaceAI companies64 Fair
WebMDHealth63 Fair
MITEducation62 Fair
ReplicateAI companies60 Fair
Canada.caGovernment59 Fair
GitHub DocsDeveloper docs58 Fair
The GuardianNews58 Fair
AsanaSoftware57 Fair
CDCHealth53 Fair
American ExpressFinance52 Fair
TargetEcommerce52 Fair
Stripe DocsDeveloper docs50 Fair
The Home DepotEcommerce49 Poor
AWS DocsDeveloper docs48 Poor
ReactDeveloper docs48 Poor
Python DocsDeveloper docs46 Poor
ChaseFinance44 Poor
AmazonEcommerce33 Poor
SouthwestTravel30 Poor
DuolingoEducation27 Poor
DeltaTravel24 Poor
Booking.comTravel21 Poor
Khan AcademyEducation20 Poor
Charles SchwabFinance18 Poor
AP NewsNewsNot readable. Challenged by Cloudflare
Best BuyEcommerceNot readable. No response within 8 seconds
BloombergNewsNot readable. Challenged by HUMAN (PerimeterX)
CanvaSoftwareNot readable. Challenged by Cloudflare
CNNNewsNot readable. Page larger than 5 MB
CVSHealthNot readable. Page larger than 5 MB
eBayEcommerceNot readable. Challenged by Akamai
EtsyEcommerceNot readable. Challenged by DataDome
ExpediaTravelNot readable. HTTP 429
HiltonTravelNot readable. HTTP 403
Johns Hopkins MedicineHealthNot readable. Challenged by Cloudflare
MarriottTravelNot readable. Challenged by Akamai
Mayo ClinicHealthNot readable. Challenged by Akamai
NIHHealthNot readable. Challenged by Cloudflare
NPRNewsNot readable. Connection closed by the server
OpenAIAI companiesNot readable. Challenged by Cloudflare
OxfordEducationNot readable. Challenged by Cloudflare
PerplexityAI companiesNot readable. Challenged by Cloudflare
ReutersNewsNot readable. Challenged by DataDome
Social SecurityGovernmentNot readable. HTTP 403
The New York TimesNewsNot readable. Challenged by DataDome
The Washington PostNewsNot readable. Connection closed by the server
TripadvisorTravelNot readable. Challenged by DataDome
UnitedTravelNot readable. Connection closed by the server
WalmartEcommerceNot readable. Challenged by HUMAN (PerimeterX)
WayfairEcommerceNot readable. Challenged by HUMAN (PerimeterX)
xAIAI companiesNot readable. Challenged by Cloudflare

How this was measured, and what it cannot tell you

  • The sites. We picked 100, ten in each of ten categories, because people know them. It is not a random sample and not a ranking of the web.
  • One page each. Only the home page was scanned, once, on 2026-10-11, from one location. Sites that failed for a transport reason were tried a second time.
  • The scanner. The Agent Analyzer, check version 2026.10.1. It fetches a page without running JavaScript and identifies itself as AgentExperiencesScanner. It obeys robots.txt.
  • The score. A weighted total of automated checks on access, structure, metadata, structured data, and trust signals. It measures how easily an agent can read and attribute a page. It does not measure whether AI systems cite the site, and a high score does not predict that they will.
  • Size limit. Two home pages were larger than the scanner’s 5 MB limit and count as not readable. That is our limit, though an agent with a fixed context budget faces the same problem.