AgentReady scoring methodology v0.3
AgentReady measures how usable a website is for AI agents — programs that browse sites on a person's behalf: they look for information, choose products and fill in forms. The result is a number from 0 to 100. This page explains what it is made of.
The main principle: only properties of the site itself count towards the score. Facts about the site and obstacles on the way are checked by rules, without AI, so the result does not depend on which model walks the site. An AI agent's attempt is shown in the report separately and earns no points: its outcome depends on the model, not only on the site.
Every report states its methodology version. When the rules change, the version changes, and old reports stay calculated by their own version.
Four blocks make up the score
| Block | Weight | What we check |
|---|---|---|
| Access | 20 | The site responds, does not block AI agents, and robots.txt does not shut them out |
| Readability | 25 | schema.org markup, element labels, content without JavaScript |
| Scenario passability | 45 | Four typical scenarios: what the site offers for them and what gets in the way |
| Agent interfaces | 10 | llms.txt, sitemap.xml, WebMCP or an API description |
Inside a block the score is the average of its checks: passed = 1, partly = 0.5, failed = 0. The “Semantics” check and the scenarios give a fractional score from 0 to 1.
Scoring rules:
- What does not apply is not counted. If the site has no cart, the “cart” scenario is left out. The same goes for checks and scenarios that were not run (for example, time ran out) or failed through our fault.
- One scenario weighs at most 15 points. If fewer than three scenarios apply, the “Scenario passability” block weighs 15 points per scenario, and the rest of its weight is shared among the other blocks. Otherwise a single scenario would decide almost half of the score.
- An empty block gives its weight to the others. If there is nothing to score in a block, its weight is shared proportionally among the other blocks.
- WebMCP is a bonus. If the site declares tools for agents or publishes an API description, the “Agent interfaces” block goes up; if it does not, the score is not lowered.
- A blocked bot caps the score at 30. If the site does not let our bot in (a 401/403/429/503 response or an “are you a robot?” page), the remaining checks are not run and the score cannot exceed 30 points.
- An unavailable site gets no score. If the server answers everyone with an error — plain requests and the browser alike — no score is given.
Access (20)
- The site answers a plain request — the page answers with a 2xx status to an HTTP request with our User-Agent.
- Lets AI agents in — a browser with the honest User-Agent
AgentReadyBotgets the page, not an error or an “are you a robot?” check. Partly — there is a captcha on the page, or plain requests without a browser are blocked. - robots.txt does not shut out AI agents — we look at the access of 14 known agents and crawlers (OpenAI, Anthropic, Perplexity, Google, Apple, Meta, Amazon, Common Crawl, ByteDance). Passed — none is blocked, partly — some are blocked, failed — all are blocked. A robots.txt that answers with a server error is partly (by RFC 9309, crawlers then treat the site as closed).
Readability (25)
- schema.org markup (JSON-LD) — there is at least one valid JSON-LD block with a schema.org
@contextand a@type. Partly — some blocks have errors, or the markup appears only after JavaScript runs. - Semantics and element labels — a score from 0 to 1: 70% is the share of buttons, links and fields with a clear name in the accessibility tree, 30% is the page skeleton (a
main, anav, h1–h3 headings). From 0.8 — passed, from 0.5 — partly. - Content is visible without JavaScript — the share of words of the visible text that are already in the HTML from the server. From 80% — passed, from 40% — partly.
- Contacts available without AI — the home page has
mailto:ortel:links and contacts (phone, e-mail, address) in JSON-LD. Only one of the two — partly. Any program can take such contacts without parsing the layout. This check covers the home page; in the “Find contacts” scenario the same signs are also checked on the contacts page.
Scenario passability (45)
Our robot walks the same four scenarios as the AI agent, but by rules, without a language model. It clicks links and buttons the way an agent would, and records what the site offers and what got in the way.
Scenario score = share of signals present × (1 − obstacle penalty).
Signals
| Scenario | When it applies | Signals (each is an equal share of the score) |
|---|---|---|
| Find contacts | always | the phone or e-mail is a tel: / mailto: link; contacts are in JSON-LD; the home page links to a contacts page (or shows contacts itself); the phone or e-mail is in the HTML without running JavaScript |
| Find a product and its price | the site has prices, product markup, a cart or a pricing link | the home page links to a catalog, a product or pricing; the price is in the HTML without JavaScript; a product or service with a price is in JSON-LD; for stores — a labelled search field |
| Cart | there is a cart link or an “add to cart” button | the “add to cart” button is labelled and can be clicked; there is a cart link; the cart leads on to checkout. No order is placed |
| Request form | a contact form is found on the home page, on the contacts page or one link away from them | fields have labels (label, aria-label, placeholder, autocomplete) — at least 80% of the fields; the submit button has a label; the form is not behind a captcha; the form is in the HTML without JavaScript. The form is neither filled in nor submitted |
A scenario whose subject is absent from the site (no cart, no form) is “not applicable”, not zero. Search, login, sign-up and one-field subscription forms are not counted as contact forms.
Obstacles and penalties
An obstacle is whatever made the robot's click fail or cut the path short.
| Obstacle | Penalty |
|---|---|
| Bot-protection page | 100% — the scenario cannot be walked |
| A consent banner that cannot be closed without consenting to tracking | 50% |
| The element is covered by another one (pop-up, sticky header) | 30% |
| The element is outside the visible area (carousel) | 30% |
| The element is disabled | 30% |
| The element does not respond to a click | 30% |
| The element is hidden (closed menu) | 20% |
| The path leads to another website | 20% |
Several obstacles in one scenario multiply (30% and 30% leave 49%); an obstacle of one kind counts once. A consent banner that closes with a neutral button (“reject”, “only necessary”, “close”) gives no penalty: the robot closes it and goes on. The robot never consents to tracking.
The penalties are expert estimates for now. They were set from the first observations: for example, scenarios where an agent met a covered element succeeded in 4 cases out of 23, against 121 out of 182 without obstacles. There are few observations, almost all from one model. After a calibration run on a large set of sites and several models, the penalties will be replaced with measured ones (“gets in agents' way in N% of cases”), and the methodology version will change.
How the AI agent did (outside the score)
Separately from the score, an AI agent gets the same tasks in plain language and sees the page as a text snapshot of the accessibility tree — without pre-written selectors. For every scenario the report shows the outcome, the model and the steps with screenshots. It is one agent's attempt on the model shown: another agent may walk the site differently, which is why it earns no points. It is there to show where an agent gets stuck in practice and to collect data for calibrating the penalties.
The agent's answer is checked, not taken on trust: the contacts and price it reports must appear on the pages it visited, and the checkout screen and the filled-in form are detected from the page itself. An unconfirmed answer is partly. The agent closes consent banners only with a neutral option; clicking “accept” is forbidden in code.
Limits for an agent scenario: 20 steps, 90 seconds of work, 40,000 input tokens (cached tokens of providers that give a discount count as half). Scenarios not checked through our fault (a limit or a failure of the AI model, not enough time) are retried; if the retry fails too, the result is “not checked”.
Agent interfaces (10)
- llms.txt file —
/llms.txtis served, it is markdown with a# …heading, it contains links, and the checked links (the first three) open. - sitemap.xml — the sitemap from robots.txt or at
/sitemap.xmlcontains<urlset>or<sitemapindex>. - WebMCP or an API description (bonus) — the page registers tools through
navigator.modelContext, has forms with atoolnameattribute, or an MCP or OpenAPI description is served at/.well-known/mcp.json,/.well-known/mcpor arel="service-desc"link.
How the bot behaves
- It introduces itself honestly: User-Agent
AgentReadyBot/0.1. - It obeys
User-agent: AgentReadyBot/Disallow: /in robots.txt — such a site is not checked. - At most 60 requests per check and no more than one per second; a whole check takes no longer than 5 minutes.
- It does not submit forms, place orders or sign up, types nothing but test data (
Test Agent,agent-test@example.invalid,+10000000000), and does not bypass bot protection. This applies to both the robot and the AI agent: they share the same click and request rules, enforced in code. The robot types nothing at all.
Limitations
The score reflects the state of the site on a particular date. The robot follows a typical path: if contacts or the catalog hide behind unusual names, it may not find them — the signal then counts as missing even though a person would have found it. Sites change and serve slightly different content from request to request, so a repeat check may give a different result: in our measurement two checks of the same site in a row matched exactly in 85% of cases, and the difference was usually within 5 points. The score does not depend on any AI model.
Version history
- 0.3 (2026-10-03): only properties of the site count towards the score. The “Task completion” block was replaced by “Scenario passability”: a robot walks the scenarios by rules, and the score is signals × obstacle penalties. The AI agent's outcome is shown separately and earns no points. One scenario weighs at most 15 points. The reason: in the repeatability benchmark the score of the same site differed by more than 15 points between runs for 5 sites out of 20, because the AI model sometimes found what it needed and sometimes did not.
- 0.2 (2026-10-02): a scenario passed only by opening a link by its direct address counts as partly; the “Contacts available without AI” check was added to the “Readability” block; cached tokens of providers that give a discount count as half towards the scenario limit; consent banners: the agent closes them with a neutral option, and a banner that cannot be closed is partly.
- 0.1 (2026-10-01): the first version. Reports calculated by it are marked v0.1 and are not recalculated.