> ## Documentation Index
> Fetch the complete documentation index at: https://docs.heralded.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Can AI read your site?

> What stops AI engines reading the pages that matter, and how Heralded checks it.

An engine can only cite a page it can read. **Website** in the Command Center, and **Site readability** in every [Snapshot](/concepts/snapshot) report and standalone Audit, list what stops AI engines reading your site. It sits apart from [the Heralded Score](/methodology/heralded-score) and has no score of its own.

## Which pages it checks

The checklist covers every page of yours the crawl read, your homepage always among them. Legal pages and files count toward the page checks, though they get no editorial checks. The readiness checks in Website, the standalone Audit, the Snapshot report and Home count the same pages from the same crawl. Website's **Pages** heading and pager count the inventory the table lists, including URLs the crawl found and did not read. Filtering the table changes both counts.

The site inventory reads your brand's main language and country, including pages whose URL names no locale. When the crawl leaves mapped URLs out, Website explains how many pages it read out of all the pages listed, including pages that could not be read, and how many mapped URLs it excludes for other languages or countries, non-content URLs, other domains or the page limit. The inventory count comes from the same pages as the table before filters; exclusion counts are recorded by that crawl, so an older crawl with no recorded exclusions shows no explanation.

A page engines cited in this run's answers passes, since the citation shows they read it. A citation from an earlier run doesn't count: a page cited before that is blocked now fails.

## What counts as a blocker

| Blocker | What Heralded saw |
| - | - |
| **Blocked in robots.txt** | robots.txt refuses an AI citation crawler. This shows as one row for the whole site, naming the crawlers. |
| **Tells AI engines not to use it** | the page says noindex or nosnippet |
| **Empty without JavaScript** | the HTML the server sends has almost no text |
| **Error or bot challenge** | the page answered with an error status or a bot challenge |

Each blocker comes with one line on how to fix it. A page Heralded couldn't fetch is counted as not read, never as a pass. The list gives counts, such as "2 of 150 pages blocked", not points.

An HTTP 429 response is a rate limit and stays not measured after bounded retries. A bot challenge is detected before its robots directives, so a challenge page's noindex is not reported as your page's own directive.

## The checklist in the Snapshot report

The report's verdict counts the checks below, such as "5 of 9 checks not met", and reads "Nothing to fix" only when every check is met. The line under it says whether any page is blocked from AI engines and how many pages were checked.

Under the verdict, the report groups the checks into page access, homepage content and structure, and speed, and marks each one met, not met or not measured, with what Heralded measured and the target. A blocked page and its fix sit under the page check it fails. The checks are:

* the four page checks above, each over every page it read, such as "146/150 pages pass; 4 could not be read". A page check with no failing page is met, unless most pages could not be read. The page checks read **AI crawlers allowed**, **Page not hidden from AI engines**, **Page loads for crawlers** and **Text without JavaScript**;
* the five homepage checks behind the report's technical moves: **Citation crawlers allowed**, **Main text visible without JavaScript**, and three about your homepage's structured data, **Company details for AI**, **Company and website described for AI** and **Company name, logo and profiles**.

Structured data, also called JSON-LD, is a block of code in your page that tells AI engines who you are and what you offer. Tooltips on the technical checks explain the terms.

Heralded also checks the homepage for page speed (**Main content appears quickly**, **Page responds quickly** and **Layout stays steady while loading**, known as largest paint, total blocking time and layout shift), HTTPS and HSTS, the sitemap, the canonical URL, Open Graph tags, title and meta description, and domain authority. Any of these that is not met shows with the checks above; the rest sit under **Also checked**, collapsed on the page and open in the PDF. The verdict counts every check, collapsed or not.

The sitemap check reads `/sitemap.xml` first. If nothing there reads as a sitemap, it fetches the `Sitemap:` lines in your robots.txt, up to ten, and reports a missing sitemap only when it fetched every one and none reads as a sitemap. If robots.txt names more than ten and none of the first ten reads as one, the check shows as not measured. The **Company and website described for AI** check is met once the homepage's structured data carries Organization and WebSite. A company description that lacks a name, URL, logo or two social profiles (`sameAs` in the code) shows what it lacks and keeps the points for the parts it has. Only a homepage with no company description reads as having none.

The standalone Audit shows the same checklist under **Site readability**, after its page checks. Its page table shows each crawled page's access gates and status. The report links to the Audit under the checklist. Each technical move in the report's **Priority actions** links to the check it would fix.

With the [WordPress Connector](/concepts/connections#wordpress), Heralded can fix robots.txt rules for AI crawlers itself; see [Actions](/concepts/actions).

llms.txt isn't a check here. Heralded offers it as a kit you can take at any time.

## Website in the Command Center

The **Website** page shows your latest crawl. Its sentence counts the checks that need a fix as issues, and how many of those Heralded can draft a fix for, such as "5 issues · 2 Heralded can draft". It also counts checks Heralded could not measure. The clock beside the title shows the crawl date; its tooltip also gives the next crawl date.

**Site checks** has four groups: **Page access**, **Homepage content and structure**, **Speed** and **Other checks**. Failing and unmeasured checks show the target and measured value. Passing checks fold under **Also checked** in each group. A page count opens the failing pages, and **Fix** opens the action when one is available.

**Pages** lists every page in the crawl. Beside **Access** and **Status**, it shows **Search clicks**, **Crawled by AI**, **Fetched live**, **Cited** and **Visits from AI**. Each has an info button explaining the figure. These five cover one 28-day window ending two days before the latest pinned weekly report started. The line above the table names its dates. An unmeasured figure or a URL without a page record shows –. A measured zero shows 0. When a column is – on every row shown, the line under the dates names the connections that fill it: Search Console for **Search clicks**, Cloudflare or the WordPress Connector for **Crawled by AI** and **Fetched live**, and Google Analytics for **Visits from AI**.

**Crawled by AI** counts served requests from AI training and AI search bots. **Fetched live** counts served requests made by AI assistants fetching the page for a user. Both exclude requests known to be unverified. WordPress requests with unknown verification count. **Cited** counts citations in weekly reports completed inside that window. **Visits from AI** is a floor because visits without a referrer count as direct. WordPress crawl counts are a floor because caches can answer before WordPress sees a request.

The five access marks are AI crawlers allowed, Not hidden from AI, Loads for crawlers, Text without JavaScript and Served to AI crawlers. The key under the table gives each one an info button. A mark can pass, fail or be unmeasured. The status is **Readable**, **Blocked by robots.txt**, **noindex**, **Doesn't load**, **Needs JavaScript**, **AI crawlers refused** or **Not read**. **Last checked** gives the date the crawl last read the page. On a narrow screen, scroll the table sideways to reach every column.

**Served to AI crawlers** reads the same recorded counts. It fails as **AI crawlers refused** when [Cloudflare](/concepts/connections#cloudflare) verified that an AI crawler was refused on the page, or got not found from a page that is live. It passes when Cloudflare's counts cover the page and show neither. With only the WordPress Connector it stays unmeasured, because the plugin never sees a request your CDN or firewall answered. Search engine bots and unverified requests never fail it. A refused page also raises a site fix naming the crawlers and their counts, with where to find the rule in Cloudflare's Security Events and AI Crawl Control. You change that rule yourself; no connection applies it. It counts toward the limit on site fix suggestions in [Actions](/concepts/actions#kinds-of-work).

Filter the table by **Status**, **Cited** or **Type**, or use **Search pages**. The Cited filter and default sort use citations in the latest weekly report. **Sort** keeps those citations, Page, Status, all served bot requests and Last checked as its choices; **Order** reverses the direction. These sorts use their existing report figures, so they can differ from the funnel figures. The five funnel columns have no sort control. The table has ten rows per page. These controls change the table alone; the crawl sentence and site checks stay the same.

Open a row for the shared page pane. **Page funnel** shows the same five figures and dates. Under it, a **Trend** chart per figure shows the same five figures for each weekly report in your plan's [history](/concepts/usage#history), each point a rolling 28 days. The bot breakdown groups training, AI search, live fetch and search-engine bots. These counts include served requests with verified or unknown verification and exclude known-unverified requests. **Search-engine crawls** and **Unverified AI requests** sit in separate rows below them. The pane also shows **Answers its topic** and **Substance**, which belong to [Content](/concepts/content).


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.