Checkers
robots.txt Checker
What your robots.txt actually allows, group by group
robots.txt looks simple and is not. Rules are grouped by user-agent, the most specific matching group wins outright, `Allow` can override a broader `Disallow`, and a crawler that matches no group is unrestricted. Reading the file top to bottom will tell you the wrong answer.
- Free, no signupNo account, no card
- No AI in the score986 deterministic rules
- Nothing publishedYour scans stay yours
- Answers in secondsQuick scan, no browser
This parses it the way a crawler does - resolving groups, specificity and Allow overrides - and reports what each agent can actually reach. That resolution matters: a naive check once reported this product's own robots.txt as blocking everything, because `Disallow: /scan/` contains the string `Disallow: /`.
It also checks the things that quietly break crawling: a sitemap directive that 404s, a crawl-delay no major engine honours, and blanket disallows that were meant to be scoped.
Why this matters
You edited robots.txt. What checks it now?
robots.txt is the one file where a single character decides whether search engines and AI assistants may read your site, and nothing errors when it is wrong. Google retired its Search Console robots.txt Tester in December 2023; the report that replaced it shows which robots.txt files Google fetched, when, and any errors - it does not answer whether a given agent is allowed. Google's own guidance points you to URL Inspection instead, which needs a verified property and a login.
What it costs you when this is wrong
A staging block ships and nothing errors
A `Disallow: /` carried up from a staging config removes the whole site from every compliant crawler while the shop carries on taking orders as normal. No alert fires; you find out weeks later from a traffic chart, after the pages have already gone.
Your robots.txt returns 200 and no rules
A single-page app that serves index.html for any unknown path answers /robots.txt with a page of HTML and a 200, so crawlers read a file containing no valid directives and every rule you wrote is simply absent. A 5xx is worse than a 404: while robots.txt is erroring, Google stops crawling rather than assuming the site is open.
Ads disapproved, link previews blank
The same Disallow lines apply to fetchers nobody thinks about: Google's ads landing-page crawler, and the agents that build link previews for X, Facebook, LinkedIn, Slack and WhatsApp. The symptom arrives as a campaign rejected for an unreachable landing page, or a link pasted into a chat as a bare URL with no title and no image.
Why run it here
Resolved by a parser, agent by agent
Group specificity and Allow overrides are settled by a conformance parser following RFC 9309, not by reading the file top to bottom, and the same resolution runs for 81 named agent tokens in one scan. It answers allow or deny at the site root for each token; it is not a per-path tester, and does not pretend to be.
Allowed, or just not mentioned?
Each agent's result records whether a rule actually matched it, or whether it is allowed only because your file says nothing about it. Silence is not a decision - the next wildcard rule someone adds will change the answer without anyone touching that agent.
The file, the tag and the header
A crawl block and an index block are separate failures, and both are reported in the same run: robots.txt directives, the robots meta tag, and the X-Robots-Tag response header. The header is the one that catches people out, because it never appears in the page source.
This check runs on its own here, and as part of the full audit alongside the other 985 rules. Either way the findings carry their evidence and their citation, and the report is never published or kept.
How to fix robots.txt problems
robots.txt is resolved by rules most people have never read, and reading the file top to bottom gives the wrong answer surprisingly often.
Serve it at exactly https://yourdomain.com/robots.txt with a 200 status and a text/plain content type.
Understand that a crawler uses only the most specific group matching its user-agent and ignores all others, including the wildcard.
Remember that the longest matching path rule wins, and that an Allow can override a Disallow regardless of order in the file.
Never use robots.txt to keep a page out of search results — a disallowed page can still be indexed from external links. Use a noindex meta tag, which requires the page to be crawlable.
Declare your sitemap with an absolute URL, and re-check after any edit.
What happens when you run it
Enter your address
Just the domain is enough — example.com, with or without the https. No account, no card, no crawl of your whole site.
We render it like a visitor
Your page opens in a real browser and JavaScript runs to completion, so what gets audited is the page a person actually sees — not the raw HTML your server sent. A screenshot comes back with the report as evidence of exactly what we measured.
You get findings you can act on
Up to 986 checks run against that rendered page. Each finding names the rule, quotes the evidence found on your page, cites the public specification behind it, and says what to change. Nothing is a judgement call — run it twice on an unchanged page and the score is identical.
What runs this check
- RFC 9309 robots parsing
Named because provenance is the product. Every finding in your report says which of these produced it, so you can check the reasoning rather than take a score on faith.
What this checks
20 rules · 12 affect your scoreA `Disallow: /` under the wildcard user-agent removes the entire site from every compliant crawler. This is the single most destructive line that can appear in a robots.txt, and it is usually left behind from a staging deployment.
GEO_ROBOTS_NO_BLANKET_DISALLOW scoredA noindex robots meta tag removes the page from search results entirely. When unintentional this single tag outweighs every other SEO signal on the page.
SEO_ROBOTS_META_INDEXABLE scoredAn X-Robots-Tag: noindex response header removes the page from search results just as effectively as the meta tag, but is invisible in the page source.
SEO_XROBOTS_INDEXABLE scoredrobots.txt exists
MediumWithout a robots.txt every crawler applies its own defaults and you have no way to express per-bot policy - which is the only lever you have over AI crawler access.
GEO_ROBOTS_TXT_EXISTS scoredNo snippet suppression
Mediumnosnippet removes your text snippet from results and blocks AI Overviews from quoting the page.
GEO_ROBOTS_NOSNIPPET scoredA Sitemap: line in robots.txt is the canonical discovery route for crawlers that never see your HTML.
GEO_ROBOTS_SITEMAP_REF scoredLinks are followable
MediumA page-level nofollow stops link equity flowing to everything you link to, including your own pages.
GEO_ROBOTS_NOFOLLOW scoredImages remain indexable
Mediumnoimageindex removes the page's images from image search entirely.
GEO_ROBOTS_NOIMAGEINDEX scoredWithout a wildcard group, every unnamed crawler falls back to its own defaults.
GEO_ROBOTS_WILDCARD_AGENT scorednoarchive stops search engines showing a cached copy, which also removes a fallback route to your content.
GEO_ROBOTS_NOARCHIVE scoredTranslation not blocked
Mediumnotranslate prevents search engines offering a translated version, narrowing your reachable audience.
GEO_ROBOTS_NOTRANSLATE scoredNo expiry date set
Mediumunavailable_after schedules the page's removal from the index on a fixed date.
GEO_ROBOTS_UNAVAILABLE_AFTER scoredmax-image-preview:large is required for large image thumbnails in Discover and results.
SEO_ROBOTS_MAX_IMAGE detectionmax-snippet controls how much text search engines may show, which matters for paywalled and licensed content.
SEO_ROBOTS_MAX_SNIPPET detectionCrawl-delay throttles compliant crawlers. Google ignores it; Bing and Yandex honour it.
GEO_ROBOTS_CRAWL_DELAY detectionA legacy Yandex directive naming the preferred mirror.
GEO_ROBOTS_HOST detectionExplicit control over how much text may be shown in a result or AI answer.
GEO_ROBOTS_MAX_SNIPPET detectionControls how long a video preview may be in results.
GEO_ROBOTS_MAX_VIDEO_PREVIEW detectionThe noai directive asks AI systems not to train on this page. Non-standard but increasingly honoured.
GEO_ROBOTS_NOAI detectionAsks AI systems not to train on this page's images. Non-standard.
GEO_ROBOTS_NOIMAGEAI detection
Questions
- Does robots.txt keep a page out of Google?
- No - it stops the page being crawled, not indexed. A blocked URL with inbound links can still appear in results with no description. `noindex` in a meta tag or header is what removes a page, and Google has to be able to crawl the page to see it.
- Why does my Disallow not work?
- Usually group specificity. If you have a `User-agent: *` group and a `User-agent: Googlebot` group, Googlebot obeys only the second one and ignores the first entirely - including rules you assumed applied to everyone.
Also checks
This page answers these searches too — they are the same job, so they share one page rather than being split across near-identical ones:
- robots.txt validator
- robots.txt tester
- robots txt syntax checker
- crawler access checker
Other checkers
- HTML ValidatorW3C markup conformance, checked by a real parser
- CSS ValidatorStylesheets parsed and checked against the W3C property grammars
- Schema Markup ValidatorJSON-LD parsed, expanded to real schema.org IRIs, and checked
- Rich Results TestWhich search features your structured data qualifies for
- Website Scam DetectorThe trust signals people and payment processors check before believing a site
- AI ValidatorWhether AI models can find, read, quote and operate your site
- Broken Link CheckerEvery link on the page requested, with blocked and broken told apart
- SEO CheckerTitles, meta, canonicals, indexability and internal linking
- Accessibility CheckerTwo WCAG engines — one in a real browser, one on your markup
- Security Headers CheckerResponse headers, CSP quality, and the TLS certificate underneath
- Core Web Vitals CheckerMeasured for real, plus the page-weight causes behind the numbers
- AI Agent Readiness CheckerCan an agent find your buttons, read your labels and finish the task
- W3C ValidatorMarkup and stylesheet conformance in one pass
- Answer Engine CheckerWhether AI answer engines can quote you, and attribute it
- AI Crawler CheckerWhich AI crawlers your robots.txt actually lets in
- SSL Certificate CheckerA real handshake: certificate, expiry, protocol and cipher
- Meta Tag CheckerTitles, descriptions, canonicals and robots directives
- Open Graph CheckerHow your link looks when someone shares it
- Image SEO CheckerAlt text, dimensions, formats and lazy loading
- Mobile-Friendly TestViewport, tap targets and responsive behaviour
- Website CheckerEvery module, one scan, one prioritised fix list
Run the robots.txt checker now
No signup, no credit card. Quick answers in seconds; a full audit takes about a minute.
Run a free scanNo signup. No credit card. Nothing stored but the result.