Blog
Specifics, not listicles. Everything here comes out of rules we had to write, edge cases that broke them, or disagreements between tools we had to resolve.
Whether ChatGPT, Claude, Gemini and Perplexity can reach your site, quote it accurately, and act on it. It is three separate capabilities, and a site can pass one while failing the others.
ReadThey crawl for different reasons, obey different rules, and reward different things. Optimising for one does not optimise for the other, and the settings that control them are entirely separate.
ReadA proposed file that points language models at your best content in plain Markdown. It is not a standard, it is not robots.txt, and it does not control access — here is what it does and does not do.
ReadThree acronyms for overlapping but genuinely different goals — ranking in a list, being quoted in an answer, and being represented in a generated response. Here is where each one applies.
ReadMaking your content likely to be retrieved, understood and represented accurately by generative AI. It starts with permission, not with writing — and most sites fail at the first step without knowing.
ReadCitation depends on three things you control: whether the assistant may fetch you, whether your answer can be cleanly bounded, and whether there is anything to attribute it to.
ReadFour types carry most of the weight for AI and answer engines, and a long tail of speculative markup carries none. Here is what to add, in what order, and what to stop adding.
ReadA validator will hand you 200 errors. Perhaps six of them affect a real visitor. Here is how to tell which, and why the rest are still worth clearing eventually.
ReadCSS fails silently and generously. A single malformed declaration causes the browser to discard the rest of its block, which is why the symptom is usually 'this style does nothing'.
ReadMost structured data problems are not syntax errors. They are markup that parses perfectly and describes a page that does not exist.
ReadValid markup makes you eligible for a rich result. Four other things decide whether one appears, and only some of them are yours to control.
ReadVisitors decide whether to trust a site in seconds, using a handful of signals. Most are trivial to add, and their absence is what makes a real business look like a template.
ReadWhether AI assistants can reach your site, quote it accurately and operate it are three separate questions. Most sites pass one and fail the others, with no symptom either way.
ReadNot every reported failure is a broken link. Bot protection, rate limits and redirect chains all look like breakage, and fixing them wastes the time the real 404s deserve.
ReadOn-page SEO is a small number of things done consistently. Here is the short list that accounts for most of what an audit finds, in the order worth fixing it.
ReadAutomated testing finds roughly a third of WCAG failures. Knowing which third — and which two thirds still need a human — is the difference between compliance and a clean report.
ReadSix headers, each defending against a specific attack. Here is what each one prevents, in the order worth adding them, and which one to deploy carefully.
ReadA lab measurement and your real users' experience are different numbers, and only one of them affects ranking. Here is how to read both.
ReadAgents book, buy and fill in forms on people's behalf. They do not see your page — they read the accessibility tree, which is why most sites are a dead end for them.
ReadConformance matters not because a badge says so, but because browsers recover from broken markup by guessing — and they do not all guess the same way.
ReadBeing crawlable is not being quotable. An engine has to bound your answer and attribute it, and most pages make both impossible without realising.
ReadReading robots.txt top to bottom gives the wrong answer surprisingly often. Group specificity and Allow precedence decide what a crawler may fetch — not the order of the lines.
ReadA certificate can be valid, current and still broken for a share of your visitors. Chain problems and protocol configuration fail selectively, which makes them hard to notice.
ReadOne line in robots.txt can remove a site from search entirely, and the resolution rules mean reading the file rarely tells you what it does.
ReadTitles and descriptions are what a result is assembled from, and duplicates across templated pages are the most common serious on-page problem there is.
ReadA preview is generated at share time by a scraper that does not run JavaScript and gives up quickly. Almost every broken preview comes down to one of four causes.
ReadImages are usually most of a page's weight and most of its accessibility failures. Five changes fix ranking, load time and screen reader usability at once.
ReadGoogle removed its mobile-friendly test in 2023. The requirements did not change, and mobile-first indexing means the mobile version is the one being ranked.
ReadA complete audit spans indexability, accessibility, security, structured data, performance and AI visibility. Here is what each contributes, and the order worth fixing them in.
ReadAlmost always group specificity. A crawler obeys exactly one group - the most specific one that matches its name - and ignores every other rule in the file.
ReadAlmost always a relative og:image URL. Scrapers fetch your image cold from a different network with no page context, so the URL has to be absolute.
ReadRoughly a third of WCAG issues. Automated tools find machine-checkable failures reliably and are structurally unable to judge whether alt text is meaningful.
ReadIt was removed in December 2023. Mobile did not stop mattering: Google indexes the mobile rendering of your page by default, so it is the version that ranks.
ReadBecause a large share of the web answers automated clients with 403 or 429 while serving humans normally. That is bot protection, not a broken link.
ReadAEO is making content usable by systems that answer directly instead of returning links. The mechanics differ: structure and attributability matter more than keywords.
ReadBecause browsers are required to repair broken markup rather than reject it. The page renders, and the DOM you get is not the one you wrote.
ReadCSS has no error reporting. A browser drops any declaration it cannot parse and continues, so a single typo produces a layout bug with no error anywhere.
ReadValid and eligible are different tests. Each rich result type has required properties, and markup can pass every syntax check while qualifying for nothing.
ReadNo single signal proves it. Legitimate sites accumulate verifiable details a fraudulent one has no reason to produce - and the absence of several together is the real signal.
ReadAgents drive real browsers and read the accessibility tree. A site can be perfectly crawlable and completely unusable by one - and most sites have never been tested.
ReadEvery visitor gets a full-page interstitial they must click through. The site is not slow or degraded - it is unreachable to anyone who trusts their browser.
ReadGPTBot, OAI-SearchBot, ClaudeBot and Google-Extended are separate tokens with separate consequences. Blocking one does not block the others, and allowing Googlebot does not allow any of them.
ReadTwo tools measuring the same page routinely disagree by thirty points. Almost always it is throttling, cold cache, or lab versus field data — not a bug.
ReadValid JSON-LD is not the same as eligible JSON-LD. The gap is usually required properties, mismatched content, or markup describing something the page does not show.
Read986 deterministic rules across eleven modules. Every finding comes with evidence, a citation and a fix preview.
Run a free scanNo signup. No credit card. Nothing stored but the result.