robots.txt
Disallow limits crawling, but it is different from noindex and does not guarantee URL removal from an index.
TECHNICAL SEO
A page indexability check looks at a URL from a search crawler's perspective: whether it can be reached, crawled and considered for an index. HTTP responses, robots.txt, page-level robots directives and canonical signals are reviewed together.
Crawlable does not mean guaranteed indexing. Search engines also evaluate content quality, duplication, site authority and internal links. This checker focuses on technical indexing blockers.
Disallow limits crawling, but it is different from noindex and does not guarantee URL removal from an index.
noindex can be declared in HTML Meta or the X-Robots-Tag response header, so both locations matter.
Canonical identifies a preferred URL. Invalid, duplicate or unintended cross-page values can replace the current URL in results.
A template release, environment mistake or SEO plugin change can put noindex into production or point canonical tags at the wrong URL. Finding these signals early prevents long-term indexing loss, canonical replacement and wasted crawl activity.
Identify 4xx/5xx responses, robots.txt blocks and noindex.
Check that canonical is unique, valid and intentional.
See the final URL and every redirect hop.
The server safely fetches the target page once and makes one additional request for the site's root robots.txt. Every page-level signal comes from the same HTML response.
Reject local, private and reserved targets and limit redirects, timeouts and response size.
Record each 3xx hop and verify the final status and URL.
Read Meta Robots, X-Robots-Tag, canonical and hreflang values.
Check whether robots.txt allows the current path and return prioritized fixes.
Search engines combine several signals to decide whether to crawl, index and select a URL as canonical. Interpret each result in the context of the page's purpose.
The final page should normally return a stable 200. Long chains, loops and 4xx/5xx responses impede crawling.
Disallow limits crawling, but it is different from noindex and does not guarantee URL removal from an index.
noindex can be declared in HTML Meta or the X-Robots-Tag response header, so both locations matter.
Canonical identifies a preferred URL. Invalid, duplicate or unintended cross-page values can replace the current URL in results.
No. It means no obvious technical blocker was found. Indexing also depends on content quality, duplication, internal links, site reputation and search engine scheduling.
Crawl permission is only one prerequisite. Check noindex, canonical, HTTP status, content quality and whether internal links expose the page.
Not necessarily. A search engine may retain the URL from external signals. For removal, allow crawling so noindex can be read, or use the search platform's removal process.
No. Results stay only in the current page state, are not written to the database and do not create a public history URL.
Continue with focused checks for crawlability, canonicalization and search presentation.