Page Indexability Checker

Example: The target page is fetched once, plus one robots.txt request. Results are not saved.

Page indexability results

No check started

Enter a page URL to see whether search engines can crawl and index it.

TECHNICAL SEO

What is a page indexability check?

A page indexability check looks at a URL from a search crawler's perspective: whether it can be reached, crawled and considered for an index. HTTP responses, robots.txt, page-level robots directives and canonical signals are reviewed together.

Crawlable does not mean guaranteed indexing. Search engines also evaluate content quality, duplication, site authority and internal links. This checker focuses on technical indexing blockers.

robots.txt

Disallow limits crawling, but it is different from noindex and does not guarantee URL removal from an index.

Meta Robots

noindex can be declared in HTML Meta or the X-Robots-Tag response header, so both locations matter.

Canonical

Canonical identifies a preferred URL. Invalid, duplicate or unintended cross-page values can replace the current URL in results.

Why check page indexability?

A template release, environment mistake or SEO plugin change can put noindex into production or point canonical tags at the wrong URL. Finding these signals early prevents long-term indexing loss, canonical replacement and wasted crawl activity.

Find hard blockers

Identify 4xx/5xx responses, robots.txt blocks and noindex.

Verify the canonical page

Check that canonical is unique, valid and intentional.

Remove redirect waste

See the final URL and every redirect hop.

How does the checker work?

The server safely fetches the target page once and makes one additional request for the site's root robots.txt. Every page-level signal comes from the same HTML response.

  1. 01

    Validate a public URL

    Reject local, private and reserved targets and limit redirects, timeouts and response size.

  2. 02

    Follow redirects

    Record each 3xx hop and verify the final status and URL.

  3. 03

    Parse index directives

    Read Meta Robots, X-Robots-Tag, canonical and hreflang values.

  4. 04

    Evaluate crawl rules

    Check whether robots.txt allows the current path and return prioritized fixes.

Which signals affect indexing?

Search engines combine several signals to decide whether to crawl, index and select a URL as canonical. Interpret each result in the context of the page's purpose.

01

HTTP status and redirects

The final page should normally return a stable 200. Long chains, loops and 4xx/5xx responses impede crawling.

02

robots.txt

Disallow limits crawling, but it is different from noindex and does not guarantee URL removal from an index.

03

Meta Robots and headers

noindex can be declared in HTML Meta or the X-Robots-Tag response header, so both locations matter.

04

Canonical

Canonical identifies a preferred URL. Invalid, duplicate or unintended cross-page values can replace the current URL in results.

Indexability checker FAQs

Does a passing result guarantee indexing?+

No. It means no obvious technical blocker was found. Indexing also depends on content quality, duplication, internal links, site reputation and search engine scheduling.

Why is a page not indexed when robots.txt allows it?+

Crawl permission is only one prerequisite. Check noindex, canonical, HTTP status, content quality and whether internal links expose the page.

Will a robots.txt block remove a page from the index?+

Not necessarily. A search engine may retain the URL from external signals. For removal, allow crawling so noindex can be read, or use the search platform's removal process.

Are results saved?+

No. Results stay only in the current page state, are not written to the database and do not create a public history URL.