Home » SEO » Crawl Dependency Map: Can Search Engines Reach Your Pages Through Real Links?
Crawl Dependency Map: Can Search Engines Reach Your Pages Through Real Links?
A crawler is a bit like a visitor who can only walk through open doors. It starts from a known page, reads the links it finds, and follows them one by one. If the only route to a page runs through something a crawler can't operate, that page may never be found, however good it is.
A crawl dependency map answers a simple question: for each important page, what path does a crawler have to take to reach it, and does every step along that path actually work?
What breaks a path
- Menus built with click handlers. A dropdown that opens on click but contains no real links in the HTML.
- Links that live in JavaScript. The URLs exist only as data inside a script.
- Search boxes and filters as the only route. Crawlers don't type into forms.
- Infinite scroll or "load more" buttons. Items beyond the first batch are never exposed as links.
- Login walls. Content behind a login can't be crawled at all.
- Fragment-only URLs. Addresses that depend on a # fragment aren't treated as separate pages.
Step 1: crawl the site without running scripts
Use a crawler such as Screaming Frog or Sitebulb in plain HTML mode. Start at the homepage and let it follow only the links present in the source. This is roughly what a basic crawler sees.
Step 2: record the click depth
For every URL, note the shortest number of clicks from the homepage. A tidy structure might look like this:
| Depth | Typical pages |
|---|---|
| 0 | Homepage |
| 1 | Main category and service pages |
| 2 | Individual service, product or article pages |
| 3 | Supporting detail pages |
Anything important sitting at depth five or six is a candidate for better linking.
Step 3: compare against your list of key pages
Take the list of pages that matter to the business and mark which ones the plain crawl didn't find. These are your dependency failures. For each one, work out which step in the path is broken.
Step 4: repair the paths
- Replace click-handler menus with standard anchor tags. A script can still handle the animation.
- Give paginated lists real "next page" links, such as
/blog/page/2/. - Add the missing pages to a visible hub page or footer section.
- Add contextual links from related articles. This is also where orphan pages get rescued.
- Keep the XML sitemap up to date as a backup route.
A small example
A furniture store loaded its category pages through a mega menu that was built by JavaScript. The raw HTML had no category links at all. A plain crawl found only the homepage and five pages. After switching the menu to ordinary links, the crawl found 340 URLs, and previously unindexed categories began appearing in search.
Common mistakes
- Testing the crawl with JavaScript enabled and concluding everything is fine.
- Relying on the sitemap alone to prove a page is reachable.
- Hiding links in the footer behind accordions that don't exist in the HTML.
- Letting faceted filters generate endless link combinations, which burns crawl time on pages nobody needs.
Good crawl paths are boring: plain links, sensible hierarchy, nothing clever. That is exactly what you want.
Frequently Asked Questions
What counts as a crawlable link?
An anchor tag with an href attribute that contains a real URL, such as <a href="/services/">. Links that rely only on JavaScript click handlers, or have no href, usually can't be followed.
How many clicks from the homepage is too many?
There's no fixed number, but important pages should be reachable within about three clicks. Pages buried deeper tend to be crawled less often and get less internal authority.
Is a sitemap enough to get pages discovered?
A sitemap helps discovery, but it doesn't replace internal links. Pages with no links pointing to them send a weak signal about their importance.