Crawled Isn’t Indexed: A Large-Ecommerce Indexation Workflow
Define the indexable product set, segment outcomes by template and reason, close faceted crawl traps, and improve selection signals before requesting more crawling.
Direct answer: if Google crawls many ecommerce URLs but indexes only a small subset, asking for faster crawling can amplify the wrong inventory. Define the product pages that deserve discovery, segment exclusion reasons by template, control faceted URL spaces, strengthen internal selection signals, and test representative pages before expanding the indexable set.
A current r/TechSEO thread reported millions of products, a large crawled set and a much smaller indexed set. Those counts describe one anonymous site and do not establish Google’s reason. They do expose a common category error: crawled is an observation, not an indexing approval.
Define the indexable product universe
Start with a business inventory rather than every generated URL. A candidate page should represent a real, available or intentionally retained product; satisfy a distinct search or customer job; return a stable indexable response; use a defensible canonical; appear in an accurate sitemap; and receive crawlable internal links.
Separate current products, temporarily unavailable products, discontinued products with replacements, historical products with useful demand, variants, internal search results, filter combinations and empty states. Each class needs an explicit URL policy. Five million database rows do not automatically create five million useful landing pages.
| URL class | Default review | Evidence to preserve |
|---|---|---|
| Distinct available product | Index candidate | Demand, availability, content, canonical, links |
| Color or size variant | Consolidate unless the page job differs | Search demand and product distinction |
| Filter combination | Block or constrain by default | Unique demand, inventory and stable URL rule |
| No-result filter | Return an honest empty-state status | Requested filter and result count |
| Discontinued product | Keep, replace, redirect or remove by reader job | Replacement equivalence, links, demand, legal need |
Segment indexing outcomes by template and reason
Do not work from a site-wide indexed percentage. Export the largest Page Indexing buckets and split them by product family, template version, creation month, stock state, canonical target, sitemap set, depth and internal-link source. Sample URLs from both indexed and excluded groups under the same template.
For each sample, record final status, robots controls, canonical declared and selected, rendered main content, product availability, price, structured data validity, internal links, sitemap membership, last modification accuracy and near-duplicate cluster. Include failures and URLs that change state during the review.
The existing crawled-currently-not-indexed matrix provides a smaller-site diagnostic. At catalogue scale, the additional requirement is segmentation: a technically valid product can still compete with thousands of near-equivalent URLs.
Close infinite or low-value crawl spaces
Google’s faceted-navigation guidance warns that filter parameters can create a practically infinite URL space. Prevent unnecessary crawling at the source where possible. Use stable parameter rules, avoid crawlable links to low-value combinations, return 404 for impossible filters, and keep indexable category pages finite and internally linked.
Canonical tags and nofollow are not equivalent to preventing discovery. A crawler can still spend resources fetching variants before consolidating them. Review pagination, sort orders, tracking parameters, session identifiers, alternate currency URLs, search endpoints and product variants together because their combinations multiply.
Keep clean sitemap sets limited to the URL classes you actually want indexed. Google says accurate lastmod values can help schedule recrawling; changing every timestamp on every deployment destroys that signal.
Improve selection before volume
- Choose one high-value product segment and remove accidental URL variants.
- Link the segment from useful category, brand and editorial paths.
- Make product information specific enough to complete the customer decision: compatibility, dimensions, availability, delivery, returns, evidence and limitations.
- Publish one clean sitemap for the segment with accurate modification dates.
- Track crawl, canonical selection, index status and impressions for a declared window.
- Expand only if the segment shows better selection without creating another URL explosion.
This workflow does not promise that Google will index a target percentage. It produces a better answer to the useful question: which page classes are selected, which are excluded, and what controllable difference separates them?
Use the page-retention decision guide for expired products and the Googlebot byte audit when product content appears late in oversized HTML.
Ask a question or join the discussion