Crawled Isn’t Indexed: A Large-Ecommerce Indexation Workflow

Define the indexable product set, segment outcomes by template and reason, close faceted crawl traps, and improve selection signals before requesting more crawling.

Sonar the Answer Whale guides product pages through discovery, crawl, canonical, and index selection checkpoints.

Direct answer: if Google crawls many ecommerce URLs but indexes only a small subset, asking for faster crawling can amplify the wrong inventory. Define the product pages that deserve discovery, segment exclusion reasons by template, control faceted URL spaces, strengthen internal selection signals, and test representative pages before expanding the indexable set.

A current r/TechSEO thread reported millions of products, a large crawled set and a much smaller indexed set. Those counts describe one anonymous site and do not establish Google’s reason. They do expose a common category error: crawled is an observation, not an indexing approval.

Define the indexable product universe

Start with a business inventory rather than every generated URL. A candidate page should represent a real, available or intentionally retained product; satisfy a distinct search or customer job; return a stable indexable response; use a defensible canonical; appear in an accurate sitemap; and receive crawlable internal links.

Separate current products, temporarily unavailable products, discontinued products with replacements, historical products with useful demand, variants, internal search results, filter combinations and empty states. Each class needs an explicit URL policy. Five million database rows do not automatically create five million useful landing pages.

Choose the page policy before asking Google to crawl more
URL classDefault reviewEvidence to preserve
Distinct available productIndex candidateDemand, availability, content, canonical, links
Color or size variantConsolidate unless the page job differsSearch demand and product distinction
Filter combinationBlock or constrain by defaultUnique demand, inventory and stable URL rule
No-result filterReturn an honest empty-state statusRequested filter and result count
Discontinued productKeep, replace, redirect or remove by reader jobReplacement equivalence, links, demand, legal need

Segment indexing outcomes by template and reason

Do not work from a site-wide indexed percentage. Export the largest Page Indexing buckets and split them by product family, template version, creation month, stock state, canonical target, sitemap set, depth and internal-link source. Sample URLs from both indexed and excluded groups under the same template.

For each sample, record final status, robots controls, canonical declared and selected, rendered main content, product availability, price, structured data validity, internal links, sitemap membership, last modification accuracy and near-duplicate cluster. Include failures and URLs that change state during the review.

The existing crawled-currently-not-indexed matrix provides a smaller-site diagnostic. At catalogue scale, the additional requirement is segmentation: a technically valid product can still compete with thousands of near-equivalent URLs.

Close infinite or low-value crawl spaces

Google’s faceted-navigation guidance warns that filter parameters can create a practically infinite URL space. Prevent unnecessary crawling at the source where possible. Use stable parameter rules, avoid crawlable links to low-value combinations, return 404 for impossible filters, and keep indexable category pages finite and internally linked.

Canonical tags and nofollow are not equivalent to preventing discovery. A crawler can still spend resources fetching variants before consolidating them. Review pagination, sort orders, tracking parameters, session identifiers, alternate currency URLs, search endpoints and product variants together because their combinations multiply.

Keep clean sitemap sets limited to the URL classes you actually want indexed. Google says accurate lastmod values can help schedule recrawling; changing every timestamp on every deployment destroys that signal.

Improve selection before volume

  1. Choose one high-value product segment and remove accidental URL variants.
  2. Link the segment from useful category, brand and editorial paths.
  3. Make product information specific enough to complete the customer decision: compatibility, dimensions, availability, delivery, returns, evidence and limitations.
  4. Publish one clean sitemap for the segment with accurate modification dates.
  5. Track crawl, canonical selection, index status and impressions for a declared window.
  6. Expand only if the segment shows better selection without creating another URL explosion.

This workflow does not promise that Google will index a target percentage. It produces a better answer to the useful question: which page classes are selected, which are excluded, and what controllable difference separates them?

Use the page-retention decision guide for expired products and the Googlebot byte audit when product content appears late in oversized HTML.

Primary documentation

Community discussion

Discuss: Crawled Isn’t Indexed: A Large-Ecommerce Indexation Workflow

Have a question, a useful example, or a different perspective? Join the discussion, share evidence, and help other readers reach a better answer.

0 replies Moderated
No replies yet.

Be the first to ask a focused question, share a practical example, or add useful evidence.

Ask a question or join the discussion

Share evidence, a useful example, or a clear question. Be specific, stay on topic, and challenge ideas without attacking people. First-time replies may be held for moderation.