Ecommerce SEO Audit Checklist: 7 Steps for 2026
An ecommerce SEO audit checklist is a repeatable pass over whether Google can find, keep, and understand catalog URLs — products, collections, facets, canonicals, and product facts — not a blog-content review with a shopping bag in the hero image. The seven tools on this page (Botify, Lumar, Semrush, Ahrefs, Ryte, Alli AI, Search Atlas) all crawl or score some slice of that job. None of them is Shopify, none of them is Google Merchant Center, and none of them is a ticket queue. The category decision is which crawler you buy for the catalog, and which SKU or facet fail you will actually assign. US Tech Automations sits after the crawl as a proposed ticket layer, not as an eighth storefront.
TL;DR: treat collection clones and faceted URLs as first-class audit objects, fail SKUs that are out of stock with leftover indexable HTML, and only then argue about blog copy. Newly tracked pages took an average of index lag: 27.4 days according to IndexCheckr (16 million pages, updated 28 Feb 2025; 14.00% indexed within 7 days, 64.86% within 30 days, not indexed: 61.94% in the same corpus). A new colorway that ships today is not in the unpaid index this afternoon. Public crawler prices are contact vendor on this page.
Catalog pages are not blog pages
A product URL has a SKU, a price, a stock state, and a set of variants. A blog URL has a paragraph. Auditing them with the same "are there 800 words?" checklist is how a store ships 4,000 near-duplicate collection clones and then wonders why Google picked one. Ecommerce SEO is mostly information architecture plus crawl waste, with content as a supporting actor on category and guide URLs.
US e-commerce is large enough that a catalog leak is a commercial leak: e-commerce has accounted for US e-commerce share: >15% according to US Census Bureau (Quarterly E-Commerce Report series; re-check the latest release before you quote a board slide). That share is why a facet parameter that duplicates every collection is not a "nice to clean up" item. Ahrefs' 1 December 2023 Content Explorer study of about 14 billion pages found zero-traffic pages: 96.55% according to Ahrefs (1.94% received between one and ten monthly visits). Collection clones live in that mass.
BEST_OF earn rate: 15.2% according to US Tech Automations (12,514 live pages, counted 2026-08-24). seo_automation is not in the counted vertical earn-rate table; treat it as the mix-config neutral default: 10, not a vertical earn rate. 7 Best titles: 25.5% vs 14.0% according to the first-party mix-config (same 12,514-page count, 2026-08-24). This URL is a 7-step store checklist that names seven tools; the 25.5% figure is context for numbered titles in that corpus, not a promise that a badge lifts this page.
Google Search Central's large-site crawl-budget guidance still starts at crawl-budget sites: 10,000+ daily according to Google Search Central (also 1 million+ unique pages that change weekly). A store with color, size, sort, and filter parameters can manufacture that volume without meaning to.
Cost of the automation layer around a store is a different shopping trip than this crawler checklist; see ecommerce marketing automation cost. CMS-side SEO apps versus ops tools belong on SEOmatic vs AirOps for ecommerce stores. Citation in answer engines is a third job; compare that on best tools to get cited by Perplexity and come back here for the catalog audit.
Faceted navigation is the silent index leak
Faceted navigation is the filter UI that turns one collection into thousands of URLs (color, size, price, brand, in-stock). Some facet URLs deserve to rank ("red running shoes"); most deserve to be noindex, canonicalized, or parameter-blocked. The audit fail is not "we have facets." The audit fail is indexable facet URLs that duplicate a cleaner collection, eat crawl budget, and never earn a unique fact.
A practical facet rule: if the facet combination has search demand and unique products, give it a real H1, a unique intro fact, and an indexable canonical to itself. If it does not, noindex it and keep it out of the sitemap. Do not let the platform mint ?color= and ?sort= as indexable siblings of the same 48 SKUs. Log-aware crawlers (Botify, Lumar) earn their seat here because Googlebot's real path through facets will not match a desktop crawl you ran once on the homepage.
Key Takeaways
An ecommerce SEO audit checks SKU URLs, collections, facets, canonicals, stock state, and product facts — not whether the brand blog has 800 words.
IndexCheckr's 16-million-page study put average time-to-index at 27.4 days and 61.94% of pages not indexed; a new SKU is not "live in Google" at publish.
Ahrefs found 96.55% of pages in a ~14 billion page study received zero organic traffic; collection clones and leftover out-of-stock HTML join that set.
BEST_OF pages earned 15.2% on a 12,514-page first-party corpus counted 2026-08-24; seo_automation uses the mix-config default 10, not a vertical earn rate.
Fail the row when a facet is indexable without demand, or when
inventory_quantityis 0 and the URL is still indexable, and put that fail in a ticket.
Seven ecommerce checks
Indexable SKU URLs. Each buyable product has one canonical URL, a Product-shaped HTML (or structured data that matches the HTML), and a unique fact (spec, bundle, constraint) a sibling SKU cannot copy.
Collections. Category URLs have a purpose, an H1, and pagination that does not mint infinite indexable copies.
Facets. Demand-backed facets are indexable; sort and junk parameters are not.
Canonicals and parameters. Fail self-canonical mismatches, HTTP/HTTPS duplicates, and tracking parameters that inherited indexability.
Stock state. Out-of-stock SKUs get a deliberate policy (noindex, 301 to successor, or keep with availability markup) — not leftover indexable HTML.
Sitemaps and internal links. Money SKUs and collections appear in the sitemap and have HTML inlinks; orphans fail.
Crawl/log join. Confirm Googlebot is hitting the templates you want and not wasting budget on facets you noindexed in the UI but not in HTML.
A short recipe for a new collection: pick the facet combinations that deserve a URL, block the rest, ship 20 SKUs with unique facts, submit the sitemap, wait through IndexCheckr-scale lag, then inspect coverage. Do not buy a second crawler because the first 20 SKUs are still "Crawled – currently not indexed" on day three.
| Check | Weight | Hours to verify | Min score (0–5) |
|---|---|---|---|
| Indexable SKU URLs | 20% | 2 | 4 |
| Collections | 15% | 1 | 3 |
| Facets | 20% | 2 | 4 |
| Canonicals and parameters | 15% | 2 | 4 |
| Stock state | 15% | 1 | 4 |
| Sitemaps and internal links | 10% | 1 | 3 |
| Crawl / log join | 5% | 3 | 3 |
Store-crawler matrix
The matrix below is a normalized feature view plus one proprietary operating column. The USTA column is not an eighth crawler; it is this site's BEST_OF-template earn rate and corpus size.
| Capability | Botify / Lumar | Semrush / Ahrefs | Ryte / Alli AI / Search Atlas | USTA operating number |
|---|---|---|---|---|
| Class | Log-aware / cloud crawl | Suite crawl | Cloud SEO / automation | Ticket ingest |
| Facet / parameter handling | Yes (product point) | Partial | Vendor-dependent | n/a |
| Native BEST_OF earn rate | n/a | n/a | n/a | 15.2% (12,514 pages, 2026-08-24) |
| Neutral seo_automation default | n/a | n/a | n/a | 10 |
7 Best vs 5 Best title earn | n/a | n/a | n/a | 25.5% vs 14.0% |
| IndexCheckr lag context | 27.4 days | 27.4 days | 27.4 days | 27.4 days |
| Replaces Search Console | No | No | No | No |
| Weekly export | Yes | Yes | Yes | Ticket, not a crawl |
Read the USTA column as context for why this URL is a BEST_OF checklist, not as a claim that orchestration indexes SKUs. The 15.2% figure is an earn rate on this site's BEST_OF template. The 10 is the mix-config default.
Evaluation weights for a catalog crawl
| Criterion | Weight | Hours to verify | Min score (0–5) |
|---|---|---|---|
| Facet and parameter detection | 25% | 2 | 4 |
| Log-file / Googlebot path through collections | 20% | 3 | 3 |
| JS rendering of product HTML | 20% | 2 | 3 |
| Export that a ticket tool can ingest | 20% | 2 | 3 |
| Cost at SKU volume you actually crawl | 15% | 1 | 3 |
Botify and Lumar are the log-aware class for stores that can drop CDN or origin logs. Semrush and Ahrefs are suite crawlers the team may already pay for. Ryte, Alli AI, and Search Atlas are additional cloud SEO seats; confirm current catalog features on each vendor's site rather than assuming they match Botify's log join. Search Console remains free and mandatory regardless of which paid crawler you pick.
Botify and Lumar on the log path
Best fit: stores that can drop logs and need to see whether Googlebot is wasting budget on facet combinations you thought you blocked. Botify and Lumar both sit in the cloud-crawl class with log and JS-rendering workflows. They win when the question is "does Googlebot ever hit this collection template, and is it hitting 8,000 sort URLs instead?"
Limitations: these are paid platforms. List prices are contact vendor; this page does not invent a crawl-credit dollar. They do not replace Search Console coverage, Merchant Center, or the stock system of record. Confirm current connectors on Botify and Lumar the week you buy. Human review sits on any money collection Googlebot never requests, and on any facet Googlebot requests that you intended to noindex.
Who should pick which: Botify if the stack is already enterprise crawl-and-log and the team will staff it; Lumar if scheduled cloud crawls and issue grouping are the centre of the job without moving the whole SEO stack into a keyword suite. Those are qualitative fits, not a bake-off score. Primary evidence is each vendor's own site.
Suites and newcomers: Semrush, Ahrefs, Ryte, Alli AI, Search Atlas
Semrush Site Audit and Ahrefs Site Audit are suite crawlers. Best fit: teams that already pay for the suite and will actually open the audit on collection templates, not only on the blog. Limitations: neither suite is a substitute for log files at Botify depth. Site Audit is 1 cloud crawl product inside the suite rather than a log-file crawler, according to Semrush. Confirm current audit limits on that product page and on Ahrefs. Ahrefs still matters here as the publisher of the 96.55% zero-traffic study.
Ryte, Alli AI, and Search Atlas are additional cloud SEO seats a store might already own or be sold. Best fit: teams that will confirm, on the vendor's own site, that the product actually crawls faceted catalogs rather than only scoring blog copy. Limitations: contact vendor for current plans on Ryte, Alli AI, and Search Atlas. This page does not invent a feature those vendors do not document. Human review sits on any "AI optimization" pass that rewrites 4,000 SKU titles into the same pattern.
Implementation for this class: crawl collections and SKUs, export indexability / canonical / status / parameter, join to Search Console coverage, and fail the row. Do not request recrawl on a facet you still allow to index.
Inventory-aware tickets
The failure mode is not "we lack a seventh crawler." The failure mode is an out-of-stock SKU that is still indexable, or a facet that is still in the sitemap. A proposed agentic workflow on US Tech Automations could take a nightly Shopify (or other commerce) inventory export plus a weekly crawler export as the trigger, sync rows where inventory_quantity is 0 while the crawler still lists the URL as indexable, and route each fail into a queue with the SKU, the quantity, the canonical, and the export date — configurable, not a measured customer deployment. Prerequisites are Admin API or CSV access to inventory, a crawler export, and a human review point before anyone noindexes or 301s a SKU that is coming back in 48 hours. Output in the user's hands would be a ticket per failing SKU, not a new crawl credit.
Worked example, not a live case: a 6,400-SKU catalog with 820 indexable collection URLs, 140 SKUs at inventory_quantity 0 for 14 days, and 35 facet URLs still in the sitemap, could open 175 tickets (140 stock + 35 facets) and hold a 2-day merchandising SLA on the 140 instead of waiting for IndexCheckr-scale 27.4-day lag to recycle dead HTML. The 6,400 SKUs, 820 collections, 140 zeros, 14 days, 35 facets, 175 tickets, and 2-day SLA are scenario numbers meant to size the workflow; they are not a promised index rate and not a customer result. The backticked inventory_quantity field is a Shopify Admin variant field (or the equivalent stock field in another commerce platform), not a Google ranking factor.
A second proposed US Tech Automations path, still configurable, is a webhook from the same join into the team's existing tracker so the queue is not a screenshot of Site Audit. The trigger is the inventory drop or API pull; the action is a deduped ticket keyed on SKU plus export day; the output is a list a human can noindex, 301, or snooze. Idempotency matters: the same SKU must not open a second ticket if quantity is still 0 and the policy did not change. Access control matters: the commerce token stays in a secret store. Retention matters: inventory snapshots are not a substitute for Search Console coverage, so keep both.
What 12 months of catalog crawling costs
Print only figures a named source already published. Where a cell is not on that source, it reads contact vendor. Hours in the table are scenario loads for a weekly catalog-fail ritual, not a promise of what your team will spend.
| Cost line | Botify / Lumar | Semrush / Ahrefs | Ryte / Alli AI / Search Atlas | Ticket layer (proposed) |
|---|---|---|---|---|
| Licence | contact vendor | contact vendor | contact vendor | contact vendor |
| Weekly crawls | 52 | 52 | 52 | 52 |
| Analyst hours / month if joined by hand | 12 | 10 | 10 | 4 |
| SKUs in the watch list (scenario) | 6400 | 6400 | 6400 | 6400 |
| Stock-fail tickets / month (scenario) | 140 | 140 | 140 | 140 |
| Facet-fail tickets / month (scenario) | 35 | 35 | 35 | 35 |
| IndexCheckr average lag (context) | 27.4 days | 27.4 days | 27.4 days | 27.4 days |
| First-party BEST_OF earn | n/a | n/a | n/a | 15.2% (12,514 pages) |
| Neutral seo_automation default | n/a | n/a | n/a | 10 |
| Zero-traffic backdrop | 96.55% | 96.55% | 96.55% | 96.55% |
Twelve months of a cloud crawl you cannot date is still contact vendor times 12, which is not a number. Do not build a board slide from a remembered Botify SKU. Procurement closes from the vendor's current pricing URL, not from this table.
Who a store audit is for
This page is for ecommerce SEO leads, merchandisers who own URL policy, and agencies who already ship catalog templates, already can verify Search Console, and need a weekly pass so facet clones and out-of-stock HTML cannot hide behind a healthy homepage.
Red flags: skip a paid cloud crawler if the catalog is a handful of SKUs someone already checks by eye. Skip a second suite audit if Semrush or Ahrefs Site Audit is already staffed on collections. Skip a ticket layer if a spreadsheet already fails inventory_quantity 0 rows and someone actually noindexes them. Skip recrawl-as-a-strategy if the facet is still indexable in HTML.
The honest alternative to a dedicated ticket layer is not "do nothing." It is a Zapier, Make, or n8n scenario that watches an inventory CSV and a crawl export, retries on a 429, branches when inventory_quantity is 0 on an indexable URL, and writes an audit row. Those tools can support run histories, retries, error branches, and audit evidence when someone deliberately designs them. The buyer then owns observability, idempotency, escalation, access controls, retention, and maintenance for as long as Shopify or the crawler change an export shape.
A proposed US Tech Automations design does not remove those needs. It would configure the same trigger → join → ticket path with named human review points. If your n8n flow already fails out-of-stock indexable SKUs, closes duplicates, and pages a merchandiser, you do not need a new system of record. When NOT to use US Tech Automations: when Search Console coverage on one store is the whole job; when Botify already has an owner closing facet waste; or when a staffed n8n scenario already owns the inventory join, the threshold, and the ticket with retries and an audit log you trust.
FAQ
What is an ecommerce SEO audit checklist?
It is a repeatable pass over SKU URLs, collections, facets, canonicals, stock state, sitemaps, and crawl/log join so you fail catalog waste, rather than a blog-word-count review.
How long does Google take to index a new SKU?
IndexCheckr's study of 16 million pages (updated 28 Feb 2025) found an average 27.4 days to index, with 61.94% of pages not indexed in that corpus; a store can be faster or slower.
Should out-of-stock products stay indexed?
Only if you have a deliberate policy (availability markup, successor 301, or keep-with-intent); leftover indexable HTML on inventory_quantity 0 is how dead SKUs linger in the 96.55% zero-traffic mass.
Do I need Botify if I already use Semrush Site Audit?
Not if the suite audit is already staffed on collections and facets; Botify and Lumar earn their seat when Googlebot's real path through parameters is the question.
Are faceted URLs always bad for SEO?
No. Demand-backed facets with unique products can be indexable; sort parameters and junk combinations should not be.
Can n8n replace a ticket layer for stock and facet fails?
Yes, if you deliberately design retries, error branches, audit evidence, idempotency, and an owner; a ticket layer is optional packaging for teams whose inventory CSV never becomes a queue.
Close the SKU, not the blog
An ecommerce SEO audit is seven checks on catalog objects, not a longer blog style guide. Facets and stock state are the silent leaks. IndexCheckr's 27.4-day average and 61.94% not-indexed share are why a publish button is not a strategy. Ahrefs' 96.55% zero-traffic share is why collection clones are not "more content." The useful operating loop is dated crawl + inventory → fail → human policy (noindex, 301, or unique fact) → recrawl — in that order. If you want the exception path packaged as a configurable trigger, queue, and webhook with a human review point, compare that design on the agentic workflows page and on pricing against the Zapier, Make, or n8n flow you could own yourself. Read more on the resources blog if you are still mapping the rest of the store stack around coverage.
About the Author

Helping businesses leverage automation for operational efficiency.