Skip to content
SEO & Growth

7 Ecommerce SEO Crawlers 2026 [Benchmarks Inside]

Sep 7, 2026

An ecommerce SEO crawler is software that fetches product, category, and facet URLs the way a search engine would, then tells you which ones are indexable, blocked, redirected, duplicate, or too slow. It is not a merchandising engine, it is not a content writer, and it is not Google Search Console. The seven crawlers in this shortlist are Screaming Frog, Sitebulb, Botify, Lumar, Oncrawl, Ahrefs, and JetOctopus. US Tech Automations sits above them as the ticket layer that opens a row when a recrawl still shows noindex, a redirect chain, or a dropped SKU after updated_at changed.

TL;DR: Buy Screaming Frog or Sitebulb if a desktop crawl of the catalog still fits on one machine and one analyst. Buy Botify, Lumar, Oncrawl, or JetOctopus if you need log files, JS rendering at catalog scale, and a cloud schedule. Buy Ahrefs if Site Audit already lives next to Keywords Explorer and you will not stand up a second crawl platform. Skip all of them if you have 40 products and can click the sitemap by hand.

What a product-page crawler has to see

Google Search Central treats crawling and indexing as the prerequisite layer before ranking, and it publishes crawl-budget guidance for large sites, according to Google Search Central (2026). Ecommerce makes that layer mean: SKUs that appear, disappear, and change price without a human publishing a blog post. A crawler that cannot render the JS product grid, follow the canonical, and notice that /products/blue-shirt?color=blue is a duplicate of /products/blue-shirt is the wrong crawler.

Faceted navigation is the usual failure. A 4-filter category (color, size, price, brand) can explode into thousands of URLs that all share one unique fact: the same 12 products. The crawler should count those URLs, flag parameter copies, and keep crawl budget on the canonical category plus in-stock SKUs. The merchandiser still owns which facets are indexable. The crawler does not pick the merchandising strategy; it reports whether the live HTML matches the strategy you wrote down.

On our own corpus this page is a BEST_OF. BEST_OF earn rate: 15.2% according to US Tech Automations (2026), counted on 12,514 live pages on 2026-08-24. '7 Best' titles: 25.5% vs 14.0% according to US Tech Automations Phase 1 count (2026), on 247 versus 322 pages. Those are page-shape earn rates, not crawl-budget percentages.

Catalog scale elsewhere is a reminder that URL count is a first-class input. Canva feature pages: ~310 URLs according to Ahrefs (2026), with roughly 13 million estimated monthly organic visits, while NutriScan sits at about 4,000 pages and 182,000 estimated monthly visits (Ahrefs programmatic SEO guide, updated 2026-09-02). An ecommerce catalog is often larger than 4,000 SKUs; the crawler has to finish, not merely start.

Who this is for

This page is for in-house ecommerce SEO, marketplace catalog managers, and agencies who recrawl after feed updates and need a ticket when indexability flips.

Red flags: Skip Botify if you will never upload log files and the catalog fits in Screaming Frog. Skip Screaming Frog if the store is JS-rendered behind a queue you cannot finish overnight on a laptop. Skip Ahrefs Site Audit if the only question you have is "did last night's price feed change canonicals?" and you already have Lumar on a schedule.

If the real gap is a content grader, stop and read Surfer vs Clearscope for SaaS. If the gap is suite versus grader, use Semrush vs Surfer for SaaS. The ticket layer above a grader (not a crawler) is Surfer vs the platform layer.

Screaming Frog vs Botify (then the five around them)

Screaming Frog

Best fit: an analyst who can run a desktop crawl, export, and ticket. Screaming Frog is a desktop crawler (JS rendering on the paid licence, custom extraction, sitemap compare). The free edition is limited to 500 URLs. Screaming Frog free cap: 500 URLs according to Screaming Frog (2026). A mid-size catalog is already past that cap, so the paid licence is the ecommerce default.

Limitations: public paid-licence dollars are contact vendor on this page (we did not retrieve a 2026-09-07 price list). Desktop crawls do not replace log-file analysis. Implementation: crawl after the nightly feed, export indexability, join to SKU updated_at, ticket the diffs. Linked primary evidence: Screaming Frog.

Sitebulb

Best fit: the same desktop job when you want hints and a more opinionated audit UI than a raw crawl table. Sitebulb is a desktop crawler.

Limitations: contact vendor for dollars. Sitebulb is not Botify log analysis. Implementation: same recrawl-to-ticket loop as Screaming Frog. Linked primary evidence: Sitebulb.

Botify

Best fit: enterprise catalogs that need crawl + logs + JS at a scale a laptop will not finish. Botify is a cloud SEO platform built on crawl and log-file analysis.

Limitations: contact vendor. Botify is wasted if nobody will look at log lines. Implementation: schedule the crawl, join logs, watch whether Googlebot actually hits the new SKUs. Linked primary evidence: Botify.

Lumar

Best fit: the same enterprise crawl job when Lumar (formerly Deepcrawl) is already in the stack. Lumar is a cloud crawler and site intelligence platform.

Limitations: contact vendor. A cloud crawl that nobody tickets is a report. Implementation: schedule, alert on noindex deltas, name the owner. Linked primary evidence: Lumar.

Oncrawl

Best fit: crawl + log analysis when Oncrawl is the platform you already know. Oncrawl is a cloud crawler with log integration.

Limitations: contact vendor. Oncrawl is not a Shopify merchandising app. Implementation: same as Botify/Lumar — schedule, join, ticket. Linked primary evidence: Oncrawl.

Ahrefs

Best fit: teams whose weekly habit is Site Explorer and who will run Site Audit rather than stand up Botify. Ahrefs is an index plus an online crawler.

Limitations: contact vendor for suite dollars. Site Audit is not a replacement for server logs. Implementation: keep Site Audit on a schedule; do not pretend organic_traffic explains a noindex you never crawled. Linked primary evidence: Ahrefs.

JetOctopus

Best fit: cloud crawl, JS rendering, and log analysis without going to Botify. JetOctopus is a cloud crawler.

Limitations: contact vendor. JetOctopus will not write unique category copy. Implementation: recrawl after feed runs, export, ticket. Linked primary evidence: JetOctopus.

Evaluation criteria

CriterionWeightEvidence you must keep
Finishes the catalog25%1 completed crawl vs sitemap count
JS rendering when the grid needs it20%1 rendered vs unrendered diff
Log-file join (cloud tools)20%1 Googlebot hit vs new SKU
Public price you can audit15%1 dated $ or "contact vendor"
Recrawl-to-ticket loop20%1 named owner per delta

Feature matrix

VendorDesktop or cloudLog filesFree/public cap on this pagePublic $
Screaming Frogdesktopno500 URLs freecontact vendor (paid)
Sitebulbdesktopnocontact vendorcontact vendor
Botifycloudyescontact vendorcontact vendor
Lumarcloudyescontact vendorcontact vendor
Oncrawlcloudyescontact vendorcontact vendor
Ahrefscloud Site Auditnocontact vendorcontact vendor
JetOctopuscloudyescontact vendorcontact vendor

Dated TCO and counted rates

VendorEval weightPublic $ or capRetrieval date
Screaming Frog16%500 URL free cap2026-09-07
Sitebulb12%contact vendor2026-09-07
Botify16%contact vendor2026-09-07
Lumar14%contact vendor2026-09-07
Oncrawl14%contact vendor2026-09-07
Ahrefs14%contact vendor2026-09-07
JetOctopus14%contact vendor2026-09-07
MetricValueWindow
BEST_OF earn rate15.2%12,514 pages, 2026-08-24
'7 Best' earn rate25.5%247 pages, 2026-08-24
'5 Best' earn rate14.0%322 pages, 2026-08-24
Screaming Frog free cap500 URLsvendor public limit
Canva ranking URLs (Ahrefs)~310guide updated 2026-09-02
Canva estimated monthly visits~13 millionguide updated 2026-09-02
NutriScan pages~4,000guide updated 2026-09-02
NutriScan estimated monthly visits~182,000guide updated 2026-09-02

seo_automation uses the mix-config neutral default 10, not a vertical earn rate. A content grader is a different SKU: Surfer Standard: $99/mo yearly according to Surfer (2026) buys a Content Editor, not a 12,400-SKU recrawl.

Recrawl recipe after a feed update

Worked example: a Shopify catalog with 12,400 SKUs, 86 categories, a nightly feed at 02:00, and a 6-hour crawl window. The feed writes Shopify updated_at on 1,180 products. Screaming Frog (paid licence, not the 500-URL free cap) recrawls those 1,180 plus the 86 categories. 41 URLs newly return noindex, 17 are 301 chains, 9 are out-of-stock canonicals pointing at themselves. A Make scenario can start the crawl export and retry on 5xx. US Tech Automations opens 67 tickets keyed by SKU handle + updated_at and will not close the night until a human accepts "out of stock, leave noindex" on the 9. The crawler did not decide merchandising; it listed the diffs.

If you stitch this in Zapier, Make, or n8n, you can keep run histories, retries, error branches, and audit evidence when configured. You must own observability, idempotency (one ticket per handle + updated_at), escalation, access controls, retention, and maintenance. A proposed US Tech Automations design would use that same key, require a merchandiser review point on canonical changes, and still need Screaming Frog, Botify, or Lumar as the crawl prerequisite. It would not crawl Shopify for you.

Desktop versus cloud is the other fork. If 12,400 SKUs plus faceted copies will not finish on the laptop before the next feed, stop arguing about Screaming Frog hints versus Sitebulb hints and buy Botify, Lumar, Oncrawl, or JetOctopus. If logs show Googlebot wasting budget on facet copies, that is a robots/canonical job the desktop export already hinted at — the cloud join just proves Googlebot followed the waste.

Log lines versus HTML: what each crawler actually proves

A desktop crawl (Screaming Frog, Sitebulb) proves what your machine fetched: status code, canonical, noindex, rendered HTML if JS rendering is on. It does not prove Googlebot fetched the same URL. That is why Botify, Lumar, Oncrawl, and JetOctopus sell log joins. If logs show Googlebot spending budget on ?color= copies while the 1,180 updated SKUs get two hits a week, the ticket is robots/canonical, not "buy more content."

JS rendering is the other split. A Shopify product grid that hydrates client-side will look empty to a crawler with rendering off. You then ticket 12,400 "thin pages" that are not thin in Chrome. Turn rendering on for a sample of 50 SKUs, compare rendered vs unrendered word counts, and only then decide whether the whole catalog needs JS rendering. Cloud crawlers will bill you for that; a laptop will just not finish.

Sitemaps versus crawl-finished count is the weekly benchmark. If the Shopify sitemap lists 12,400 products plus 86 categories and the crawl finishes 9,100, you do not have a ranking problem yet. You have a crawl-budget or blocker problem. Join the 3,300 misses to updated_at. If last night's 1,180 updates are inside the 3,300 misses, the feed won and the crawler lost. That is a schedule problem: start the crawl after the feed, not at 01:00 while the feed still writes.

Facets: write the policy before the crawl. Example policy you can steal: collection URLs indexable; single-facet color URLs canonical to the collection; two-or-more-facet URLs noindex. The crawler then counts violations. Without the policy, the export is just a large CSV.

Common crawl mistakes on Shopify-class catalogs

Crawling once a month while the feed runs nightly.

Using the Screaming Frog free 500-URL cap on a 12,400-SKU store and calling it a full audit.

Rendering off, then wondering why the product grid is empty.

Treating Ahrefs organic_traffic as proof the new SKU is indexable.

Letting Make retry a 200 export that missed the 41 new noindex rows.

Buying Botify and never uploading logs.

Key Takeaways

  • Crawl after the feed, not after the monthly SEO meeting.

  • Screaming Frog and Sitebulb win on desktop catalogs; the free Frog cap is 500 URLs.

  • Botify, Lumar, Oncrawl, and JetOctopus win when logs and JS at catalog scale are the job.

  • Ahrefs wins when Site Audit already sits next to the keyword suite.

  • Google Search Central still puts crawl and index before rank; the crawler is that layer, not a grader.

  • Tickets key on SKU + updated_at. Retries without that key duplicate work.

When NOT to use US Tech Automations

Skip the ticket layer if Screaming Frog already dumps diffs into a sheet one analyst clears before standup. Skip it if Botify alerts already page the merchandiser and the only missing piece is a CSS change. Skip it if 40 SKUs fit in a sitemap you read by eye. Those are the cases where the simpler existing tool wins.

The homepage is the company. Next step with a price list: pricing.

FAQs

What is an ecommerce SEO crawler?

It is software that fetches product, category, and facet URLs and reports indexability, status codes, canonicals, duplicates, and speed. It is not a merchandising app and it is not Search Console.

Is Screaming Frog or Botify better for Shopify?

Screaming Frog is better if a desktop paid crawl finishes after the nightly feed and one analyst tickets the diffs. Botify is better if you need logs, JS at catalog scale, and a cloud schedule. The free Frog cap of 500 URLs is not a Shopify catalog crawl.

Do we still need Google Search Console?

Yes. Search Console is the Google-side ledger (impressions, coverage, sitemaps). The crawler is your fetch of the HTML you actually shipped. Use both. Search Central's crawling and indexing docs are the prerequisite layer, not a vendor.

Can Ahrefs Site Audit replace Lumar?

It can replace a second cloud crawler if Site Audit already runs on a schedule and your question is "what does Ahrefs fetch?" It cannot replace log-file proof of Googlebot. If logs are the question, stay on Lumar, Botify, Oncrawl, or JetOctopus.

How should we treat faceted URLs?

Count them, canonical them, and decide which ones are indexable in writing. The crawler reports explosion; merchandising owns the policy. Do not leave 4-filter copies indexable by accident.

Will Zapier recrawl the store for us?

Zapier, Make, or n8n can start exports, retry 5xx, and file rows when configured. They will not render JS or parse canonicals unless you add a crawler. Use them as the pipe from crawl export to tickets.

What is the first benchmark to put on the weekly slide?

Put sitemap count, crawl-finished count, new noindex count, and new 3xx count, joined to how many SKUs had updated_at that night. Do not lead with a domain rating.

About the Author

Garrett Mullins
Garrett Mullins
Workflow Specialist

Helping businesses leverage automation for operational efficiency.