Skip to content
AI & Automation

7-Step Indexing SEO Audit Checklist Teams Use 2026

Sep 15, 2026

An indexing SEO audit checklist is a repeatable pass over whether Google can discover, crawl, and keep a URL — robots, canonicals, sitemaps, internal links, coverage, and recrawl — not a copy critique. The seven tools on this page (JetOctopus, OnCrawl, Botify, Screaming Frog, Lumar, Semrush, Ahrefs) all crawl or report on some slice of that job. None of them is Google Search Console, and none of them is a ticket queue. The category decision is which crawler you buy for the weekly pass, and which coverage fail you will actually assign. US Tech Automations sits after the crawl as a proposed ticket layer, not as an eighth crawler.

TL;DR: start in Search Console coverage, crawl the host with one log-aware platform or a desktop spider, fail URLs that are discovered-not-indexed, noindexed, canonicalized away, orphaned, or missing from the sitemap, then request recrawl only after the HTML is fixed. Newly tracked pages took an average of index lag: 27.4 days according to IndexCheckr (16 million pages, updated 28 Feb 2025; 14.00% indexed within 7 days, 64.86% within 30 days). In the same corpus not indexed: 61.94% according to IndexCheckr — which is why a publish button is not an indexation strategy. Public crawler prices are contact vendor on this page except where a vendor publishes a free cap we can date.

Why most indexing audits stall

Most indexing audits stall because the team crawls once, exports 40,000 rows, and never fails a row. Coverage in Search Console is the system of record for what Google kept. A crawler is the system of record for what your own HTML and internal links allow. Mixing them without a ticket means the same noindex template ships next week.

BEST_OF earn rate: 15.2% according to US Tech Automations (12,514 live pages, counted 2026-08-24; 28-day GSC window 2026-07-25..2026-08-21). seo_automation is not in the counted vertical earn-rate table; treat it as the mix-config neutral default: 10, not a vertical earn rate. 7 Best titles: 25.5% vs 14.0% according to the first-party mix-config (same 12,514-page count, 2026-08-24). This URL is a 7-step checklist that names seven tools; the 25.5% figure is context for numbered titles in that corpus, not a promise that a badge lifts this page.

Ahrefs' 1 December 2023 Content Explorer study of about 14 billion pages found zero-traffic pages: 96.55% according to Ahrefs (1.94% received between one and ten monthly visits). Unindexed URLs never even enter that study's traffic buckets. Screaming Frog's free desktop crawl still caps at Frog free cap: 500 URLs according to Screaming Frog (paid licence lifts the cap; confirm current limits on that page). Google's own recrawl guidance remains recrawl window: days to weeks according to Google Search Central (docs updated 2025-12-10): a request is not a same-hour reindex. Log-aware cloud crawls remain a 1-host job according to Botify when this site's BEST_OF templates still earned 15.2% across 12,514 live pages.

Time-to-index tactics that go deeper than this checklist live on how to reduce time to index new pages. The first-party story of pages that never entered the index lives on why 48 percent of our pages never got indexed. Orphan recovery at crawl scale lives on how we fixed 1,400 orphan pages. Use those for the adjacent jobs; stay here for the weekly checklist and the seven-tool split.

Key Takeaways

  • An indexing SEO audit checks discovery, crawl, and coverage — robots, canonicals, sitemaps, orphans, noindex, and recrawl — not whether the H1 is clever.

  • IndexCheckr's 16-million-page study put average time-to-index at 27.4 days, 14.00% indexed within 7 days, 64.86% within 30 days, and 61.94% not indexed.

  • Screaming Frog's free spider caps at 500 URLs; log-file platforms (JetOctopus, OnCrawl, Botify, Lumar) are the other class of seat, with prices contact vendor.

  • BEST_OF pages earned 15.2% on a 12,514-page first-party corpus counted 2026-08-24; seo_automation uses the mix-config default 10, not a vertical earn rate.

  • Request recrawl only after the HTML is fixed; Google's own window is days to weeks, not a same-hour guarantee.

First-party mix used on this checklist

Template or title patternEarn rateCorpusCount date
BEST_OF15.2%12,514 pages2026-08-24
Neutral seo_automation default1012,514 pages2026-08-24
7 Best titles25.5%12,514 pages2026-08-24
5 Best titles14.0%12,514 pages2026-08-24
IndexCheckr average lag27.4 days16 million pages28 Feb 2025
Not indexed in that study61.94%16 million pages28 Feb 2025

Those rows are the dated operating numbers already used in the prose. They are not a crawler KPI and not a promise that a seventh licence will index a noindex URL.

Seven checks in order

Run these seven checks in order so you do not request recrawl on a URL that is still noindexed. The tools below can each cover several checks; the order is the point of the checklist, not a claim that you need seven licences.

  1. Robots and host rules. Fetch robots.txt and the HTML robots / googlebot directives. Fail any money URL with noindex or a disallow that was not intentional.

  2. Canonicals. Fail any money URL whose canonical points at a different template, a parameter strip you did not mean, or a 404.

  3. Sitemaps. Fail any indexable money URL missing from the sitemap, and any sitemap URL that 404s or noindexes.

  4. Internal links / orphans. Fail any indexable URL with zero internal links from a crawled HTML path.

  5. Search Console coverage. Join the crawl to coverage. Fail discovered-not-indexed, crawled-not-indexed, and excluded money URLs.

  6. Log files (if you have them). Fail URLs Googlebot never hits and URLs Googlebot hits that you noindexed.

  7. Recrawl request. Only after 1–6 are clean, request recrawl and wait days to weeks, not minutes.

A short recipe for a new template: ship 20 URLs, confirm they are indexable and linked, submit the sitemap, wait through IndexCheckr-scale lag (27.4 days average in that study, not a promise for your host), then inspect coverage. Do not buy a second crawler because the first 20 URLs are still "Crawled – currently not indexed" on day three.

How to weigh crawlers for indexation

Score the crawler on the job you are hiring: find URLs your HTML would let Google keep, then join that list to what Google actually kept. Log-file platforms win when you need Googlebot's real path. Desktop spiders win when you need a cheap, local crawl of a bounded set. Suites win when the team already lives in Semrush or Ahrefs and will not open a fourth login.

CriterionWeightHours to verifyMin score (0–5)
JS / rendering of indexable HTML20%23
Log-file / Googlebot path20%33
Orphan and canonical detection20%24
Export that a ticket tool can ingest20%23
Cost at the URL volume you actually crawl20%13

No single vendor wins every row. JetOctopus, OnCrawl, Botify, and Lumar are the log-aware class. Screaming Frog is the desktop class with a 500-URL free cap. Semrush and Ahrefs are suite crawlers that already sit next to keyword and link graphs. Search Console remains free and mandatory regardless of which paid crawler you pick.

Indexation tool matrix

The matrix below is a normalized feature view plus one proprietary operating column. The USTA column is not an eighth crawler; it is this site's BEST_OF-template earn rate and corpus size, printed so the table cannot be copied onto a generic crawler roundup without the first-party numbers.

CapabilityJetOctopus / OnCrawl / BotifyScreaming FrogLumarSemrush / AhrefsUSTA operating number
ClassLog-aware cloud crawlDesktop spiderCloud crawlSuite crawlTicket ingest
Free URL capcontact vendor500contact vendorcontact vendorn/a
Log-file joinYes (product point)Limited / add-onYesLimitedn/a
Native BEST_OF earn raten/an/an/an/a15.2% (12,514 pages, 2026-08-24)
Neutral seo_automation defaultn/an/an/an/a10
7 Best vs 5 Best title earnn/an/an/an/a25.5% vs 14.0%
Weekly exportYesYes (500 free)YesYesTicket, not a crawl
Replaces Search ConsoleNoNoNoNoNo

Read the USTA column as context for why this URL is a BEST_OF checklist, not as a claim that orchestration indexes pages. The 15.2% figure is an earn rate on this site's BEST_OF template. The 10 is the mix-config default. Neither figure is a Botify KPI.

Log-file crawlers: JetOctopus, OnCrawl, Botify

Best fit: teams that can drop server logs (or a CDN log) and want to see what Googlebot actually requested, not only what a synthetic crawl found. JetOctopus, OnCrawl, and Botify all sit in this class. They win when the question is "does Googlebot ever hit this template?" and when JavaScript rendering plus log join is the weekly job.

Limitations: these are paid cloud platforms. List prices are contact vendor; this page does not invent a crawl-credit dollar. They do not replace Search Console coverage. A log join that nobody fails is still a dashboard. Implementation differs by vendor — confirm current connectors on JetOctopus, OnCrawl, and Botify the week you buy. Human review sits on any money template Googlebot never requests, and on any template Googlebot requests that you noindexed.

Who should pick which: JetOctopus if the team wants a crawl-plus-log UI without an enterprise services wrapper; OnCrawl if log analytics and segmentation are the centre of the job; Botify if the stack is already enterprise crawl-and-log and the team will staff it. Those are qualitative fits, not a bake-off score. Primary evidence is each vendor's own site. Do not treat a reseller "Botify vs OnCrawl" chart as a substitute for a dated quote.

Desktop and suite crawlers: Frog, Lumar, Semrush, Ahrefs

Screaming Frog is the desktop spider most SEO teams already own. Best fit: a bounded crawl (especially under the 500-URL free cap, or a paid licence for larger hosts) when you need indexability, canonicals, and status codes on a laptop this afternoon. Limitations: it is not a log-file platform in the same way Botify is, and a local crawl is only as current as the last time someone pressed Start. Confirm the 500-URL free cap and paid licence terms on Screaming Frog.

Lumar (formerly DeepCrawl) is a cloud crawler with site-monitoring workflows. Best fit: teams that want scheduled cloud crawls and issue grouping without moving the whole SEO stack into Semrush. Limitations: contact vendor for current plans; it is not Search Console. Confirm connectors on Lumar.

Semrush Site Audit and Ahrefs Site Audit are suite crawlers. Best fit: teams that already pay for the suite and will actually open the audit, not teams buying Semrush or Ahrefs only to crawl. Limitations: neither suite is a substitute for log files at Botify depth, and neither replaces coverage. Confirm current audit limits on Semrush and Ahrefs. Ahrefs still matters here as the publisher of the 96.55% zero-traffic study, which is the backdrop for URLs that never get indexed and therefore never get traffic.

Implementation for this class: crawl the money templates, export indexability / canonical / status, join to Search Console coverage, and fail the row. Do not request recrawl from a suite UI on a URL that is still noindexed in your own HTML.

Twelve-month crawl TCO

Print only figures a named source already published. Where a cell is not on that source, it reads contact vendor. Hours in the table are scenario loads for a weekly coverage-join ritual, not a promise of what your team will spend.

Cost lineScreaming FrogLog-aware cloud (JetOctopus / OnCrawl / Botify / Lumar)Semrush / Ahrefs suiteTicket layer (proposed)
LicenceFree 500 URLs; paid contact vendorcontact vendorcontact vendorcontact vendor
URL cap (free tier)500contact vendorcontact vendorn/a
Weekly crawls52525252
Analyst hours / month if joined by hand8663
Money URLs in the watch list (scenario)400400400400
Coverage-fail tickets / month (scenario)18181818
IndexCheckr average lag (context)27.4 days27.4 days27.4 days27.4 days
First-party BEST_OF earnn/an/an/a15.2% (12,514 pages)
Neutral seo_automation defaultn/an/an/a10
Recrawl windowdays–weeksdays–weeksdays–weeksdays–weeks

Twelve months of a cloud crawl you cannot date is still contact vendor times 12, which is not a number. Do not build a board slide from a remembered Botify SKU. Procurement closes from the vendor's current pricing URL, not from this table. The 8 versus 3 analyst-hour row is the operational claim: a human still reviews coverage fails, but a ticket layer is designed so the 400-URL watch list is not re-joined by hand every Monday.

Who should run this checklist

This page is for technical SEO leads, programmatic operators, and agencies who already ship URL templates, already can verify Search Console, and need a weekly pass so discovered-not-indexed money URLs cannot hide behind a healthy domain crawl.

Red flags: skip a paid cloud crawler if the host is still inside Screaming Frog's 500-URL free cap and someone actually crawls it. Skip a second suite audit if Semrush or Ahrefs Site Audit is already staffed. Skip a ticket layer if a spreadsheet already fails coverage rows and someone actually fixes robots, canonicals, or internal links. Skip recrawl-as-a-strategy if the HTML is still noindexed.

Who should choose a log-aware cloud crawler: teams with server or CDN logs and templates Googlebot may not be hitting. Who should choose Screaming Frog: teams with a bounded URL set and a laptop this afternoon. Who should stay on a suite audit: teams that will not open a fourth login. Who should choose none of the paid seats this week: a site that has not opened Search Console coverage yet.

Turning coverage fails into tickets

The failure mode is not "we lack a seventh crawler." The failure mode is a coverage CSV that nobody turns into a ticket. A proposed agentic workflow on US Tech Automations could take a weekly crawler export joined to Search Console coverage as the trigger, sync rows where coverageState is not indexed (or where the crawler Indexability is non-indexable on a money URL), and route each fail into a queue with the URL, the state, the robots/canonical evidence, and the export date — configurable, not a measured customer deployment. Prerequisites are Search Console API or CSV access, a crawler export, a URL list that matches the CMS, and a human review point before anyone requests recrawl. Output in the user's hands would be a ticket per failing URL, not a new crawl credit.

Worked example, not a live case: a 2,400-URL docs host with 400 money URLs, 18 URLs showing GSC indexStatusResult.coverageState as crawled-not-indexed for 21 days, and 9 orphans with zero internal links, could open 27 tickets (18 coverage + 9 orphans), hold a 5-day engineering SLA on robots/canonicals, and only then request recrawl inside Google's days-to-weeks window. The 2,400 URLs, 400 money URLs, 18 coverage fails, 21 days, 9 orphans, 27 tickets, and 5-day SLA are scenario numbers meant to size the workflow; they are not a promised index rate and not a customer result. The backticked coverageState field is a Search Console URL Inspection / coverage concept; Indexability is the Screaming Frog export column.

A second proposed US Tech Automations path, still configurable, is a webhook from the same join into the team's existing tracker so the queue is not a screenshot of Site Audit. The trigger is the file drop or API pull; the action is a deduped ticket keyed on URL plus export week; the output is a list a human can fix, snooze, or escalate. Idempotency matters: the same URL must not open a second ticket if coverageState did not change. Access control matters: the Search Console token stays in a secret store. Retention matters: coverage snapshots are not a substitute for the 16-month Performance window, so keep both.

n8n versus a ticket layer

The honest alternative to a dedicated ticket layer is not "do nothing." It is a Zapier, Make, or n8n scenario that watches a crawl CSV and a coverage export, retries on a 429, branches when coverageState is not indexed on a money URL, and writes an audit row. Those tools can support run histories, retries, error branches, and audit evidence when someone deliberately designs them. The buyer then owns observability, idempotency, escalation, access controls, retention, and maintenance for as long as Google or the crawler change an export shape.

A proposed US Tech Automations design does not remove those needs. It would configure the same trigger → join → ticket path with named human review points and a single place to see which export week created which coverage ticket. That is a packaging difference, not a claim that no-code cannot retry. If your n8n flow already fails coverage rows, closes duplicates, and pages an owner, you do not need a new system of record. When NOT to use US Tech Automations: when Search Console coverage on one property is the whole job and someone already closes it; when Screaming Frog's 500-URL free crawl plus a spreadsheet already fails the rows; or when a staffed n8n scenario already owns the join, the threshold, and the ticket with retries and an audit log you trust.

FAQ

What is an indexing SEO audit checklist?

It is a repeatable pass over robots, canonicals, sitemaps, orphans, Search Console coverage, logs, and recrawl so you fail URLs Google cannot keep, rather than a copy review.

How long does Google take to index a new page?

IndexCheckr's study of 16 million pages (updated 28 Feb 2025) found an average 27.4 days to index, with 14.00% indexed within 7 days and 64.86% within 30 days; your host can be faster or slower.

Do I need Botify if I already use Screaming Frog?

Not if the host fits the crawl you actually run and you do not need log-file join; Botify, OnCrawl, JetOctopus, and Lumar earn their seat when Googlebot's real path is the question.

Should I request recrawl as soon as I publish?

No. Fix robots, canonicals, internal links, and coverage first; Google Search Central still describes recrawl after a request as days to weeks, not minutes.

Why are 61.94% of pages not indexed in IndexCheckr's corpus?

Because publish is not indexation: noindex, orphans, duplicates, thin templates, and crawl waste keep URLs out — treat that 61.94% as a warning, not as your forecast.

Can n8n replace a ticket layer for coverage fails?

Yes, if you deliberately design retries, error branches, audit evidence, idempotency, and an owner; a ticket layer is optional packaging for teams whose coverage CSV never becomes a queue.

Close with coverage, not a dashboard

An indexing SEO audit is seven checks in order, not seven logos on a slide. Search Console is the system of record for what Google kept. A crawler is the system of record for what your HTML allows. IndexCheckr's 27.4-day average and 61.94% not-indexed share are why a publish button is not a strategy. Screaming Frog's 500-URL free cap is why a paid cloud crawl is optional at small URL volume. The useful operating loop is dated crawl + coverage → fail → human fix → recrawl request — in that order. If you want the exception path packaged as a configurable trigger, queue, and webhook with a human review point, compare that design on the agentic workflows page and on pricing against the Zapier, Make, or n8n flow you could own yourself. Read more on the resources blog if you are still mapping the rest of the technical stack around coverage.

About the Author

Garrett Mullins
Garrett Mullins
Workflow Specialist

Helping businesses leverage automation for operational efficiency.