Crawl and data integrations

Every tool on this list is only as good as what it reads. The integrations are the connections to Search Console, analytics, crawl data, keyword tools and live search results, plus the rules about which source is trusted for what.

Built in-house. It is my own tooling rather than a reseller’s dashboard, which is why it checks the things I act on rather than everything a generic tool can measure.

What it does

It connects the data sources directly rather than through exported spreadsheets, so a check runs against current data and can go deeper than an export allows. It also holds the awkward knowledge: which source caps its reports, which one inverts a comparison, and which number is a floor rather than a total.

That second part is the reason this page exists. Most bad SEO reporting comes from treating a figure as more solid than it is.

The sources

Source What it’s used for What it’s not used for
Google Search Console Impressions, clicks, position, indexing status, live URL inspection Total traffic, since it covers Google organic only
GA4 Sessions, conversions, channel behaviour including the AI assistant channel Ranking data
Crawl data Site structure, status codes, canonicals, duplicates, internal links Anything about demand
Keyword tools Search volume, competitive context, competitor visibility Precise traffic forecasts
Live search capture What a results page actually looks like, including AI Overviews and features Volume
Raw HTML capture The head section, schema and what a crawler receives before JavaScript Rendered behaviour

Why raw HTML matters

Many tools convert a page to plain text before reading it, and that strips the head section entirely. The title, meta description, canonical, hreflang and schema all live there, so any check built on converted text is blind to the elements that decide how a page appears in results.

Reading the raw HTML is the only reliable way to report on them, and it’s why a separate capture route exists rather than relying on whatever a text conversion returns.

The caps and quirks worth knowing

Search Console truncates. Query and page reports have row limits, so on a large site a total is a floor rather than a complete figure. A report claiming an exact sitewide number from a capped export is overstating what it knows.

Long URLs can truncate in the interface, which makes them look like different pages than they are.

Period comparisons need checking for direction. More than one tool presents a comparison in the opposite sense to the one you’d assume, and a reversed sign turns a decline into a recovery in a client report.

Crawlers get rate limited. A 4xx from a crawl is sometimes the server defending itself rather than a broken page, so status codes get rechecked before they’re reported.

Keyword volume comes from keyword tools, not from live results. Live capture shows what a results page looks like, and it can’t tell you how many people searched.

What I review

Any figure going in front of a client is checked against its source and, where it’s a claim about a live page, against the live page. Where a number is a floor rather than a total, the document says so rather than presenting it as exact.

Where two sources disagree, the document says which was used and why, rather than picking the more flattering one.

Limitations

  • Access depends on the client. Search Console and GA4 both need granting, and without them the work runs on external estimates.
  • Third-party estimates are estimates, useful for relative comparison and poor as absolutes.
  • Live capture is a snapshot. Results pages vary by location, device and over time.

Services it powers

The integrations sit underneath everything, and most directly behind reporting and competitor reporting. See the rest of the tools on the tech stack page.

FAQ

Do you need access to my Search Console?

It makes the work considerably better, because it’s the only source for what your site actually earns in Google rather than what a third party estimates. Without it the analysis runs on external tools, which are useful for comparison but poor as absolute figures.

Why do two SEO tools give different numbers?

Because most of them estimate. Keyword volumes and traffic figures from third-party tools are modelled rather than measured, so they differ by method. Search Console and GA4 report your own data, which is why they’re preferred whenever a figure needs to be accurate.

Is Search Console data complete?

No, and that’s worth knowing before quoting it. Its query and page reports have row limits, so on a large site the totals you can export are a floor rather than the full picture. Any report I produce says so rather than implying precision.

See the whole toolkit

The tech stack page lists every tool, and the free assessment uses these sources directly.

Get started

Proof, not promises.

Tell me what's going on and I'll come back with what I'd do first. No obligation, and no report you need a translator for.

Over 925,000 impressions and 13,795 clicks in a year, from people asking one awkward question: what size do I need?condoms.uk, past 12 months

Brands and sites I have worked on

Brands worked on across nearly twenty years, in agency roles and directly.