NegativeSEO.ICU logo — negative SEO reference and recoveryNegativeSEO.ICUNegative SEO reference & recovery
Detecting an attack

What each category of tool can see, and the three things none of them can

Most confusion in this subject comes from asking a tool the one question it cannot answer. Here is which question each category actually answers.

The organizing question

Every tool in this subject answers exactly one of three questions, and almost all the confusion comes from asking one of them a question belonging to another.

  1. What did Google decide? Only Google can tell you. First-party consoles, and nothing else.
  2. What exists in the world? Third-party crawls, backlink indexes, archives.
  3. What actually happened at my server? Logs, uptime records, file integrity.

No tool answers all three, and no third-party tool answers the first one at all. That sentence is the most useful thing on this page, and it disposes of an entire product category on its own: any product claiming to detect a Google penalty is inferring from ranking movement, and ranking movement cannot distinguish a penalty from an update, a technical fault or a competitor improving.

Tools are named below only where the name is informative about a category. Nothing here is rated, ranked or recommended, no purchase is advised, and feature sets change constantly enough that any specific claim about a commercial product should be checked against the live tool before you rely on it.

First-party consoles: the only things that see Google's decisions

Google Search Console is free and unsubstitutable, and these are the observations that exist nowhere else:

  • Manual actions. Google's documentation states that it "issues a manual action against a site when a human reviewer at Google has determined that pages on the site are not compliant with Google's spam policies," and that in that case "some or all of that site will not be shown in Google search results." No third-party tool can detect one. Not by inference, not by pattern, not at all.
  • Security issues - Google's determination that "your site was hacked, or that it exhibits behavior that could potentially harm a visitor or their computer."
  • Google-selected canonical - "The page that Google selected as the [canonical] URL when it found similar pages on your site." This single field is the most decisive piece of attack evidence available anywhere, because it is Google stating in the first person which URL it treats as the authority for your content. If it names a host you do not control, you have a hijack, and no crawler or link index can produce that finding.
  • Google's own click and impression counts - not a sample, not a model, and unaffected by whatever your analytics is doing.
  • Crawl stats and robots.txt fetch history - ninety days of requests, response codes, download size and average response time, plus which robots.txt files Google found and when it last crawled them.
  • Removal requests filed by non-owners, in the Removals report's Outdated content tab.

Bing Webmaster Tools belongs in the same tier for a reason that has nothing to do with Bing traffic: it is a first-party report from a different index. If a pattern appears in both, the cause is likelier to be on your site; if only Google moved, the cause is likelier to be a Google ranking change. That comparison costs ten minutes of verification and no money.

Where Search Console falls short, and owners over-trust it

It is not an oracle, and four of its limits cause real misdiagnosis.

The Links report is a sample. Google says it "isn't a comprehensive list of every link on your site," tables cap at 1,000 rows against a 100,000-row landing-page export, and "The report doesn't specify if a link is marked as nofollow," so follow status cannot be derived from it at all. It also retains links whose pages have long since gone.

There is no competitor data, which is why the decisive "did competitors move too" test needs a tool from another category entirely.

There is no history before verification. The bulk data export does not backfill: historical data preceding your setup requires other methods. If you verify the property after the incident, the console cannot tell you about the incident.

It reports outcomes and never causes, and its most recent days are provisional - "The newest data can be preliminary, meaning it's still being collected and might change in the next few hours." Average position is an average, not a position, so it will not settle an argument about where a page ranked on a given morning in a given city.

Third-party backlink indexes

A backlink index is a vendor's own continuously refreshed crawl of the web, from which it reconstructs which pages link to which. Ahrefs, Semrush, Majestic and Moz each maintain one; Common Crawl is an open crawl that others build on. Named as category exemplars, with nothing implied about their relative merit - because their differences matter far less than what they all share.

What they can see: links Search Console's sample does not surface; anchor text across the whole discovered profile, which is the report that actually matters for triage; competitor link profiles, which Search Console cannot show at all; and a discovery timeline useful for building a dated record.

What they cannot see, and this list is the point:

  • Google's index. No vendor knows what Google has crawled, kept, consolidated or discarded. Every index is a sample of its own crawl.
  • Whether Google counts a link. Given the documented shift to devaluation in 2016 and the nullification of spam links since, the population of links that exist and the population of links that matter are different populations, and nothing outside Google can tell them apart.
  • Creation dates. The date axis is discovery - when the vendor's crawler found it. A vertical spike is frequently a crawl batch rather than an event, and this is the leading cause of false alarms in the whole subject.
  • Whether a link is disavowed. Disavow files are not published and are visible to no one outside Google.
  • Each other. Two vendors will report materially different referring-domain counts for the same site. Neither is wrong; they crawled different things. A report citing one vendor's count as the number is presenting a sample as a census.

The toxicity score, stated plainly

Several link vendors attach a per-link or per-domain toxicity score, spam score or equivalent, and build monitoring and cleanup products around it. Start from what Google's policy actually prohibits, because the mismatch is structural rather than a matter of accuracy.

Google's spam policies define link spam as "the practice of creating links to or from a site primarily for the purpose of manipulating search rankings," and every example listed is conduct by the site owner: buying or selling links, excessive exchanges, "Using automated programs or services to create links to your site," advertorials, widget links, "Widely distributed links in the footers or templates of various sites," and "Forum comments with optimized links in the post or signature."

So the policy addresses who created a link and why. A toxicity score grades how a link looks to a vendor's crawler. Those are different objects, and no arithmetic converts one into the other. A link built by your competitor to frame you and a link built by you to cheat can look identical to a crawler and are treated oppositely by the policy - which is precisely the distinction the score cannot make and the policy is entirely about.

These numbers are vendor opinions, computed from the vendor's own crawl using the vendor's own undisclosed heuristics, with no defined relationship to anything Google publishes. Google defines no per-link or per-domain toxicity metric in its documentation, publishes none, and exposes none in Search Console. That is an absence-of-evidence finding rather than a quotation, and it should be read as one. The nearest Google comes to describing its own assessment is in the disavow documentation: "In most cases, Google can assess which links to trust without additional guidance, so most sites will not need to use this tool" - an internal assessment, not a published score.

The most direct on-record comment from Google about the businesses built on these scores is John Mueller's, on 31 January 2023: "That's all made up & irrelevant. These agencies (both those creating, and those disavowing) are just making stuff up, and cashing in from those who don't know better," with the advice to spend the time building the site up instead (reported by Search Engine Journal, 2 February 2023).

Two things follow. A rising toxicity score is not evidence of a problem, and a falling one is not evidence of a recovery - both are evidence about a model. And acting on the score with a disavow file is the mechanism by which a non-event becomes real damage, because Google says that tool "can potentially harm your site's performance in Google Search results" if used incorrectly. The same reasoning applies to third-party authority metrics - Domain Authority, Domain Rating, Trust Flow and their relatives. Useful as rough comparative shorthand, meaningless as evidence of a Google judgment, and absent from every Google report.

Use these tools to build a list to look at. Never use one to build a list to act on.

Rank trackers and volatility monitors

What they see: positions for specified queries, from specified locations and devices, on a schedule - and, in the ones configured to store it, the whole result set. That stored top ten is what answers "did competitors move too," which is the decisive test when separating an update from an attack, and it is the reason a tracker is worth running before anything goes wrong rather than after.

What they cannot see: your users' actual results. Every reported position is a sample from one configuration, stripped of personalization, history and context. Two trackers will disagree about the same query on the same day and both will be defensible.

Volatility trackers measure the vendor's own fixed keyword panel. They answer "did this market move on that date," which is a real question. They are not announcements, they say nothing about your site specifically, and treating a volatility spike as confirmation of an unannounced update is inference stacked on inference. The authoritative list of announced updates is Google's Search Status Dashboard, and the procedure for using it is in attack or algorithm update.

The configuration decision that determines whether a tracker is useful at all: one that stores only your own position is nearly worthless for this subject; one that stores the top ten is a contemporaneous record of the market.

Crawlers, log analyzers and change monitors: aimed at the attacks that work

This is the group the protection market talks about least, because these tools find your own problems rather than someone else's malice - and it is where the vectors that reliably work become visible.

Site crawlers and auditors. Screaming Frog and Sitebulb are desktop crawlers; several hosted platforms include auditors. They surface pages you did not create, redirect chains and unexpected off-domain redirects, off-domain canonical tags, stray noindex directives and response-code faults. Crucially, configured with Googlebot's user agent they show the difference between what a browser receives and what a crawler receives, which is how server-side cloaking on a compromised site becomes visible. What they cannot see is anything about Google's decisions: a crawler tells you what your site says, and only the console tells you what Google did about it.

Log analyzers see ground truth, and they are the only tools here that produce evidence in the forensic sense: which address, with which user agent, requested which URL, at what time, and what status code came back. Use them to date a crawl-side fault, find requests for URLs that do not exist on your site, measure request floods, and separate real from forged crawlers. Verifying Googlebot has two documented methods - a reverse DNS lookup on the accessing address confirming it resolves to googlebot.com, google.com or googleusercontent.com followed by a forward lookup confirming it returns the same address, or matching the address against Google's published ranges. Anything failing both is not Googlebot, which is the whole of the spoofed Googlebot question. Read the current range-file names off Google's live documentation rather than any list republished elsewhere; they have changed.

The constraint that defeats this category is never technical. It is that the logs have usually already been rotated away by the time anyone thinks to look.

Uptime, response-time and change monitors see availability, latency and modifications to files, templates, configuration and rendered pages. Visual and HTML diff monitors catch conditionally served content that file monitoring misses. What none of them can see is intent: a change monitor reports that something changed, never who changed it or why.

Copy detection, blacklist scanners, and the tools counsel cares about

Copy detection. The free version is a verbatim phrase search in quotation marks on a distinctive sentence from the affected page, and it resolves most cases; Copyscape is the long-standing commercial example of the same idea at scale. What none of them can see is which copy Google treats as canonical, which is the only question that matters - and that answer lives in the Google-selected canonical field and nowhere else. The uncomfortable finding on content scraping is that a scraper outranking an original is usually a symptom of the original's weakness rather than proof of an attack.

Malware and blacklist scanners. For Google specifically these are redundant: Search Console's Security Issues report is first-party and free, and Google's own site-status lookup sits in its Transparency Report. What is not redundant is the non-Google lists - browser vendors, mail reputation lists, hosting-side blocklists - which Search Console does not cover and which have separate delisting processes.

Evidence and archive tools are not SEO tools and matter most if a case ever becomes legal. The Internet Archive's Wayback Machine holds dated snapshots of your pages, competitor pages and attacking pages, and is frequently the only surviving record of a page that has since been deleted. The Lumen Database publishes takedown notices submitted to participating platforms, including the notice text, which is how a fraudulent copyright takedown is identified and attributed. Google's Transparency Report covers copyright removals and Safe Browsing status. And your own dated exports and screenshots outrank all of them, because nothing produced after the fact is worth as much as something captured before it.

The category that is not a tool

Products marketed as negative SEO protection, toxic backlink protection or penalty insurance are a marketing category rather than a technical one. There is no mechanism by which a subscription prevents a stranger from publishing a link to your site. What genuinely reduces exposure is ordinary security hardening of your own site - patching, credential hygiene, access control, backups, file integrity - which is security spend under its correct name. Which continuous services earn their keep is set out in negative SEO monitoring.

The recurring mistakes, in the order they cost the most: asking a third-party tool what Google decided; treating a toxicity or authority score as a Google signal; acting on a vendor score with a disavow file; reading a link tool's date axis as creation dates; citing one vendor's referring-domain count as a fact; using site: as an index measurement when Google's own documentation disqualifies it; reading URL Inspection's default result as a live test when Google says plainly it is not; buying a hosted auditor and never once running a crawl as Googlebot, so never seeing the cloaked version of your own site; and losing the logs, which is the one irreversible tooling failure in the subject.

This site sells negative SEO recovery, which is a commercial interest in the opposite conclusion from most of the above. The conclusion stands anyway, and if a tool vendor's dashboard is the only thing telling you that you are under attack, you are not yet entitled to believe it.

Frequently asked questions

Is there a tool that tells me whether Google has penalized my site?

Only Search Console, and only for the thing Google actually calls a penalty. A manual action exists in Google's systems and is reported to the verified owner in the Manual actions report and the message center. Every third-party penalty checker is inferring from ranking movement, which cannot distinguish a penalty from an update, a technical fault or a competitor improving.

Which backlink tool is the most accurate?

The question does not have an answer, because there is no shared reference to be accurate against. Each vendor reports its own crawl, the counts differ substantially between them, and none of them knows what Google kept or discarded. Use one for anchor text at scale and for a dated discovery record, and treat any single count as one sample rather than a fact.

My tool gives every link a toxicity score. Can I trust it?

Trust it as a way to sort a long list into an order for a human to look at. Do not trust it as a Google signal: Google defines no such metric, publishes none, and exposes none in Search Console, and its policy turns on who created a link and why rather than on how the link looks to a crawler. No vendor has published a reproducible study connecting a toxicity score to a measured Google outcome.

What free tools actually matter?

Search Console, Bing Webmaster Tools as a second first-party opinion from a different index, your own server logs, a verbatim phrase search for duplicated content, and the Wayback Machine for dated evidence. Between them they answer what Google decided, what happened at your server and what a page said last year - which is most of the diagnosis.

Can a tool tell me who attacked me?

Almost never. Link indexes show pages, not people; change monitors report that something changed, never who changed it. The two places attribution genuinely surfaces are server logs, which record the requesting address and agent, and published takedown notices, which carry the submitter's own claims. Everything else is inference, and inference is not something to put in a legal letter.

Top