NegativeSEO.ICU logo — negative SEO reference and recoveryNegativeSEO.ICUNegative SEO reference & recovery
Abstract nested ring illustration representing Canonical Hijacking
Content & Platform AttackYour content and who owns it

Canonical Hijacking

Situational Works only under specific conditions, and rarely otherwise.

A tag that is a hint rather than a rule can be aimed at you - but the same fact is why it usually fails.

What canonical hijacking is

Canonical hijacking is an attack in which a copy of your page, hosted on a domain the attacker controls, uses the rel="canonical" element to contest which version of the content a search engine treats as the authoritative one. Hijacking, in the search sense, means taking over the identity of something you do not own - here, the identity of a page - without ever touching the server it lives on.

The tag itself is ordinary infrastructure. A URL - the address at which a page is reachable - is not the same thing as the content at that address, and the same content is usually reachable at several addresses at once: with and without parameters, on a print version, on a syndication partner's site. The rel="canonical" element is a line in a page's <head> naming which of those addresses is the preferred one, so the set is consolidated rather than all indexed separately. When the named address sits on a different domain from the page declaring it, that is a cross-domain canonical, and Google supports it for legitimate syndication.

Two directions of abuse exist, and they are constantly confused with one another. Keep them apart, because they need different diagnoses and different responses:

  1. Attribution theft. The attacker copies your page onto their domain and points the canonical at their own URL, or omits it entirely, hoping Google clusters the two pages and selects theirs. If Google does, your URL is the one that drops out of the results.
  2. Signal poisoning, or the reverse canonical. The attacker copies your page including your canonical tag, so their spam page declares your URL as the canonical. The hoped-for effect is that the spam page's characteristics attach to your address. This is the variant reported in 2018, and it is the one Google disputes on mechanism.

Neither requires access to your server. That is the appeal: the entire attack runs on infrastructure the attacker owns. It is also, as it happens, the reason your defenses are mostly on your own side of the fence.

This page covers deliberate manipulation of the canonical element. If a copy is simply outranking you, that symptom has its own page; if the question is whether being copied costs you anything, see content scraping; if the copy updates the instant you edit your site, it is reverse-proxy hijacking.

Hint, not a rule - the sentence that cuts both ways

Google's canonicalization documentation says this, and it is the most consistently misread sentence on the topic:

"You can indicate your preference to Google using these techniques, but Google may choose a different page as canonical than you do, for various reasons. That is, indicating a canonical preference is a hint, not a rule."

That is quoted from Google Search Central's canonicalization documentation, and it is normally produced as evidence that the attack works. It is evidence of exactly half of that. Yes: because the tag is only a hint, Google can override your canonical, which is what makes the attack conceivable at all. But the same sentence means Google can override the attacker's, and that is what happens in the overwhelming majority of cases.

Think of it as a contest rather than an instruction. Google groups duplicate pages into a cluster and picks the one that, on the signals its indexing collected, is the most complete and useful. It weights those signals unequally - redirects strongest, rel="canonical" annotations strong, sitemap inclusion weak, plus site-level preferences such as HTTPS. A third party's tag is one input into a contest your original normally wins on age, internal linking, inbound links and host authority. An attacker asserting a canonical is not issuing an order. They are filing a claim, and claims get weighed.

It is worth adding what John Mueller of Google said about the losing case, reported in September 2019: he framed selection as two questions - which URL does it look like the site wants Google to use, and which version would be most useful to searchers - and said there is no negative impact on rankings when Google picks a canonical other than the one you would prefer. Most people who arrive at that Search Console message are not under attack and are not being penalized.

The 2012 experiment, including the ending nobody quotes

The only rigorous, method-disclosed public demonstration of this attack is fourteen years old. In November 2012 Dan Petrovic copied the HTML of several target pages onto subdomains of his own domain and pushed internal link equity at them. The copies displaced the originals in Google's results - including, notoriously, for the query "Rand Fishkin," where a days-old test page took the top position across four countries. Three other sites were hijacked the same way.

His conclusion was that Google takes rel="canonical" as a hint rather than a directive, and that where the same content is crawled under several URLs, the one with the highest link equity is the one that appears. Note the detail most often dropped: a canonical on the original seemed to prevent the hijack succeeding fully, but not in every case. Partial protection, not immunity.

And then the ending, which is almost always omitted when this experiment is cited as proof the attack is live. Google detected it and acted against the researcher. His domain received a search quality notification for copied content and the hijacking pages were removed from the index. So the experiment demonstrated two things at once: that the mechanism can work, and that Google's spam enforcement reverses it and penalizes the party running it. Anyone citing the first half without the second is selling something.

Fourteen years of scraper detection have been built since. No comparably rigorous replication showing this attack succeeding against an established site has been published in the intervening period. Absence of published replication is not proof the attack is dead - nobody has run the controlled test either way - but it is the reason this page carries a situational verdict rather than a documented-threat one, and I would rather state the reason than hide behind the pill.

The 2018 case I published, and Google's public disagreement with it

This one is mine, so I will state its limitations as plainly as its findings.

On 19 April 2018 I published an account of a negative SEO attack I had found while investigating a client's sudden ranking loss. A spam page on a domain I did not control had copied the victim page's entire <head> section, the victim's own rel="canonical" element included, so that the spam page declared the victim's URL as its canonical. The rest of that page was unrelated content, much of it adult-themed. My hypothesis was the signal-poisoning direction described above: that Google would associate the two pages through the canonical and carry the spam page's characteristics across.

How it was found is the methodologically interesting part, because the attack left almost no trace in the usual places. Search Console showed the ranking loss but nothing about its cause. A source-code search engine did not surface the relationship at all. A backlink index did, because it indexes canonical-tag data alongside link data, and because the attacker had linked outward to other domains they controlled. One operational lesson worth keeping: when a copy is invisible to the tools that crawl pages, the tools that crawl links may still see it.

Google's response was a rebuttal, not a confirmation, and the honest version of this story says so. John Mueller responded publicly within days. Search Engine Roundtable dates the first reply to 21 April 2018: "The premise that rel canonical combines pages is wrong. That's not how it works, it's either one or the other." The second, dated 22 April: "The rel canonical has been around for over a decade, people have tried lots of things with it. It's a signal for canonicalization; one URL wins, the others' crawls get dropped." Both are reported by Search Engine Journal, 20 April 2018 and by Search Engine Roundtable, 23 April 2018.

Two further things belong here rather than in a footnote. Search Engine Journal, which credited me with the finding, also stated plainly that the exploit had been documented but never tested or verified experimentally. And Barry Schwartz was skeptical of the framing, writing that the tactic was not actually novel. Both are fair. I forwarded details to Google and received no response; there is no bug report, no bounty and no fix to point at.

The numbers I reported at the time - a burst of new links on 11 April 2018 to a site that ordinarily accrued a fraction of that daily, carrying zero trust against a substantial citation score, and one keyword falling from first position to the seventies - are my own observations of one site, never independently verified by anyone, including the outlets that repeated them.

In April 2019 I published a follow-up on a different case, in the opposite direction: there a copy on a more authoritative domain caused Google to select the attacker's URL as canonical, deindexing the original - attribution theft rather than signal poisoning. The useful artifact from that one is a Search Console string worth memorizing: "Duplicate, submitted URL not selected as canonical."

What the disagreement actually turns on

The 2018 exchange is worth understanding rather than picking a side in, because the two positions are not describing the same thing.

Mueller's rebuttal is substantive, and its load-bearing phrase is "either one or the other." In Google's account, canonicalization selects a URL from a cluster; it does not merge two pages' properties into a single record. If that is right, there is no channel through which the spam page's characteristics could flow into the victim's URL - the spam page simply loses the cluster, its crawls get dropped, and nothing about it reaches you. Under that model, a canonical aimed at your URL is not an attack. It is an attacker forfeiting.

What I had was a correlation: a page that lost rankings, and a spam copy declaring that page as canonical, found in the same investigation. A correlation of one, in a live environment where dozens of variables move at once, is not a mechanism. That is the honest weakness in my own account, and it is why the exchange was never resolved - neither side ran a controlled test, and nobody has published one since.

So this page does not tell you Google confirmed anything, and it does not tell you Google fixed anything. No source says either. What the record supports is narrower and more useful: the attribution-theft direction is demonstrated, has clear conditions, and has a diagnostic screen you can check yourself; the signal-poisoning direction is a disputed claim that Google's named spokesperson denied on mechanism, on the record, with a dated explanation of why. If you are triaging a site, act on the first and treat the second as unproven.

The two fields that settle it in your own account

Almost every false alarm on this subject starts with somebody reading a summary and stopping there. The diagnosis lives in two fields, side by side.

Open Search Console, go to Pages, and look for "Duplicate, Google chose different canonical than user." That message means Google saw your declared canonical and overrode it. It does not mean you were attacked. Then inspect an affected URL with the URL Inspection tool and read two fields together: user-declared canonical, which is what your tag says, and Google-selected canonical, which is what Google decided. The relationship between those two fields is the entire diagnosis.

  • Google-selected canonical is your own URL. Nothing has been hijacked, whatever any third-party tool says. A scraped copy declaring a canonical somewhere else has done nothing to you.
  • Google-selected canonical is a different URL on your own site. Internal duplication - parameters, a staging domain left indexable, an HTTP and HTTPS split, a www and non-www split, a country variant, a CDN address. This is by far the most common cause and it is self-inflicted.
  • Google-selected canonical is on a domain you do not control. Now you are looking at a live canonical hijack, or a legitimate syndication partner who is winning the cluster. Establish which before you treat anyone as an attacker.

Corroborate with a search: put a long, distinctive verbatim sentence from the page in quotation marks. If a copy on another domain returns and your original does not, the cluster has resolved against you. If your original returns first, there is no hijack. Check the manual actions report too - it should be empty, because a canonical hijack produces no manual action against the victim.

In my judgment the large majority of reported canonical hijacks are one of the mundane causes in the second bullet. Ruling them out costs an hour. Acting on the wrong one costs a quarter.

The conditions under which it genuinely works

This is a contest of authority between two copies, so it succeeds only where the attacker can win that contest. That is a much shorter list than the fear around this subject suggests:

  • The target page is new, with little link equity and no crawl history. There is nothing yet for Google to prefer it on.
  • The target is thin, or the content is short, generic and close to a dozen other pages. A cluster of near-identical thin pages is decided on host signals, and the attacker may have better ones.
  • The attacker's host is substantially more authoritative. The 2012 results depended on precisely this asymmetry, and a hacked page on a strong domain supplies it for free.
  • The original is slow to be crawled and the copy is discovered first. For a new page on an infrequently crawled site this is a real risk, and it means indexation speed protects you more than any tag does.
  • Your own signals contradict each other - no self-referencing canonical, the same content on several of your URLs, a sitemap and an internal link structure naming different addresses. Google is then guessing, and a guess is a coin you handed to somebody else.

One more condition belongs on the list and is not an attack at all: legitimate syndication where the partner is stronger than you. This is the most common real-world instance of "a copy is outranking me," and Google made it more consequential in May 2023 when it withdrew its recommendation to manage syndicated non-news content with cross-domain canonicals, on the ground that syndicated pages are often very different, and pointed publishers toward blocking indexing instead. One useful side effect: with the legitimate use of cross-domain canonicals narrowed, an unexpected inbound one is more anomalous than it used to be, and easier to argue about.

Winning the cluster back

In order, and the order matters because steps two and three are what actually decide the contest:

  1. Confirm with URL Inspection that Google's selected canonical is a URL on a domain you do not control. Without that, stop - you are about to remediate the wrong problem.
  2. Fix your own signals first. Every page carries a self-referencing canonical, written as a fully-qualified address rather than a relative one. Internal links, the XML sitemap and the canonical all name the same URL. No contradictory signals anywhere. The page is reachable, fast, and returns a normal response. Most losses in a duplicate cluster are self-inflicted, and this step fixes them.
  3. Strengthen the original's claim. Get the page crawled promptly, and make new content discoverable quickly. Being indexed first is worth more than any tag you can add afterward.
  4. Report the copy as scraped content under Google's spam policies. The 2012 case shows Google does act against copying sites, on its own timetable and with no case-by-case feedback.
  5. If the copy reproduces your content wholesale, a copyright removal request against the copy is the fastest lever available. It removes the competing URL from the index rather than arguing about canonical selection. Use it where the copying is genuine and you own the copyright - and note that the same statute abused in fraudulent DMCA takedowns is the legitimate remedy here, with the same liability attaching to a notice you know to be false.

What does not help: disavowing links from the copying domain, because canonical selection is not a link-penalty mechanism; filing a reconsideration request, because no manual action exists against you; adding noindex to your own page; removing your canonical tag; and rewriting the page, which surrenders your position in the cluster and can cost you rankings you still hold. Broad crawler blocking in robots.txt is the most damaging overreaction of the set: it can stop Google seeing your canonical at all, which makes your claim strictly weaker.

And doing nothing is frequently correct. If your original still ranks and Search Console's Google-selected canonical is your own URL, a scraped copy pointing a canonical anywhere it likes has done nothing to you and needs no response. Copies of any page that performs well are permanent background noise on the web. Treating each one as sabotage is a reliable way to spend a quarter and change nothing.

Where the false alarms come from

The errors on this subject are unusually consistent, and most of them cost money.

  • Reading "Duplicate, Google chose different canonical than user" as proof of an attack. It overwhelmingly means internal duplication, URL parameters, or a syndication partner you agreed to.
  • Believing that a canonical tag is an instruction. It is a hint - stated as such in Google's own documentation - which is why an attacker's tag is not an order either.
  • Believing a duplicate content penalty exists. There is no such penalty. Losing a duplicate cluster is a selection outcome, not a sanction, and nothing about it appears in the manual actions report.
  • Citing the 2012 experiment as proof the attack works today while omitting that Google reversed it and penalized the researcher.
  • Retaliating by copying the attacker or canonicalizing at them. That is the conduct Google's spam policies target, and 2012 shows who gets penalized when it does.
  • Publishing the same content across several domains you own without deciding which one is canonical, then reporting the entirely predictable result as sabotage.
  • Buying protection against it. There is no product that stops a third party from writing a tag on their own server. What actually reduces exposure is unglamorous and free: self-referencing canonicals, consistent internal signals, and getting new pages indexed quickly.

Frequently asked questions

Another site has a canonical tag pointing at my URL. Is that hurting me?

On the available evidence, probably not, and the burden of proof sits with anyone claiming otherwise. Google's position, stated by John Mueller in April 2018, is that canonicalization selects one URL from a cluster rather than combining two pages, so a spam page declaring your URL as canonical is forfeiting rather than attacking. I published the case that prompted that exchange and I still cannot demonstrate the mechanism - nobody has run a controlled test either way. Check the Google-selected canonical for your own page. If it is your URL and your rankings are intact, treat the copy as noise.

How do I tell an attack from a syndication partner outranking me?

Ask whether you agreed to it. A syndication partner is a copy you authorized, usually on a stronger site, and the fix is a conversation plus a request that they block indexing of the syndicated version - which is what Google steered publishers toward in May 2023 when it withdrew its cross-domain canonical recommendation for non-news content. An attack is a copy you never agreed to on a domain with no relationship to you. The Search Console symptom is identical in both cases, which is exactly why you have to check who owns the winning URL before deciding what happened.

Should I disavow the domain hosting the copy?

No. The disavow tool addresses links being counted toward your site; canonical selection is a completely different mechanism and disavowing cannot influence it. You would be spending effort on the wrong system and taking on the risk of a carelessly built disavow file for no possible benefit.

Should I remove my canonical tag so nobody can copy it?

No - that makes your position worse. A self-referencing canonical is one of the signals telling Google which URL you want indexed, and removing it means Google decides without your input. If someone copies your head section wholesale, the tag they take is evidence of copying, not a weapon you handed them.

How fast does the original come back once the copy is gone?

It depends on recrawl, and there is no published timeline. Google has to fetch both URLs again before it can revise the cluster, so the pace is set by how often it crawls the pages involved - days for a frequently crawled site, considerably longer for one it visits rarely. Requesting indexing for the affected URL is worth doing; assuming a date is not.

Is any of this worth paying someone to fix?

Sometimes, and the deciding question is what URL Inspection shows. If the Google-selected canonical is your own URL, there is nothing to fix and no engagement worth buying. If it names a domain you do not control, the work is real - establishing authorship, correcting your own signals, and running a copyright removal or spam report against the copy - and it is a matter of weeks rather than months. If it names another URL on your own site, that is a technical cleanup, and it is the most common outcome by a wide margin.

Top