What is actually happening when a copy outranks you
Plagiarism - the reproduction of somebody else's words as your own - becomes a search problem at exactly one point: when the copy appears above the original for a query the original was written to win. That is the failure everything else in this subject area is a preamble to, and it is the one that costs money.
The mechanism is canonical selection. Google groups pages it judges to be duplicates or near-duplicates into a cluster, then picks one URL - the canonical - to represent that cluster in results. The rest of the cluster is suppressed for that content. So when a copy wins, your page is not ranked lower. It is filtered out for that body of text, while continuing to rank normally for everything else it is about. From the outside those two outcomes look identical, and treating a filtering event as a penalty is how people end up filing reconsideration requests against a manual action that was never issued.
The copy may belong to a scraper, a competitor, a freelancer reselling text you paid for, a syndication partner behaving exactly as your contract permits, or a machine-assembled rewrite close enough to cluster with yours. Only some of those are attacks. All of them produce the same symptom, which is why the diagnosis below comes before the response.
The Google-selected canonical: the one screen that settles it
Search Console's URL Inspection tool reports, for any URL on a property you own, a field labeled Google-selected canonical. It tells you which URL Google actually chose to represent the cluster your page is in. Alongside it sits User-declared canonical: the URL your own page nominates through its rel="canonical" element. The relationship between those two values is the diagnosis, and there are only three answers:
- Both name your URL. Google agrees with you. Whatever moved your rankings, it was not a copy, and every hour spent on takedowns from here is an hour wasted.
- Google-selected names another URL on your own site. This is self-inflicted duplication - a template canonical pointing everything at the home page, a faceted or parameterized URL competing with the page it was generated from, two of your own pages saying the same thing. Nobody attacked you and the fix is on your own server.
- Google-selected names a URL on a domain you do not control. This is the confirmation. Google has put your page in a cluster it does not head, and awarded the position to somebody else's copy.
The same finding surfaces in aggregate in the Page indexing report, where the statuses to read verbatim are "Duplicate, Google chose different canonical than user" - which Google glosses as your page being marked canonical while Google considers another URL the better one and has indexed that instead - and "Duplicate without user-selected canonical", which is what you see when your page declares no preference at all (Google's page indexing report documentation). Use URL Inspection for a specific page and the indexing report to find out how many pages are affected.
Two limits are worth knowing. Search Console only ever shows you properties you have verified, so the copy never appears by name anywhere except in this one field - which is why sites lose clusters for months without noticing. And the field reflects a decision, not a rule: Google's documentation states that indicating a canonical preference is a hint rather than a rule, and that Google may choose a different page for various reasons (Google Search Central, canonicalization). That cuts both ways, and both directions matter: Google can override your declaration, and it can override the copier's.
The documented cases, and what they have in common
This is the one topic on this site with genuinely documented instances, so the evidence deserves to be stated precisely rather than gestured at.
Google has acknowledged the failure mode. In February 2014 it launched a Scraper Report form for exactly this complaint, with Matt Cutts inviting people to report a scraper URL outranking the original source. The form today returns a message that it is no longer accepting responses. Google built a channel for the problem, ran it, and closed it without a replacement - which tells you the outcome is real and that there is no longer a queue to join.
The clearest case on the record is dated 18 September 2019, six days after Google announced changes to elevate original reporting. Search Engine Journal documented Yahoo News syndications outranking the original publishers, Footwear News and The Blast, in Top Stories, for identical content published at the same time (Search Engine Journal, 18 September 2019). Google's Danny Sullivan responded on the record, saying that where people deliberately choose to syndicate their content it becomes difficult to identify the originating source, and that this is why Google recommends the use of canonical or blocking. Barry Schwartz carried the same explanation at Search Engine Land the following day.
That episode is the strongest evidence available anywhere that a copy on a stronger domain can outrank an original. Note what it is not: it is not an attack. Every documented instance I have been able to find is either syndication the publisher agreed to, or opportunistic scraping that happened to sit on a much stronger domain. I found no case in which someone engineered a copy that displaced an original. That absence is the reason this page carries a situational verdict rather than a documented-threat one.
Google's response to the 2019 episode is also worth reading for its scope. On 12 September 2019, Richard Gingras, its VP of News, described ranking changes intended to keep original reporting visible for longer, and updates to the Search Quality Rater Guidelines rewarding reporting that reveals information which would not otherwise have been known, and crediting publishers with a history of original reporting. That intervention is framed entirely around news and Top Stories. There is no published equivalent for ordinary commercial pages, and it would be wrong to tell a manufacturer or a law firm that Google has a mechanism protecting their originality.
Why this is hard to weaponize
The negative-SEO theory here is that domain strength decides the cluster, and that an attacker's domain can be made stronger than yours. That theory is coherent, and it contains its own limit: an attacker who already owns a domain strong enough to beat yours has considerably more profitable things to do with it than republish your product page.
Google's published canonicalization signals reinforce that. Redirects and rel="canonical" annotations count as strong signals; inclusion in a sitemap is weak; HTTPS preference and language clustering also play a part. What is absent from that published list is instructive: publication date, first-crawled date, and domain authority are not named as canonicalization signals at all. Practitioners believe "first indexed wins" almost universally, and Google has never confirmed it.
So the realistic threat model is not sabotage. It is a strong publisher, an aggregator, a marketplace or a large user-generated-content platform holding a copy of your text with no signal pointing home - most often because you or a colleague put it there.
The conditions under which it really happens
- Syndication without a cross-domain canonical. By a wide margin the most common real cause, and it is self-inflicted rather than an attack. The partner publishes your article, declares nothing, and their domain wins.
- A copy on a substantially stronger domain - news aggregator, national publisher, marketplace, high-authority forum.
- An original that is crawled and indexed late, typically because the site is small, slow, or poorly internally linked.
- An original that is thin or templated, so it is already a near-duplicate of forty of its own siblings and has a weak claim to head anything.
- Manufacturer-supplied product descriptions used verbatim by fifty retailers. Nobody is attacking anybody; the strongest retailer takes the cluster, and the only fix is to write your own copy.
If none of those describe you, the copy you have found is very unlikely to be the reason your traffic moved.
Ruling out the things this is mistaken for
Before treating a copy as the cause, confirm the copy is actually the cause. In order of how often each turns out to be the real answer:
- Verify the copy genuinely outranks you for a query somebody types - in an unpersonalized search, for a real commercial query, not for a twelve-word exact-match phrase with no search volume. Most reported "outranking" is a copy surfacing for a snippet string nobody has ever searched.
- Read the Google-selected canonical for the affected URL. One screen eliminates most of this list in seconds.
- Compare the shape of the loss in Search Console Performance. A canonical loss looks like impressions for one URL collapsing to near zero while the rest of the site is unchanged. A core update looks broad and gradual and touches many pages. They are not easily confused once you have looked at both.
- Rule out your own accidents: two of your pages competing for one topic, a canonical error shipped in a template, a
noindexleft behind, a migration that lost its redirects.
Establish the date of the drop before you accept any explanation for it. The scraper you found today may have existed for two years.
Getting the cluster back
- Strengthen your own claim before attacking theirs. A self-referencing canonical using an absolute URL on every page, the page present in your sitemap, internal links to it from pages Google crawls often, a visible publication date and author, images you own, and structured data. These act on the signals Google says it actually uses, and unlike the rest of the list they are entirely within your control.
- If it is syndication, fix the agreement rather than the search results. Require partners to point a cross-domain
rel="canonical"at your URL, or to keep their copy out of the index. This is Google's own stated recommendation from 2019 and it is the only reliably effective remedy in this entire topic. Be careful with the word "reliably": the canonical is a hint, so partners who implement it correctly usually resolve the problem, and "usually" is the honest word. - If it is theft, remove the copy. The host removes the content; Google's legal removal process removes the URL from search results and leaves the page online. The full procedure, in order, is on the page about recovering from content scraping.
- Ask for a re-crawl of your URL once the copy is gone, then watch the Google-selected canonical rather than watching your traffic. Recovery appears there first. How long a cluster takes to re-form is not something anyone has published methodically; reported experience runs from days to months, so treat any specific promise about the timeline with suspicion, including one from me.
- Consider doing nothing. Where the copy outranks you only for strings nobody searches, the loss is theoretical, and the correct action is to note it and move on. This is the commonest right answer.
What does not help: disavowing the copying domain, since canonical selection is not a link-penalty mechanism; rewriting your page so it differs from the copy, which throws away everything the page had earned; publishing a post announcing that yours is the real version; and asking Google publicly to intervene in a specific ranking.
What people get wrong
They call it a penalty. It is a selection outcome. The distinction is not pedantic: penalties have appeals, and selection outcomes have signals. Filing a reconsideration request against a cluster loss burns the one formal channel Google gives you and cannot possibly work, because there is nothing on file to reconsider.
They send takedown notices to their own syndication partners. A copy published under a license you signed is not infringing, and a sworn notice claiming that it is exposes you to liability for knowing material misrepresentation under 17 U.S.C. §512(f) and gets your account flagged for misuse. Check the contract before the paperwork.
They file in volume. Google announced in August 2012 that sites with high numbers of valid copyright removal notices against them may appear lower in results. That signal points at the accused, not the accuser - which is precisely why mass filing against a competitor is itself a recognized attack, and why your own notices need to be accurate and documented.
They rewrite the page. It is the intuitive response and it is close to the worst one: you surrender the history, links and engagement that gave your version its claim on the cluster in the first place, and a live copy simply follows you.
They assume malice. Most of the time it is a partner deal nobody documented, a manufacturer feed, or a freelancer selling the same words twice. Those have remedies, and none of them are search remedies.
Frequently asked questions
How do I know for certain that a copy is outranking my original?
Inspect your URL in Search Console and read the Google-selected canonical field. If it names a URL on a domain you do not control, Google has chosen that copy to represent the cluster and your page is being filtered out for that content. If it names your own URL, the copy is not the reason for your ranking change, however alarming it looks in a search result.
Does Google penalize me for having duplicate content out there?
No. Google's canonicalization documentation states that some duplicate content is normal and is not a spam-policy violation. What can happen is that Google selects a different URL to represent the duplicate cluster, which suppresses your page for that text without demoting your site. There is no manual action for having been copied, and no appeal to file.
A syndication partner is outranking my own article. What do I ask them to do?
Ask for a cross-domain canonical on their copy pointing at your URL, or for their copy to be kept out of the index. That is the remedy Google itself recommended when this happened publicly in 2019. Put it in the contract for future pieces rather than negotiating it once per article, and remember the canonical is a strong hint rather than a directive, so verify the outcome in Search Console afterward instead of assuming it.
Can a competitor deliberately do this to me?
In principle, and it is much harder than it sounds. Displacing you requires publishing your content on a domain that beats yours on the signals Google uses, and anyone holding a domain that strong has better uses for it. Every documented instance I could find is syndication or opportunistic copying rather than an engineered attack. Guard against the conditions rather than the adversary: absolute self-referencing canonicals, distinctive content, and normal crawl coverage.
How long does it take to get my ranking back after the copy is removed?
Nobody has published a methodical answer, and reported experience ranges from days to months. What I would watch is not traffic but the Google-selected canonical for the affected URLs, because the cluster reassigns before the traffic returns. Anyone quoting you a firm timeline for this is quoting a number that does not exist.