What content scraping is, and the two reasons it happens
Content scraping is the automated, wholesale copying of a website's pages onto a domain somebody else controls. A script reads your URLs, frequently straight out of your own XML sitemap, fetches each page, and republishes the body text, the headings, often the images and sometimes the entire template somewhere else, usually with no attribution. It costs the copier nothing to do it to ten thousand pages rather than one, which is why it is almost never done to one.
Two motives sit behind nearly every case, and separating them is the first thing to do:
- Commercial parasitism. The copy exists to earn advertising or affiliate revenue from writing somebody else paid for. You are not the target of an attack; you are collateral in a business model. This is the overwhelming majority of what I am shown.
- Deliberate sabotage. The copying is meant to injure the original - the attacker hoping a search engine will read your page as the duplicate and the copy as the source.
Only the second is negative SEO, it is far rarer than the panic around it suggests, and the two are indistinguishable at the moment you discover them. That is precisely why so much money is spent on the wrong response: the discovery is alarming, the diagnosis gets skipped, and a cleanup is bought for a problem that was never costing anything.
Does being copied actually hurt your rankings?
Usually not, and Google's own documentation is unusually blunt about it. The canonicalization page in Google Search Central, last updated 20 August 2026, states that "Some duplicate content on a site is normal and it's not a violation of Google's spam policies" (Google Search Central, canonicalization). The fear that sends most people to this page is answered, in one sentence, by the party whose opinion actually decides the question.
What Google does with duplicates is not punish them - it chooses among them. Where several near-identical pages exist, they are grouped into a cluster and one URL is selected to represent that cluster in results. The rest are not demoted; they are simply not shown for that body of text. There is no penalty in the mechanism at all, and there is equally no rule in it that rewards whoever published first. That symmetry is the thing most defenders misread in both directions.
On the record, John Mueller of Google was quoted in May 2019 saying that other sites copying your content would not be something that negatively affects your website (Search Engine Journal, 7 May 2019); his stated reasoning was that scraper sites generally fail to rank for anything competitive in the first place. That quotation is second-hand from an office-hours session, so weigh it as a reported statement rather than published documentation - but it points in the same direction as the documentation does.
Then there is the absence, which I think is the strongest single fact here. I could not find one documented case of a site demoted because it was copied. No Google statement describes such a demotion, no manual action in Search Console corresponds to having been plagiarized, and no disclosed, reproducible experiment shows mass-copying a site's pages causing the original to lose rankings. Every claim to the contrary I have read is a vendor assertion with no method attached. For a claim this widely repeated, that evidentiary record is remarkably empty, and it should weigh heavily against anyone quoting you a price to protect you from it.
Read the direction of Google's spam policy
Google's spam policies, last updated 28 August 2026, do carry a section on scraping, and people cite it at me as proof that copying is dangerous to the victim. Read which way it points. It defines scraping as taking content from other sites, often automatically, and hosting it in order to manipulate rankings, and its examples are republishing other sites' content without adding value or citing the source, and copying content and altering it only slightly (Google's spam policies). Every clause describes the conduct of the copier. Not one creates any exposure for the site that was copied.
There is exactly one published route by which copying-related signals demote a website, and it also runs away from you rather than toward you. In August 2012 Google announced that the number of valid copyright removal notices it received about a site would become a ranking signal, and said that sites with high numbers of removal notices may appear lower in its results - the change is commonly called the Pirate update. That demotes the site accused of hosting copies.
The implication is uncomfortable and worth stating plainly: in this entire subject area, your documented ranking exposure comes from being accused of hosting copies, not from having been copied. If your site legitimately carries syndicated, licensed or manufacturer-supplied material, the useful defensive work is keeping those permissions documented - not worrying about the scraper.
Who is genuinely at risk
The verdict on this page is situational rather than mythical, and these are the situations. Every one of them is checkable in an afternoon:
- A new, thin or weakly linked original. Canonical selection leans on signals a young site has very little of. If your page has three links and the copy sits on an older domain with three thousand, the cluster can genuinely resolve the wrong way.
- An original that is crawled slowly. If a copy is fetched and indexed before your page has been seen at all, the ordering can favor it until the cluster re-forms.
- A site that is thin or templated to begin with. A page that is near-duplicate of forty other pages on your own site has a weak claim to head any cluster, with or without a scraper.
- A server that cannot take the load. Aggressive scraping is a capacity problem long before it is a ranking problem, and a site that goes down under it loses rankings for availability reasons that have nothing to do with duplication.
Notice what is not on that list: an established site with unique content, ordinary links and normal crawl coverage. If that describes you, a copy on some domain you have never heard of is background noise.
Two adjacent situations have their own pages because the mechanism differs. If the copying pipeline is fed by your RSS feed rather than by a crawler, the timing problem is much sharper and belongs under autoblogging. If a copy is outranking you right now for a query that matters, that specific failure - and the one Search Console screen that confirms it - is the subject of the page on plagiarism that outranks you.
Finding the copies
Diagnosis here is cheap, which is one more reason not to buy it as a service.
- Exact-phrase search. Take a distinctive eight to twelve word sentence from the middle of a page - never the opening paragraph, which legitimate quotation reproduces - put it in quotation marks and exclude your own domain from the results. Repeat across three or four pages in different sections. One hit is noise; the same pattern across sections means a crawler walked the whole site.
- Duplicate-finding services run that phrase matching across the index on a schedule and keep the history. They do nothing you could not do by hand; what you are buying is coverage and a timestamped record, and the record is the part that matters later, when you have to show a host that a copy existed on a given date.
- Server logs. A scrape has a shape: one IP address or a narrow range, requesting your URLs in sitemap order at even intervals, with an unusual or absent user-agent, fetching HTML and never touching your stylesheets, images or scripts. Line up the timing of requests for your sitemap against the sequence of page requests that follows it.
- Image referrer logs catch hotlinking, where the copy displays your images from your server. This is often the first hard evidence a copy exists, and it usually arrives as a bandwidth bill.
- Alerts on your brand name or a phrase you always use will catch new copies as they get indexed.
What it gets mistaken for. A ranking drop discovered in the same week as a scraper is very rarely caused by it. The usual real causes are a core update, duplication the site created for itself through faceted URLs or session parameters, a botched migration, a noindex left behind after a staging push, or ordinary competitive movement. Establish the exact date the traffic changed, then ask what changed on that date. Doing it in the other order is how people lose a month.
What scraping really costs you
An honest accounting, because the ranking loss everyone fears is usually the smallest item on it:
- Bandwidth, capacity and crawl budget. Scrapers are indiscriminate and repetitive. On a metered host that is an invoice; on a fragile one it is downtime; and a server busy answering bots is slower answering Googlebot.
- Image bandwidth. Hotlinked images are served by you, to somebody else's visitors, at your expense, indefinitely.
- Licensing value and brand dilution. If your text is worth paying for, a free copy is a lost sale whether or not either version ranks - and your writing appearing beside content you would never publish is a cost no ranking report will show you.
- Your time. In most cases this is the largest cost by a wide margin, and it is entirely elective.
- Frequently, nothing at all. A copy on a domain with no traffic, ranking for nothing and hotlinking nothing, has cost you zero and deserves exactly that much attention.
What to do, in order
- Establish whether anything was lost before responding to anything. Compare impressions for the affected pages across the date of the drop. If your pages still rank and still draw impressions, the copy is a nuisance and you can stop at step three.
- Put your own house in order first. Every page self-canonicalizing to a fully qualified URL, an accurate sitemap, no unintentional duplication of your own making. Fixing your own duplicate-URL sprawl improves canonical selection more than anything you can aim at a scraper.
- Get crawled faster. A current sitemap, a fast server response, and internal links to new pages from the pages Googlebot already visits often. Being seen first is the only defense in this category that works on the mechanism itself rather than on the symptom.
- Only then act on the copy - and only where it is outranking you for a query you care about, or costing you bandwidth. Everything else is optional.
- Take it down through the host, then the registrar, then Google, remembering that a search removal takes the URL out of results and leaves the page on the web. I have set out that procedure, the order to run it in and its risks on the page on recovering from content scraping.
- Block at the edge for load, not for rankings. Rate limits, challenges for abusive clients, and a referrer rule to stop image hotlinking.
What does not help, at all: disavowing the copying domain, because disavow addresses link signals and copied text is not a link; rewriting content that was ranking perfectly well on the theory that it must be differentiated from the copy; filing a reconsideration request when Search Console shows no manual action to reconsider; adding noindex to your own pages to make them less worth stealing; and blocking crawlers with broad rules that catch Googlebot along with the scraper. That last one is the most common self-inflicted injury in this whole category, and unlike the scraping it will definitely cost you traffic.
What people get wrong about this
They look for a reporting channel that no longer exists. In February 2014 Google opened a Scraper Report form, and Matt Cutts invited people to use it when a scraper URL outranked the original. That form now returns the message that it is no longer accepting responses. The dedicated channel for this exact complaint was retired and never replaced. Read that in both directions: it takes a tool away from defenders, and it suggests Google stopped treating the problem as one needing human triage.
They buy protection from a penalty that does not exist. There is no duplicate content penalty; Google's documentation says so in terms. Anyone selling you monitoring or "duplicate content protection" on the strength of a penalty the search engine publicly denies having is selling you something you do not need, and the fact that copies genuinely exist does not make the product less hollow.
They conflate discovery with causation. Finding a scraper feels like finding the culprit. It is usually finding a coincidence, and the weeks spent on takedowns are weeks not spent on whatever actually moved the rankings.
They file copyright notices carelessly. Sworn statements made under penalty of perjury against material that turns out to be quoted, licensed or syndicated with permission create real liability and get accounts flagged for misuse. Accuracy is not a formality.
And they ignore the costs that are real. Nobody calls me about the image bandwidth, and that is frequently the only line item the scraping was ever going to produce.
Frequently asked questions
Someone copied my whole site. Will Google penalize me for duplicate content?
No. Google's canonicalization documentation states that some duplicate content is normal and is not a violation of its spam policies. There is no duplicate content penalty and no manual action corresponding to having been copied. What can happen - rarely, and mostly to new or thin sites - is that Google picks the copy rather than your page to represent the duplicate cluster in results. That is a selection loss, not a punishment, and it is confirmable rather than something to guess at.
How do I prove I published it first?
Keep the evidence before you need it: your content management system's revision history, the page's dated appearance in the Internet Archive, your own server logs, and the sitemap entry with its original last-modified date. Search engines do not adjudicate authorship on that evidence - a hosting provider, a registrar or a court might. Capture it at the moment you find the copy, because copies get edited and deleted once you complain.
Should I disavow the domains hosting copies of my content?
No. The disavow file tells Google to ignore links pointing at your site. Copied text is not a link, so there is nothing for the tool to act on, and a hastily assembled disavow file will usually strip out legitimate links along with the intended targets. This is one of the most reliably counterproductive things a worried site owner can do.
Should I switch to a summary feed or block bots to stop the copying?
Both are trades, not wins. Blocking has to be precise enough to miss Googlebot, Bingbot and real readers, and a rule broad enough to stop determined scrapers is usually broad enough to cost you indexing. Feed changes are worth considering only when the copying is feed-driven, which is a different pattern with its own page. Fix your own canonical signals first; they cost you nothing and they work on the mechanism that decides the outcome.
When is it actually worth paying someone to deal with this?
When a copy is outranking you for a commercially valuable query, when copies are monetizing your work at scale, when hotlinked images are generating real bills, or when you need the evidence preserved properly for counsel. Outside those cases, the honest answer is that the copies exist, they are not costing you anything measurable, and the correct response is to note them and get on with your work.