What each route actually achieves
Content scraping is the automated copying of a site's text, images or feed onto another domain, usually at scale and usually without attribution. The mechanism, the prevalence and the contested question of whether a copy can outrank the original belong to the content scraping profile. This page is the recovery procedure: what to do, in what order, once you have found the copy.
The order matters more than any individual step, because the first two questions decide whether the remaining steps are worth taking, and most site owners start at step three. Before any of it, be clear about what each route delivers, because they get conflated constantly:
- Fixing your own canonicalization makes your URL the one Google treats as the original. It is the only route that affects rankings directly.
- A copyright removal request to Google removes the copy's URLs from Google Search. The page stays on the web.
- A notice to the host removes the content itself. The URL may then return a 404 everywhere.
- A notice to the registrar rarely does anything about copyright; registrars act on fraud and on false registration data. Slow, and often no action at all.
- Doing nothing costs nothing where the copy costs nothing, which is frequently the case.
Google's stated position is that it selects a canonical among duplicates rather than penalizing duplication, and the practical implication for a recovery procedure is blunt: a copy on an unknown domain, ranking for nothing and drawing no traffic, is not a search problem and does not merit a legal instrument. Scraping is real, ubiquitous and occasionally damaging, but the damage is concentrated in a small subset of cases, and firing notices at every mirror is disproportionate, slow, and carries liability the sender rarely understands.
Step 1 - establish which URL Google treats as canonical
Before anything else, find out whether Google has actually picked the copy over you. This is checkable in about four minutes, and until you have checked it you are speculating.
- Search Console, URL Inspection, on the affected URL of your own site. Read the Google-selected canonical against the user-declared canonical. If Google's selection is your URL, the copy is not displacing you at the index level and the problem is smaller than it feels.
- A
site:search on your own domain for a distinctive sentence, to confirm your page is indexed at all. An unindexed original is an indexing problem — a page Google has crawled but not stored, or never crawled — and it has a completely different fix from a theft problem. Diagnosing one as the other wastes weeks. - The same distinctive sentence in quotes, without
site:, logged out, to see the order results actually appear in rather than the order your personalized results suggest. - The copy's page source, checked for a
rel="canonical"pointing at your URL. Many scrapers copy the tag wholesale, which means the scraper is actively telling Google your page is the original. That is a good outcome and requires nothing from you. - Whether the copy is indexed at all. A scraped page Google has not indexed is invisible and inert.
Then put your own house in order, because this is the part you control. Every page self-canonicalizing; one hostname serving 200s rather than both the apex and the www variant; a current XML sitemap with accurate lastmod values; fast server responses; internal links from pages that are crawled often. Being crawled before the copy is the most effective defense a victim actually has, and fixing your own duplicate-URL sprawl does more for canonical selection than anything aimed at the scraper.
Step 2 - decide whether the copy is costing you anything
Ask these in order and answer them honestly, which mostly means answering them with data rather than with how it feels to see your own sentences on someone else's domain.
- Is it outranking you for a query you care about? Not any query — one that sends you customers.
- Is it taking traffic? Referral logs, brand search volume, and Search Console impressions for the affected URLs before and after the copy appeared.
- Is it monetizing your work through ads, affiliate links, a paywall or a lead form?
- Is it misrepresenting you — your brand name attached to content you did not write, or your content attached to something defamatory or unsafe? That is a different and more serious problem than duplication, and it belongs with SERP defamation rather than here.
- Is it costing bandwidth, through hotlinked images or an aggressive crawler? That is an infrastructure problem with an infrastructure fix — rate limits and referrer rules — and needs no legal instrument at all.
If every answer is no, stop. A copy on a domain nobody visits costs you nothing and deserves nothing. The hours spent filing notices against it are hours not spent on the site, and for the ordinary autoblog mirror described on the RSS autoblog theft page, doing nothing is the correct answer rather than the lazy one.
Step 3 - the DMCA route through Google, and what it really does
A copyright removal request to Google removes the specified URLs from Google Search results. It does not remove the page from the web, does not affect other search engines, and does not touch the copy's own direct traffic. Knowing that before you file prevents the most common disappointment.
Where it is filed. Google's legal removal troubleshooter is the only correct channel. There is no email address, and a notice sent anywhere else at Google is not a notice.
What a valid notice must contain — the statutory elements at 17 U.S.C. 512(c)(3)(A):
- a physical or electronic signature of the copyright owner or an authorized agent;
- identification of the copyrighted work claimed to be infringed;
- identification of the material claimed to be infringing, with enough information to locate it — which in practice means the specific URLs, not a domain;
- contact information for the complaining party;
- a statement of good faith belief that use of the material in the manner complained of is not authorized by the owner, its agent, or the law;
- a statement, under penalty of perjury, that the information is accurate and that the sender is authorized to act on the owner's behalf.
Timeline. Google publishes one figure: "Our average processing time across all removal requests submitted via our web form for Search is approximately 6 hours." Read it for what it is. That is an average across an enormous volume of automated filings by rightsholder agencies, and it includes refusals. It is not a promise about your individual manually reviewed request, and quoting it without that qualification is how a client ends up angry on day two.
Refusals. Google states that it declines requests for reasons including not having enough information about why the URL is allegedly infringing, not finding the allegedly infringing content referenced in the request, and detecting that the copyright removal process is being used improperly. Vague notices and domain-level notices get refused. Name the URLs.
Publication. Google forwards notices to the Lumen database, which publishes them with personal information removed, and Google may show a link to the Lumen entry in place of the removed results. Your notice becomes a public document. Assume it will be read by the other side and by anyone searching your brand, and write it accordingly. Google may also notify the site operator, who then has the counter-notice route open to them.
Step 4 - the host, the ad network, and the registrar
A search delisting leaves the page live. Removing the content itself means going to whoever hosts it, and the order of effectiveness is not the order most people try.
- Identify the host, not the CDN. A reverse-IP or WHOIS lookup frequently returns a proxy or content delivery network rather than the origin. The CDN's own abuse process will usually forward to, or identify, the host behind it. Cloud providers, shared hosts and CDNs all publish designated agent contacts, and a US service provider that wants the section 512 safe harbor has an agent registered with the Copyright Office.
- The notice to the host is the same statutory instrument — same elements, same perjury statement, same misrepresentation exposure — but the outcome differs: removal of the material rather than delisting of it. Google publishes no timeline for a host's initial response, and neither does anyone else.
- The ad network is frequently the fastest lever, and it is commercial rather than legal. An autoblog that exists to monetize stolen content has no reason to keep running once its ad revenue stops, and programmatic platforms publish policy-violation reporting channels. This route gets skipped because it feels less serious than a legal notice, and it often works better.
- The registrar is a third-tier route. Registrars generally will not act on a copyright complaint, because a domain registration is not the infringing material. Where one does act it is usually on false or unreachable registration contact data, under its accreditation obligations. Treat it as an escalation for an abandoned, anonymous domain with no reachable host, not as a primary channel.
Step 5 - the risks nobody reads before filing
The counter notification. The person you filed against can file a counter notice under 17 U.S.C. 512(g)(3), stating under penalty of perjury a good faith belief that the material was removed by mistake or misidentification and consenting to the jurisdiction of a federal district court. If they do, the host must restore the material "not less than 10, nor more than 14, business days following receipt of the counter notice" unless the original complainant files suit. Google's own statement is that when it receives a counter notification it may reinstate the material in question.
The practical consequence is worth sitting with: a takedown you cannot or will not follow into court is reversible, inside that statutory window, by anyone willing to sign a form. Filing against a determined operator without being prepared to litigate buys a two-week removal and a better-informed opponent.
Section 512(f) misrepresentation liability. This is the exposure that makes an overreaching notice dangerous rather than merely ineffective. The statute makes any person who knowingly materially misrepresents that material is infringing — or that it was removed by mistake — liable for any damages, including costs and attorneys' fees, incurred by the alleged infringer and by the service provider. Google warns about it on its own legal removals process page, telling filers they will be liable for damages including costs and attorneys' fees if they materially misrepresent that a product or activity is infringing, and suggesting that anyone unsure whether material online infringes their copyright should first contact an attorney.
Two cases define the practical shape of that exposure. In Online Policy Group v. Diebold, Inc., decided in the Northern District of California on 30 September 2004, the court found a section 512(f) violation, holding that no reasonable copyright holder could have believed that the portions of the email archive discussing possible technical problems with Diebold's voting machines were protected by copyright; Diebold agreed to pay $125,000 in damages and attorneys' fees. In Lenz v. Universal Music Corp., decided by the Ninth Circuit on 14 September 2015, the court held that a copyright owner must consider whether a use is a fair use before sending a takedown notice, reasoning that fair use is not an infringement to be excused but is not infringement at all — while adopting a subjective standard for liability, under which a sender who genuinely, even unreasonably, believed the material was infringing can escape it.
Read together: the threshold for liability is demanding but not theoretical, and the obligation to consider fair use before filing is settled law in the Ninth Circuit. Other circuits have not uniformly adopted the same reasoning, so do not treat it as a single national rule. In a negative SEO context the exposure is sharpest in the mirror image of this page — a fraudulent notice used as a weapon against a legitimate site, which is its own documented attack and is covered under fake DMCA takedowns. If you would be outraged to receive the notice you are about to send, do not send it.
And do not file on content you do not own. The perjury statement is about authorization, and quotation, commentary, a licensed stock image, a syndication agreement you forgot about, or material owned by a client rather than by you are all routes to a defective notice.
The whole procedure, in order
- Establish Google's selected canonical for your own URLs, and fix your own canonicalization, sitemap and crawl speed first.
- Decide whether the copy costs anything. If it does not, stop.
- If it does: a notice to the host removes the content, a request to Google removes the search results, and both name specific URLs.
- Escalate to the ad network if the copy monetizes; to the registrar only for an abandoned, unreachable domain.
- Rate-limit or challenge abusive crawlers for load reasons only, never with rules broad enough to catch Googlebot.
- Keep a dated record of every notice, response and reappearance. Scrapers return, and a documented history is what makes a pattern legible later.
What does not help: disavowing links from the scraper, which is the wrong instrument entirely; rewriting your own content to differentiate it from the copy; adding noindex to stop the scraping, which removes you from search and leaves the copy in place; filing a reconsideration request, because there is no manual action here; and broad robots.txt or firewall rules that also exclude Googlebot, which is the most common self-inflicted wound in this category.
One jurisdictional limit, stated plainly rather than implied: everything above describes United States procedure. Outside the US the DMCA does not apply and the equivalent notice-and-action regimes differ by country. Do not assert US procedure against a host that is not subject to it, and do not assume a template written for one regime travels.
Frequently asked questions
A site copied my content and I think it is outranking me. What do I check first?
URL Inspection in Search Console, on your own URL, and read the Google-selected canonical. Most reports of a scraper outranking the original are not borne out by that check. Then search a distinctive sentence in quotes while logged out, and see the actual order. If Google has selected your URL, the copy is not displacing you at the index level and a takedown notice is solving a problem you do not have.
How long does a Google copyright removal take?
Google publishes an average processing time of approximately 6 hours across all removal requests submitted through its web form for Search. That average is dominated by high-volume automated filings from rightsholder agencies and includes refusals, so it is not a representative figure for one manually reviewed request from a site owner. Treat it as context, not as a commitment.
Should I send the notice to Google or to the host?
They do different things. Google removes the URLs from search results and leaves the page live; the host removes the content itself. If the copy is a search problem, Google is the right recipient. If it is a copying problem regardless of search, the host is. Sending both is reasonable when the copy is doing real damage, and neither is worth sending when it is not.
Can I get in trouble for filing a DMCA notice?
Yes, under 17 U.S.C. 512(f), which makes a knowing material misrepresentation that material is infringing actionable for damages, costs and attorneys' fees. Google warns about that exposure on its own process page. The Ninth Circuit held in Lenz in 2015 that a copyright owner must consider fair use before sending a notice, though it applied a subjective standard for liability, and other circuits have not uniformly followed the same reasoning.
The scraper filed a counter notice and the content came back. What now?
That is the statute working as written: after a counter notification the host restores the material not less than 10 nor more than 14 business days later unless you file suit. There is no administrative appeal above it. Your remaining options are litigation, pressure on the ad network that funds the copy, or accepting the copy and competing — and for most sites the third is the rational one, which is why step two of this procedure comes before step three.
Should I block the scraper in robots.txt?
Only with rules narrow enough to be certain of what they catch. Scrapers routinely ignore robots.txt anyway, and the broad rules people write in frustration take Googlebot out along with the scraper. That converts a copying problem into a deindexing problem, which is a far worse day. Rate limits at the server, aimed at load rather than at copying, are the safer instrument.