NegativeSEO.ICU logo — negative SEO reference and recoveryNegativeSEO.ICUNegative SEO reference & recovery
Legal & reporting

Preserving the evidence, before it decays out from under you

Whether the evidence still exists in six months is a decision made in the first week, usually by someone who does not know they are making it.

Preservation has two parts, and conflating them serves neither

A search attack leaves evidence in perhaps a dozen places. Almost all of it is perishable, most of it is held by third parties who owe you nothing, and the parts that persist are the parts most easily challenged as unauthenticated, altered or cherry-picked.

Preservation has two distinct parts, and work aimed at one rarely satisfies the other. Reconstruction is being able, a year later, to establish what happened, when it started and in what order - what an expert witness needs in order to have an opinion at all. Admissibility is being able to put the material in front of a fact-finder over an opponent's objection, which is what the Federal Rules of Evidence govern. A screenshot pasted into a word processor usually delivers neither.

Running through both is chain of custody: a contemporaneous record of every person who handled each item, when, what they did to it, and where it was stored in between. It is what converts "here is a file" into "here is a file whose history is accounted for," and its absence is the first thing a competent opponent attacks - an attack that never requires proving anything was altered, only establishing that nobody can say it was not.

This page is general information about how electronic evidence is preserved and authenticated in United States federal practice. It is not legal advice, and it cannot tell any reader what their own obligations are. I am not a lawyer; this site does not practice law, and this page is not a preservation protocol for any particular matter. Preservation duties, spoliation consequences, admissibility and privilege are decided by a court applying the rules of its own jurisdiction to specific facts, and several states have not adopted analogues of the federal self-authentication rules below. Anyone anticipating litigation needs counsel, and where the evidence will be contested, a qualified examiner engaged early enough to collect it properly. I work as an expert witness in matters of this kind, which is why I will say plainly that the material handed over is usually thinner than the client believes.

One rule governs the rest: preserve first, then remediate. The most damaging habit in this subject is the owner who, on discovering an attack, cleans the site, deletes the injected pages, rotates credentials and disavows a thousand domains - destroying the record before anyone has looked at it. Where speed and preservation genuinely conflict, an image taken before cleanup resolves it in minutes.

The decay clock: what disappears, and how fast

This is the part a defender needs first, and the unit is days rather than weeks. Every figure below is an order-of-magnitude prompt to go and ask your own providers, not a sourced statistic - retention defaults vary by host, stack and plan.

  • Web server access logs - days to weeks by default, held by you or your host, rotated on a schedule nobody chose deliberately. Rarely recoverable once rotated.
  • CDN, WAF and reverse-proxy logs - hours to days on standard plans, held by a third party, and frequently not retained at all unless you configured streaming to your own storage.
  • The attacker's own email - deletable at their option, instantly, and never recoverable from your side.
  • A spam or grievance page - alive until the attacker removes it, after which only an archive holds it.
  • Third-party link index rows - overwritten as the index refreshes, with lost links dropping out entirely. Recoverable only from a dated export taken at the time.
  • Search Console performance data - a rolling window held by Google. Once it rolls off it is gone, and there is no appeal.
  • The Search Console links report - no history at all. It is a snapshot of now.
  • Search result positions and appearance - continuously changing, personalized, held by nobody.
  • Registration and DNS records - changed at the registrant's option, DNS in minutes; recoverable only from commercial history or passive DNS services.
  • Review text and reviewer profiles - until removed or edited by the platform or the author.
  • The text of a copyright notice - the durable exception. Notices Google shares with the Lumen Database persist and are usually retrievable later.

The practical consequence. By the time a business is angry enough to call a lawyer - typically four to eight weeks after the drop - the logs are gone, the attacker has deleted the pages, and the only surviving record is a screenshot somebody took on a phone. That is the ordinary case, and a large part of why so few matters are brought.

Logs: the paradigm case for certification

Server and CDN logs contain what nothing else does: requests, timestamps, source addresses, user agents, referrers, status codes and response sizes. They are the only source that can show a flooding pattern, a spoofed crawler pattern, the moment an injection request succeeded, or the authenticated session that made a change.

They vanish because rotation - the automatic recycling of log files on a fixed schedule - is a default rather than a decision, and nobody discovers their retention window until the day they ask for it. A written request to the host and the CDN to extend retention and preserve what they currently hold is therefore an early step, and the request is itself evidence of when preservation began.

Evidentially, a log file is a record of an automated process and the paradigm case - the textbook example a rule was written around - for Federal Rule of Evidence 902(13), "Certified Records Generated by an Electronic Process or System," which self-authenticates "a record generated by an electronic process or system that produces an accurate result, as shown by a certification of a qualified person." The same material can be reached through Rule 901(b)(9), "Evidence About a Process or System," which allows authentication by "evidence describing a process or system and showing that it produces an accurate result." The certification route is what avoids calling a systems administrator to the stand - but it needs a qualified person who can describe how the logging system works, and it carries the same advance written notice to the adverse parties that Rule 902(11) requires for certified business records.

A preservation step, concretely: copy the logs off the machine to controlled storage, hash each file at the moment of copying, record who copied them, when, from where and with what tool, and keep the written request to the provider alongside them.

Search Console, and the export that escapes the rolling window

Search Console holds Google's own view of the site, which no other source can supply: performance by query, page, country and device; the external links report; the Manual actions report; the Security issues report; index coverage; and the message center. Its limits are severe and scattered across a dozen help pages:

  • It is a rolling window. Data older than the window is gone and cannot be recovered from Google on any terms.
  • There is a reporting lag. Google states that collected data is usually available in two to three days, and an expert dating the start of an attack has to account for it.
  • Rare queries are omitted from the table but included in chart totals, so totals and row sums disagree by design. An opponent will make that discrepancy sound like manipulation unless it is explained before they raise it.
  • The links report has no history whatsoever. There is no "as of" date and no way to ask what it said last March. A links report not exported at the time is simply lost.
  • The Manual actions report is a live state, not a log. Once an action is lifted, the historical record survives only in the message center and in whatever the owner captured.

The one durable fix, and the most useful operational fact on this page: Search Console's bulk data export to BigQuery escapes the rolling window entirely. Google's own documentation says that data "will be accumulated forever for your project, unless you set an expiration time for your data". Configured before an incident, it converts the single most important evidentiary source in this subject from perishable to permanent, at the cost of an afternoon and a storage bill. Configured after, it does nothing retrospectively - which is the entire point, and the reason it belongs in a prevention conversation rather than an emergency one.

For the reports that cannot be exported at all - Manual actions, Security issues, the message center - the record is a dated full-page capture with the account and system clock visible, hashed at capture. And every export should carry its export date in the filename and in a log, because a spreadsheet with no date is not evidence of anything.

Link indexes are measuring instruments, not self-proving records

Commercial link indexes are instruments here, not remedies. They see new referring domains before Search Console does, and they are the only practical source for the shape of the referring-domain growth curve, which is what distinguishes a campaign from ordinary acquisition - a vertical step from a standing start looks nothing like the slow drift of a scraped web.

They must be captured with a date, and this is not negotiable. These are commercial crawls of a moving web: rows drop out when a link disappears, when a host goes offline, or when the vendor re-crawls and re-scores. A report run in March and a report run in September are different data about a different web, and no vendor holds an immutable historical record you can subpoena in a form a court will accept. If the export was not taken at the time, the state of the link profile at the time cannot be reconstructed - the most common irrecoverable loss in these matters, and entirely avoidable. A defensible capture is a raw export rather than a screenshot of a chart, filed with the tool name, its own index or refresh date where published, the export timestamp, the exact filter settings and the account that ran it, hashed, and repeated on a schedule so the curve exists as a series rather than two lonely points.

Anyone assuming a spreadsheet downloaded from a subscription tool is self-proving is going to be surprised. It is a third party's record. Admitting it usually means either a Rule 902(11) business-records certification from the vendor, which vendors are not set up to give routinely, or an expert testifying under Rule 901(b)(9) to how the index is produced and what it does and does not represent. Expect the numbers to be attacked as an estimate drawn from a crawl rather than a census of the web, because that is what they are.

Archives, registration data, and the disclosure route most people miss

Web archives answer the two questions an attacker's delete button otherwise closes: what did that page say, and who published first. The Wayback Machine has a long track record in litigation and stamps a capture time into the URL itself; archive.today captures on demand, which is its virtue, since a defender can force a capture of a live page now rather than hoping a crawler visited. Two cautions travel with both: a capture that does not exist cannot be created retroactively, and a noted archive URL is not a preserved copy - download the rendered page and its assets and hash them.

Registration data changed in a way most search practitioners have not registered. Public WHOIS is largely gone. Under ICANN's Registration Data Policy, effective 21 August 2025, registries and registrars publish a limited set - domain name, registrar details, creation date, expiry date, domain status, registrant country - while redacting the registrant's name, street, postal code, phone and technical contact data.

The useful half is that the same policy creates a formal disclosure route with clocks on it. Registrars and registries must publish a direct link on their homepage to the process for submitting a disclosure request, acknowledge a request within two business days, and respond without undue delay and no more than thirty calendar days from acknowledgement, absent exceptional circumstances. For urgent requests involving imminent threats the acknowledgement window is two hours and the response window twenty-four. That is a documented, time-bound channel most victims do not know exists, and counsel should hear about it early - thirty days is a long time if the request goes in during month three.

It is worth the effort because a purpose-built grievance domain's creation date establishes when a campaign began even when the registrant is masked, nameserver and address-record changes date a redirect hijack, and shared infrastructure across attacker-controlled domains is one of the few practical routes to attribution at all.

Captures, hashes, and what makes them hold up

An ordinary screenshot proves nothing about when it was taken. Clocks are trivially changed, metadata is trivially edited, and an image pasted into a document has lost whatever metadata it had. Under Rule 901(a) the proponent "must produce evidence sufficient to support a finding that the item is what the proponent claims it is" - and for a bare screenshot that means a witness swearing they took it, which invites the obvious cross-examination.

What strengthens a capture, in ascending order: the full browser window showing the URL bar, the system clock and the signed-in account, not a cropped rectangle; the page source or a full archive capture saved alongside it, so the response headers and served markup survive rather than a picture of the rendering; a hash of every file recorded at capture time in a log; a third party in the chain, such as a forced public-archive capture at the same moment; and, for anything expected to be contested, capture performed or supervised by a qualified examiner who is not the interested party. A capture produced by a documented, repeatable system is what Rule 902(13) was written to accommodate - the value of a system over a person with a phone is that a system can be certified.

Hashing is now a rule rather than a nicety. Rule 902(14), "Certified Data Copied from an Electronic Device, Storage Medium, or File," self-authenticates data copied from a device, medium or file "if authenticated by a process of digital identification, as shown by a certification of a qualified person" meeting Rule 902(11)'s certification and notice requirements. Rules 902(13) and 902(14) both took effect on 1 December 2017, so that electronic records could be authenticated by certification instead of by dragging a witness into a courtroom.

Three things are routinely got wrong. Hash at the moment of collection, because a hash taken after a file has been opened, re-saved or passed through a mail client proves the wrong thing. Record the hash somewhere other than beside the file, since a hash sitting in the same folder as the evidence, both writable by the same person, proves less than people assume. And the certification still needs a qualified person: hashing is necessary and not sufficient.

The reference standard for handling all of it is NIST Special Publication 800-86, Guide to Integrating Forensic Techniques into Incident Response (September 2006) - old, and still the routine citation. A workable chain of custody record carries the item description, its source, who collected it, the date, time and time zone, the tool and version, the hash, every transfer with date and recipient, and where it was stored. Kept contemporaneously: one reconstructed from memory is worse than none, because it invites a question about everything else in the file.

The legal hold: the duty starts before anyone files

The trigger is anticipation, not filing. Federal Rule of Civil Procedure 37(e) (law.cornell.edu/rules/frcp/rule_37), added in the 2015 amendments, applies to electronically stored information "that should have been preserved in the anticipation or conduct of litigation" and that "is lost because a party failed to take reasonable steps to preserve it, and it cannot be restored or replaced through additional discovery."

What a court can then do splits in two. Under 37(e)(1), on finding prejudice, it "may order measures no greater than necessary to cure the prejudice." Under 37(e)(2), and only on finding "that the party acted with the intent to deprive another party of the information's use in the litigation," it may presume the lost information was unfavorable, instruct a jury that it may or must so presume, dismiss the action, or enter a default judgment.

Three things follow for a business owner. The duty attaches to you, the potential plaintiff, as much as to the attacker - a victim who sues and cannot explain what happened to their own logs is the one answering for it. "Reasonable steps" in practice means suspending routine deletion for the sources that matter - log rotation, mailbox auto-purge, backup expiry, revision pruning - and documenting that you did, and when. And because the severe sanctions require intent, the ordinary risk is not a default judgment but a curative measure, plus the far larger informal cost of a case that cannot be proved.

Scope is the counterweight, and a hold that tries to preserve everything is its own failure mode. Rule 26(b)(1) limits discovery to matter "relevant to any party's claim or defense and proportional to the needs of the case"; Rule 26(b)(2)(B) recognizes information "not reasonably accessible because of undue burden or cost"; Rule 26(f)(3)(C) requires the discovery plan to address preservation and the form of production. What gets preserved, and how broadly, is a decision for counsel. A hold also does not reach third parties - you cannot impose one on Google, a link vendor, a registrar or a CDN. You can ask in writing and keep the request, export what they will give you now, and have counsel consider a preservation letter or subpoena early rather than late, because the window on third-party data is shorter than the window on your own. One omission belongs on the page rather than left silent: nothing here addresses privilege, work product, or the discoverability of an investigation's own findings, all of which are counsel's decisions and all of which can be prejudiced by how day one is structured.

What an expert needs to reconstruct this a year later

This is the acid test for a preservation effort. To reconstruct an attack from a standing start twelve months on, an examiner needs eight things, and the first is the one nobody keeps.

  1. A baseline - what traffic, rankings, index coverage and the link profile looked like before, from a source not created after the fact. Without a baseline there is no change, and without change there is no causation.
  2. A dated series rather than two snapshots. Two points describe a line; a series describes an event with a start.
  3. Raw artifacts rather than summaries. An analysis cannot be re-run on a screenshot, and an opinion that cannot be re-run is an opinion an opposing expert will dismantle.
  4. The change history of your own site - deploys, revisions, plugin and theme updates, DNS and hosting changes, migrations, and the work of any prior vendor. The first thing a competent opponent looks for is a self-inflicted explanation, and in a large share of suspected-attack matters they find one.
  5. A dated external timeline of core updates, spam updates and platform policy changes, because attributing a drop to an attack means excluding the update that landed the same week.
  6. The attacker-side artifacts - archived copies of the pages, the text of any takedown notice, review text and reviewer profiles, registration and DNS history, the emails with headers.
  7. The provenance of every item - who collected it, when, how, with what tool, and its hash. Material of unknown origin cannot be opined on, and a report resting on it gets excluded or discounted.
  8. A contemporaneous incident log, written as things happened, in the owner's own words. It is frequently the most persuasive single document in the file and it costs nothing to keep.

The honest caveat, and it belongs here rather than in a footnote. Even with all eight, an examiner can usually establish what happened and when far more confidently than who did it or that it caused the loss. Preservation makes a case possible; it does not make it strong, and no volume of careful collection turns an anonymous campaign into an identified defendant. Anyone told otherwise is being sold something. What preservation does buy is the one variable in this whole subject that is genuinely inside your control on the day you first suspect something - which is why the legal theories guide ends where this page begins.

Frequently asked questions

How long do server logs last?

Days to weeks by default, and on many content delivery networks hours - or not at all, unless streaming to your own storage was configured in advance. Nobody discovers their real retention window until they ask for it, which is usually the week after the window closed. Asking the host and the CDN in writing to extend retention is an early step, and the written request doubles as evidence of when preservation began.

Can I get old Search Console data back?

No. Performance data older than the rolling window cannot be recovered from Google on any terms, and the links report has no history at all - it shows the current state, with no way to ask what it said last March. The only durable fix is the bulk data export to BigQuery, which Google documents as accumulating data for your project indefinitely unless you set an expiration. It has to be configured before the incident; it does nothing retrospectively.

Is a screenshot enough evidence?

Usually not on its own. Under Rule 901(a) the proponent has to produce evidence sufficient to support a finding that the item is what it is claimed to be, and a cropped image proves nothing about when it was taken. Capture the full window with the URL bar and clock visible, save the page source or a full archive capture alongside it, hash everything at the moment of capture, and where possible force a public archive capture at the same time so the date does not rest solely on your own machine.

Do I need a lawyer to start preserving?

Preservation and the decision to litigate are different decisions, and the first does not wait on the second - the evidence decays on its own schedule. But scope, privilege and what a hold covers are counsel's calls, and the duty to preserve attaches when litigation is reasonably anticipated rather than when it is filed. Suspending routine deletion and capturing what exists is the part that cannot sensibly be postponed.

Should I clean up the site first or preserve first?

Preserve first. Cleaning a compromised site, deleting injected URLs, rewriting the targeted page and disavowing on suspicion all destroy the record before anyone has examined it, and the intrusion is the part with the most law behind it. Where waiting is genuinely not an option, an image or snapshot taken before cleanup resolves the conflict in minutes and preserves everything.

Does any of this apply outside federal court?

The framework described here is federal. State rules differ, sometimes materially on spoliation and on self-authentication, and several states have not adopted analogues of Rules 902(13) and 902(14). The underlying practice - date everything, hash at collection, keep the raw files, record who touched what - travels well regardless, but which rule governs your matter is a question for a lawyer in your jurisdiction.

Top