Monday, October 5, 2026
Home » How Do Archives Run Fixity Checks at Petabyte Scale?

How Do Archives Run Fixity Checks at Petabyte Scale?

Fixity checks are how digital archives prove that content has not changed. A checksum computed when a file enters the archive is compared with a checksum computed later; if they match, the file is intact. For a few terabytes, that is simple. At petabyte scale, with hundreds of millions or billions of files across several copies, fixity checking becomes one of the largest workloads an archive runs. Reading every byte takes time, consumes bandwidth and competes with ingest and access, and results must be recorded and acted on.

This article explains how archives run fixity checks at scale: the throughput math, scheduling strategies, the role of storage-level integrity checking, sampling and how to handle failures. For the wider preservation context, see our hub on digital preservation storage.

What fixity checking involves

A fixity program typically includes:

  • Checksum at ingest: compute a checksum for every file using an algorithm such as SHA-256 or MD5, and store it in preservation metadata.
  • Periodic verification: re-read each file, compute a new checksum and compare it with the stored value.
  • Verification on movement: check fixity after transfers, migrations and media refresh.
  • Recording results: log every check, including date, result and copy location, as evidence for audits.
  • Repair: when a check fails, replace the damaged copy with a verified good one.

The verification step is where scale becomes a challenge.

The throughput math

Fixity checking at scale is a reading problem. To verify a file, you must read all of it. Consider an archive holding 10 PB per copy:

  • At an aggregate verification rate of 1 GB per second, reading 10 PB takes about 10 million seconds, or roughly 116 days.
  • At 5 GB per second, it takes about 23 days.
  • With three copies, multiply by three if every copy is verified independently.

Checksum computation also takes CPU, especially for stronger algorithms. And verification competes with ingest, access and preservation actions for the same storage and network resources.

The lesson is that fixity schedules must be designed around realistic throughput, not wishful frequency. A policy that says “verify everything monthly” is meaningless if the infrastructure can only read the archive every four months.

Strategies for scale

Set frequency by risk

Not all content needs the same verification frequency. Newly ingested content, content on older media, content recently migrated and particularly valuable collections may be checked more often. Stable content on healthy storage may be checked less often. Many archives aim for every object on every copy to be verified at least once a year, with more frequent checks for higher-risk content.

Spread verification continuously

Rather than running large verification campaigns, run fixity checks continuously in the background at a controlled rate, so the full archive is covered over the target cycle without spikes that disrupt other work.

Verify close to the data

Moving every byte across the network to a central verification server is slow and expensive. Running checksum computation close to storage, on storage nodes or on servers in the same data center, reduces network load. Some storage platforms can compute checksums internally and report them.

Use storage-level integrity checking

Modern object storage platforms typically store checksums for every object or fragment and scrub data in the background, detecting and repairing silent corruption from redundant data automatically. This does not replace the archive’s own fixity checks, which verify end-to-end integrity against checksums recorded at ingest, but it greatly reduces the chance that preservation-level checks find problems and provides an additional layer of protection. See data durability in high-density storage systems and the scality.com post on object storage integrity verification.

Sampling and tiered verification

Some archives combine full verification on a longer cycle with frequent statistical sampling, checking a random selection of objects regularly to detect systemic problems early. Sampling is not a substitute for full coverage, but it gives faster warning of issues affecting many files.

Verify all copies, differently

Each copy should be verified, but not necessarily at the same frequency. An online copy may be checked continuously, while an offline tape copy is checked on a longer cycle or when tapes are migrated.

Choosing checksum algorithms

MD5 is still widely used in preservation because it is fast and well supported, and for detecting accidental corruption it remains effective. SHA-256 and other stronger algorithms are preferred where protection against deliberate tampering matters, and are increasingly used as the primary or secondary checksum. Some archives record more than one algorithm at ingest. Changing algorithms later requires reading every file again, so choose carefully. Faster modern algorithms may reduce CPU cost for very large archives.

Packaging affects fixity

How content is packaged has a big effect on fixity workload:

  • Many small files increase per-file overhead in reading, checksum computation and logging.
  • Large container files reduce per-file overhead but mean a single corrupted byte may require the whole package to be repaired.
  • Manifest-based packages, such as BagIt, record checksums for every file inside a package, allowing verification at both levels.

Consider fixity when designing packaging and storage together. See what the OAIS model asks of storage.

Handling failures

When a fixity check fails:

  • Confirm the failure by re-reading the file, since transient read errors can cause false alarms.
  • Check other copies to find a verified good version.
  • Repair the damaged copy from a good one and verify the repair.
  • Investigate the cause, such as a failing drive, faulty controller or software bug, and check whether other files are affected.
  • Record the incident and repair in preservation metadata.

Patterns of failures, such as several in the same storage pool, can reveal hardware or software problems that need attention beyond individual repairs.

Recording and reporting

Fixity results are evidence. Auditors and certification schemes look for records showing that every object has been verified regularly and that failures were handled. Store results in the preservation system’s database, report coverage and failure rates and keep historical logs. Dashboards showing the percentage of the archive verified within the target cycle help managers see whether the program is keeping up.

A worked scheduling example

Imagine an archive with 6 PB per copy across three copies: two online object storage copies at different sites and one tape copy. The online platforms can sustain about 3 GB per second of verification reads each without affecting ingest and access. Reading 6 PB at 3 GB per second takes about 23 days, so each online copy could in principle be fully verified roughly monthly. In practice, the archive runs verification at a lower background rate, completing a full cycle every quarter, and checks newly ingested content within its first week. The tape copy is verified on a two- to three-year cycle aligned with tape migrations, plus a small random sample each month. Statistical sampling of 0.1 percent of objects on each online copy runs weekly to spot systemic problems early. This combination gives quarterly full coverage online, early detection and complete coverage of every copy within the archive’s policy.

Fixity for content in transit

Many integrity failures happen not at rest but during movement: ingest from producers, transfers between sites, migrations and delivery to users. Checking fixity at each handoff, comparing checksums before and after the move, catches problems at the point they occur, when they are easiest to fix. Transfer tools and preservation systems that verify automatically during movement reduce the burden on periodic checks.

Infrastructure for fixity at scale

  • Storage with high aggregate read throughput that scales with capacity.
  • Network capacity between storage and verification processes.
  • Compute for checksum calculation, ideally near the data.
  • Scheduling software that spreads verification and prioritizes risk.
  • Storage-level scrubbing to catch corruption early.

Scale-out object storage suits this well: as capacity grows by adding nodes, read throughput grows too, keeping verification cycles stable. See object storage vs traditional storage.

Checklist: fixity checks at scale

  • Compute and store checksums for every file at ingest.
  • Calculate how long a full verification cycle takes at real throughput.
  • Set verification frequency by risk, aiming for full annual coverage at minimum.
  • Run verification continuously and close to the data.
  • Use storage platforms with background integrity checking.
  • Add statistical sampling for early warning.
  • Verify every copy, including offline copies, on defined cycles.
  • Automate confirmation, repair and recording of failures.
  • Report coverage and failure rates for audits.

Putting it together

Fixity checks at scale are a capacity planning problem as much as a preservation one. Work out how fast the infrastructure can read the archive, set verification cycles that match, prioritize higher-risk content and run checks continuously close to the data. Let the storage platform scrub and repair beneath the preservation system, and keep thorough records. Done well, fixity checking becomes a routine background process that gives the archive, and its auditors, confidence that every object is still intact.

Frequently asked questions

What is a fixity check?

A comparison between a checksum computed when content entered the archive and a checksum computed later, confirming the content has not changed.

How often should fixity be checked?

Many archives aim to verify every object on every copy at least annually, with more frequent checks for higher-risk content.

Does storage-level scrubbing replace fixity checks?

No. It catches and repairs corruption within the storage platform, but archives still need end-to-end checks against checksums recorded at ingest.

Is MD5 still acceptable for fixity?

It is effective for detecting accidental corruption and widely used. Stronger algorithms such as SHA-256 are preferred where tampering is a concern.

What happens when a fixity check fails?

Confirm the failure, find a good copy, repair the damaged copy, investigate the cause and record the incident.

Further reading

See digital preservation storage, the OAIS model and storage, digital preservation copies, research data management storage and data durability in high-density storage systems.