3 Digital preservation copies are the foundation of every long-term archive. No storage system is perfect: drives fail, software has bugs, people make mistakes, buildings flood and attackers encrypt data. The only reliable defense against permanent loss is to keep multiple independent copies, in different places, protected in different ways. The question every library, archive and repository has to answer is how many copies are enough, where they should live and how different they need to be. This article explains common guidance on digital preservation copies, the risks each copy protects against, how to choose locations and technologies and how to balance protection against long-term cost. For the broader picture, see our hub on digital preservation storage. Why copies matter more than any single system Preservation storage faces threats that a single system cannot eliminate: Media failure: drives and tapes fail, sometimes silently. System failure: controllers, software bugs or firmware problems can corrupt data across an entire system. Human error: accidental deletion or misconfiguration. Site disasters: fire, flood, power loss or building damage. Regional disasters: earthquakes, storms or wide-area outages that affect multiple nearby sites. Cyberattacks: ransomware or malicious insiders targeting storage. Organizational risk: supplier failure, funding gaps or loss of expertise. Redundancy within a storage system, such as erasure coding, protects against media failure but not against most of the other threats. Independent copies do. What preservation guidance recommends The NDSA Levels of Digital Preservation The National Digital Stewardship Alliance’s Levels of Digital Preservation are widely used as a self-assessment tool, including outside the United States. For storage, the levels progress roughly as follows: At the first level, keep at least two complete copies that are not stored in the same place. Higher levels call for at least three copies, with at least one in a different geographic location. The most mature levels call for copies in locations facing different disaster threats, with active monitoring of storage and obsolescence. The exact wording has evolved between versions, but the principle is consistent: more copies, more separation and more diversity as preservation maturity increases. The 3-2-1 principle Borrowed from backup practice, 3-2-1 means three copies, on two different types of media, with one off site. Many archives use it as a minimum, sometimes extended to include an offline or immutable copy. See the 3-2-1-1-0 backup strategy. Lots of copies keep stuff safe The LOCKSS approach, developed in the library community, takes the idea further by distributing many copies across independent institutions and using them to detect and repair damage. Collaborative preservation networks follow similar principles. How many copies? Most institutions settle on three copies of their preservation masters as a practical baseline, with some keeping more for especially valuable or irreplaceable collections. Two copies leaves the collection exposed while one copy is being repaired or migrated. Three copies allow one to be lost or under repair while two remain, and support majority comparison when fixity checks disagree. Access derivatives, such as compressed images or streaming versions, usually need fewer copies, since they can be regenerated from masters. Where should copies live? Separate buildings At minimum, copies should not share a building, power supply or network. A fire in one data center should not affect the other copy. Separate regions At least one copy should be far enough away to avoid shared regional risks such as earthquakes, floods or storms. What counts as far enough depends on geography: in some countries, sites a few hundred kilometers apart face different hazards; in others, more distance is needed. Within national boundaries National libraries and archives often must keep copies within their own country under legal deposit, public records or data protection rules. That constrains location choices but can usually be met with domestic sites in different regions. In Europe, Japan and many other countries, national institutions keep all preservation copies in-country. Different organizations Some institutions place a copy with a partner organization or a collaborative network, reducing organizational risk. How different should copies be? Technology diversity If all copies use the same hardware, software and firmware, a single bug could affect them all. Many archives keep at least one copy on a different technology, such as an object storage platform for online copies and tape for an offline copy, or two different storage products. Administrative separation Copies should not be deletable by the same credentials. Separate administrative domains, so a compromised account cannot destroy every copy, are increasingly important as ransomware targets storage. Online and offline Online copies support access and frequent fixity checking. Offline or air-gapped copies, such as tape stored on shelves or immutable object storage in a separate security domain, protect against cyberattacks that reach online systems. See air-gapped backup storage. Immutability Object lock or WORM media can prevent copies from being changed or deleted during defined periods, protecting against both mistakes and attacks. See S3 object lock: immutability and WORM. Keeping copies in sync and verified Copies are only useful if they are correct. Preservation systems should: Track every copy’s location in preservation metadata. Verify fixity of all copies regularly, not just the primary. Repair damaged copies from verified good ones. Confirm that new content reaches every copy within a defined time. Re-verify copies after migrations or media refresh. See how archives run fixity checks at petabyte scale. Balancing protection and cost Each copy adds cost: storage, power, space, operations and refresh. Ways to manage cost include: Efficient protection within each copy: erasure coding uses less raw capacity than replicating full copies within a site. See erasure coding vs replication. Tiered copies: an online copy for access and verification, and lower-cost copies on dense or offline media. Prioritizing collections: more copies for unique, irreplaceable masters; fewer for content that can be recreated or obtained elsewhere. Shared infrastructure: collaborating with other institutions for geographically distant copies. Model total cost over at least ten years, including refresh cycles for each copy. See storage cost per terabyte. Copies during migration and refresh Migrations are when copies are most at risk. When one storage platform is replaced, that copy is temporarily in transition, and any error during the move could corrupt content. Good practice is to keep the other copies untouched during a migration, verify every migrated object against its recorded checksum and only retire the old platform once verification is complete. Storage platforms that refresh hardware in place, moving data between old and new nodes automatically while verifying it, reduce the number of full migrations an archive has to perform over its lifetime. Testing recovery from copies Having copies is not the same as being able to recover from them. Archives should periodically restore a sample of content from each copy, including the offline one, and confirm it matches recorded checksums. Full disaster recovery exercises, where a site is treated as lost and services are rebuilt from the remaining copies, reveal gaps in documentation, credentials and capacity that are much better found in a test than in a real incident. A common architecture Many national institutions arrive at a similar design: Copy 1: online object storage at the primary site, used for access and frequent fixity checks. Copy 2: online object storage at a second site in another region, kept in sync and also verified. Copy 3: tape or immutable storage in a separate security domain, possibly at a third location, verified on a longer cycle. A multi-site object storage platform can provide the first two copies with geographic distribution built in, while the third adds technology and administrative diversity. Checklist: digital preservation copies Keep at least three copies of preservation masters. Separate copies by building, region and threat profile. Keep copies within required national boundaries. Use technology diversity for at least one copy. Separate administrative credentials across copies. Include an offline or immutable copy. Verify every copy regularly and repair from good ones. Prioritize collections by value and replaceability. Model ten-year cost including refresh. Putting it together Digital preservation copies are how archives turn inevitable failures into recoverable events. Three copies of preservation masters is a common and sensible baseline, spread across separate buildings and regions, kept within national boundaries where required and diversified by technology and administration. Verify them all, repair from good copies and match the number of copies to the value of each collection. That combination keeps collections safe without making preservation unaffordable. Frequently asked questions How many copies are needed for digital preservation? Three copies of preservation masters is a common baseline, with more for especially valuable collections. Do preservation copies need to be in different locations? Yes. Copies should be in separate buildings and, for higher maturity, in locations facing different disaster threats. Should all copies use the same storage technology? Ideally not. Using a different technology for at least one copy reduces the risk of a single bug or flaw affecting every copy. Is erasure coding a substitute for multiple copies? No. It protects against media failure within a system, but not against site disasters, system bugs or cyberattacks. How do archives keep copies consistent? By tracking copy locations, verifying fixity of every copy and repairing damaged copies from verified good ones. Further reading See digital preservation storage, the OAIS model and storage, fixity checks at scale, research data management storage and air-gapped backup storage.