Monday, October 5, 2026
Home » How Do Research Computing Centers Tier Data from HPC Scratch to Archive?

How Do Research Computing Centers Tier Data from HPC Scratch to Archive?

HPC data tiering is how research computing centers keep expensive, high-performance storage available for running jobs while holding years of research data on affordable capacity. Universities, national laboratories and research institutes run clusters that read and write data at enormous speeds, from simulations, genomics pipelines, imaging and increasingly AI training. The parallel file systems that feed those clusters are costly per terabyte and fill up quickly. Without a clear path from scratch to project storage to archive, centers face full file systems, frustrated researchers and data that is either lost or kept on the wrong tier forever.

This article explains the typical storage tiers in research computing, how data moves between them, the role of object storage and the policies that make tiering work. It is written for research computing and HPC storage architects. For research data governance, see research data management storage.

The tiers in a research computing center

Scratch

High-performance parallel file system storage, often on flash, used for active job input, output and intermediate files. Scratch is optimized for throughput and metadata performance. Files are typically purged after a set period, such as 30 to 90 days without access, and are not backed up.

Project or work storage

Shared storage for active research groups, holding datasets, code and results that persist beyond individual jobs. It may be on the same parallel file system as scratch or on a separate, larger-capacity tier, and is often backed up or snapshotted.

Home directories

Small, backed-up storage for user configuration, scripts and small files.

Archive or campaign storage

Large-capacity storage for data that is no longer actively computed on but must be kept for months or years: completed project data, raw instrument data, published results and data retained for reproducibility. Object storage and tape are common choices.

Preservation

For data with long-term value, preservation copies with fixity checking and multiple locations.

Why tiering matters

  • Cost: flash-based parallel file systems cost far more per terabyte than object storage or tape.
  • Performance: keeping scratch uncluttered preserves metadata and throughput performance for running jobs.
  • Capacity: research data grows faster than budgets for high-performance storage.
  • Data safety: scratch is not backed up, so valuable data must move to protected tiers.
  • Compliance: funders and institutions require research data to be retained and managed.

How data moves between tiers

Manual movement

Researchers copy data with tools such as rsync or parallel copy utilities. This is simple but relies on users remembering, and often results in data left on scratch until it is purged.

Policy-based tiering within the file system

Many parallel file systems support policies that move files based on age, size or access to lower-cost pools or external targets, keeping a stub or namespace entry so users still see the file.

Hierarchical storage management

HSM systems migrate file contents to an archive tier, often tape or object storage, and recall them on access. Parallel file systems commonly provide an HSM framework that uses data mover agents to move files.

Workflow-driven movement

Pipelines stage data in from archive to scratch before jobs and write results back out afterward, using workflow managers or data movers. This is common for instrument data and AI training.

The role of object storage

Object storage has become a central tier in research computing:

  • Capacity at lower cost than parallel file systems, scaling to many petabytes by adding nodes.
  • S3 access for modern tools, workflow managers, data portals and AI frameworks.
  • Durability with erasure coding and multi-site replication.
  • Online access, much faster to recall than tape.
  • Sharing, since objects can be shared with collaborators through presigned URLs or access policies.
  • Service model, letting centers offer S3 storage to researchers directly.

Many centers now use object storage as the campaign and archive tier, with tape as a deeper or secondary copy where very long-term, low-cost retention is needed.

Instrument and AI data

Instruments such as cryo-electron microscopes, sequencers, light-sheet microscopes and telescopes generate continuous streams of data that need to land somewhere reliable before processing. Many centers ingest instrument data directly to object storage or a landing file system, then stage it to scratch for processing. AI training adds large datasets that are read repeatedly, often staged from object storage to fast storage close to GPUs.

Policies that make tiering work

  • Scratch purge policies, clearly communicated, with warnings before deletion.
  • Project quotas that encourage moving inactive data to archive.
  • Archive allocations per project, priced or free, with clear retention terms.
  • Automation that moves data by age or project status.
  • Metadata requirements, so archived data can be found later.
  • End-of-project processes to decide what to keep, publish or delete.

A typical data journey

Follow a single dataset through a center. A research group runs a climate simulation that writes 200 TB of output to scratch over a week. Post-processing jobs reduce it to 20 TB of analysis-ready files in project storage, while the raw output is copied to the object storage archive with project metadata and the scratch copy is deleted. Over the following months, researchers query the analysis files and occasionally stage portions of the raw output back to scratch for new analyses. When a paper is published, the underlying data is registered in the institutional repository, with the archive copy serving downloads. At project end, the group decides which raw output to keep under the funder’s retention requirement, and the rest is deleted on schedule. At every step, the data lives on the tier that matches how it is used.

Measuring tiering effectiveness

Useful metrics include scratch utilization and the share of files untouched for more than 30 days, project storage growth by group, archive ingest and recall volumes, time to recall data from archive and the number of support tickets about lost or purged data. Tracking these over time shows whether policies and automation are working or whether researchers need more help moving data.

Sizing the tiers

  • Scratch: sized for active job working sets and throughput, not total data.
  • Project: sized for active datasets across groups.
  • Archive: sized for accumulated data over retention periods plus growth, which is usually the largest tier by far.
  • Network: sized for movement between tiers, particularly staging data to and from scratch.

Regional and funding considerations

Research computing in Europe, the UK, the US and Asia operates under different funding models and data policies. European national and regional centers often serve many institutions and must keep data within national or EU boundaries for some projects. Funders such as Horizon Europe, UK research councils and US agencies require data management plans that affect retention and sharing. Archive tiers must support those obligations, including data location where needed.

Checklist: HPC data tiering

  • Define scratch, project, home, archive and preservation tiers and their purposes.
  • Set scratch purge and project quota policies.
  • Use object storage as a scalable archive and campaign tier.
  • Automate movement with file system policies, HSM or workflows.
  • Plan instrument and AI data ingestion and staging.
  • Provide S3 access to researchers where useful.
  • Require metadata for archived data.
  • Size archive capacity for accumulation and growth.
  • Size networks for staging between tiers.
  • Align retention and location with funder requirements.

Putting it together

HPC data tiering keeps high-performance storage fast and affordable by moving data to the right tier at the right time. Scratch serves running jobs, project storage holds active work and object storage provides a scalable, online archive for everything that must be kept, with tape or preservation copies where needed. Clear policies, automation and S3 access make tiering part of everyday research rather than an afterthought, and keep research data safe long after the jobs that produced it have finished.

Frequently asked questions

What is scratch storage in HPC?

High-performance storage for active job data, typically on a parallel file system, purged after a set period and not backed up.

Where should HPC data go after a project ends?

To an archive tier, often object storage, with metadata and retention aligned to institutional and funder requirements.

Can object storage replace tape for HPC archives?

Often for the primary archive, because it is online and faster to recall. Many centers keep tape as a deeper or secondary copy.

How does data move from parallel file systems to object storage?

Through manual copies, file system tiering policies, HSM or workflow-driven staging.

Why not keep everything on the parallel file system?

It is expensive per terabyte, and clutter degrades performance for running jobs.

Further reading

Why HDDs still win for AI storage

The technical paper on delivering AI-grade throughput on object storage without all-flash cost.