7 Instrument data storage has become one of the biggest challenges for research computing centers. Cryo-electron microscopes, high-throughput sequencers, light-sheet and electron microscopes, mass spectrometers, synchrotron beamlines and telescopes produce data continuously, often terabytes per day per instrument. That data must be captured reliably as it is generated, processed on HPC or GPU clusters, shared with researchers and collaborators and retained for years under institutional and funder policies. Instruments are expensive to run, so storage must never be the reason an experiment stops. This article explains the data flow from instrument to archive, the storage tiers involved, the role of object storage and the policies that keep instrument data manageable. It is written for research computing architects and core facility managers. For the wider tiering picture, see our hub on HPC data tiering. Why instrument data is different Continuous generation: instruments run around the clock during sessions, writing data constantly. High volumes: modern detectors and sequencers produce large files at high rates. Cryo-EM sessions with direct electron detectors, for example, commonly generate several terabytes per day per microscope. No tolerance for loss: re-running an experiment can be expensive or impossible, as with unique samples or scarce instrument time. Processing-heavy: raw data is processed into derived data through compute-intensive pipelines, often on GPUs. Shared facilities: core facilities serve many research groups and external users, each needing access to their own data. Long retention: raw and processed data may be required for reproducibility, publication and funder compliance. The data flow Acquisition Instruments write data to local acquisition computers or storage attached to the instrument. Local storage is usually small and must be emptied quickly to keep the instrument running. Transfer to a landing zone Data is transferred continuously or after each acquisition to central storage, a landing zone sized to absorb instrument output with headroom. Transfers should be automated and monitored, with checksums to confirm integrity. Processing Data is staged to high-performance storage near compute for processing: motion correction and particle picking for cryo-EM, base calling, alignment and variant calling for sequencing, reconstruction for imaging. Intermediate files can be large and short-lived. Results and sharing Processed results are stored in project storage and shared with research groups, sometimes externally through portals or S3 access. Archive Raw data and key processed results move to archive storage for retention, often object storage, with metadata linking them to projects, samples and publications. Publication and deposition Data supporting publications may be deposited in public repositories, such as community archives for cryo-EM or sequencing data, with local copies retained according to policy. The role of object storage Object storage increasingly serves as the landing zone, archive and sharing layer for instrument data: Scale: grows by adding nodes as instruments and projects increase. Ingest throughput: handles parallel writes from multiple instruments. S3 access: modern acquisition software, data movers, workflow managers and portals can read and write directly. Durability: erasure coding and multi-site replication protect irreplaceable data. Sharing: access policies and presigned URLs let facilities share data with internal and external users securely. Metadata: object tags can record instrument, session, project and sample identifiers. A common pattern lands data on object storage, stages it to a parallel file system for processing and writes results back to object storage. Retention decisions Instrument data volumes force hard choices about what to keep: Raw data: some fields keep all raw data; others keep raw data for a limited period and retain processed outputs long term, since processing can be expensive to repeat but raw data is very large. Intermediate files: usually deleted after processing. Processed results: kept for the life of the project and beyond. Published datasets: kept per funder and publisher requirements. Agree retention policies with facility users and research leadership, and automate them with lifecycle rules. Cryo-EM specifics Cryo-EM sessions produce large numbers of movie files that are processed into micrographs and particle stacks. Facilities typically: Stream data from the microscope to a landing zone during collection. Run on-the-fly preprocessing on GPU nodes to give feedback during the session. Archive raw movies on object storage, at least until processing and publication are complete. Keep processed particle stacks and maps on project storage for further analysis. Sequencing specifics Sequencing data includes raw signal or image data, base calls and downstream alignments and variants. Facilities typically: Transfer run outputs from sequencers to central storage automatically. Process with pipelines on HPC clusters. Retain compressed read files and key derived results, while deleting large intermediates. Apply data protection controls for human genomic data, which is sensitive personal data in many jurisdictions. Sizing instrument storage Landing zone: at least several days of peak output from all instruments, to absorb processing backlogs or outages. Processing storage: working sets for concurrent pipelines, on high-performance storage. Archive: accumulated raw and processed data over retention periods, plus growth from new instruments. Network: links from instruments to the landing zone and from storage to compute sized for peak rates. Facility operations Core facilities run instruments as shared services, so storage must fit facility workflows. Booking systems can create project and session identifiers that flow into data paths and object tags automatically, linking every dataset to the right group and grant. Facilities often provide users with a time-limited period of free storage after a session, after which data moves to the group’s own allocation or is deleted. Clear communication of these rules, plus automated notifications before data expires, avoids disputes and lost data. Monitoring is essential. Track transfer queues from each instrument, landing zone capacity, processing backlogs and archive ingest. An alert when an instrument’s transfers stall gives staff time to act before local acquisition storage fills and the session has to stop. Planning for new instruments Instrument capability tends to jump with each new detector or sequencer generation, often multiplying data rates. When a facility plans a new instrument, storage and network planning should start at the same time as the purchase, not after installation. Ask vendors for realistic data rates under typical protocols, model the effect on landing, processing and archive tiers and budget storage expansion as part of the instrument project. Treating storage as part of the instrument cost avoids the common situation where a new microscope or sequencer sits underused because the data infrastructure cannot keep up. External users and collaborations Many facilities serve external academic and industry users. Provide them with secure ways to retrieve their data, such as time-limited S3 credentials or download portals, rather than shipping disks. Agree in advance how long external data is kept and who pays for longer retention. Regional and regulatory considerations Human data from sequencing and imaging is subject to data protection laws such as GDPR in Europe and equivalent rules elsewhere, and some national research programs require data to remain in-country. Facilities serving international collaborations must manage access and location carefully. On-premises object storage helps keep sensitive data within institutional and national boundaries. Integrity from the start Compute checksums on the acquisition computer before transfer and verify them on arrival, so corruption in transit is caught while the original still exists on the instrument. Checklist: instrument data storage Automate transfer from instruments to a landing zone with checksums. Size the landing zone for several days of peak output. Stage data to high-performance storage for processing. Use object storage as landing, archive and sharing layer. Tag data with instrument, session, project and sample metadata. Define retention for raw, intermediate and processed data. Automate lifecycle rules and deletion. Protect irreplaceable data across sites. Apply data protection controls for human data. Size networks for peak instrument output. Putting it together Instrument data storage must capture every byte reliably, move it efficiently to compute and keep what matters for years. Object storage has become central to that flow, serving as landing zone, archive and sharing layer, with high-performance file systems handling processing. Clear retention policies, automation and good metadata keep growing instrument estates manageable and make sure valuable experimental data is never lost. Frequently asked questions How much data does a cryo-EM microscope produce? Modern cryo-EM microscopes with direct electron detectors commonly produce several terabytes per day during collection sessions. Where should instrument data land first? On a central landing zone with enough capacity for several days of peak output, often object storage or a high-performance file system. Should raw instrument data be kept forever? It depends on the field, cost and policy. Some keep all raw data; others keep it for a defined period and retain processed outputs long term. Why use object storage for instrument data? It scales, handles parallel ingest, supports S3 access for tools and portals, protects data durably and enables secure sharing. How is sensitive sequencing data protected? With access controls, encryption, audit logging and storage locations that comply with data protection laws. Further reading HPC data tiering Parallel file system tiering to object storage University S3 storage services Genomics data storage