7 A university S3 storage service gives researchers on-demand object storage, accessed through the same S3 API used by public clouds, but run by the institution on its own infrastructure. Researchers get a familiar, scalable place to store datasets, share data with collaborators, feed pipelines and archive results. The university gets predictable costs, data that stays on campus or in-country, integration with institutional identity and a central platform it can govern. Many universities and research institutes now offer S3 alongside traditional file storage and HPC scratch. This article explains how universities design and run S3 storage as a service: the platform, identity and access, project and quota models, cost recovery, data sharing, support and governance. For broader context, see our hub on HPC data tiering and the guide to research data management storage. Why offer S3 to researchers Modern tools speak S3: workflow managers, data science tools, AI frameworks, data portals and many instrument and analysis applications read and write S3 natively. Scale: object storage grows by adding nodes, accommodating large and unpredictable research datasets. Sharing: S3 makes it easy to share data with collaborators through access policies or time-limited links. Cloud compatibility: researchers can move workflows between campus and public cloud with minimal changes. Cost: on-premises object storage typically costs less per terabyte than high-performance file storage and avoids public cloud request and egress charges. Data location: data stays within the institution and country, important for sensitive research and funder requirements. The platform A university S3 service typically includes: S3-compatible object storage deployed across racks or sites, protected with erasure coding and optional replication. Load-balanced endpoints with TLS and stable DNS names. An account and identity layer integrated with university authentication. A self-service portal or request process for creating projects, buckets and credentials. Monitoring and reporting of capacity, usage and performance per project. Identity and access Researchers should use institutional identities rather than unmanaged keys where possible: Single sign-on to the self-service portal using the university identity provider. Federated identity for external collaborators, for example through research and education identity federations. Access keys issued per user or application, scoped to specific buckets, with expiry and rotation. Group-based access reflecting research groups and projects. Bucket and object policies for fine-grained sharing. S3 access policies covers policy design. Project and quota models Most services organize storage around projects or research groups rather than individuals, so data survives when people leave: Projects with a responsible principal investigator, members and an end date. Buckets created within projects, with naming conventions. Quotas on capacity, and sometimes object counts, per project. Default allocations for all researchers, with larger allocations on request or for a fee. Cost recovery Universities fund S3 services in various ways: Free baseline allocations funded centrally, with charges above the baseline. Per-terabyte charges to grants, where funders allow storage costs. Up-front payments for a defined retention period, useful for data that must be kept after grants end. Central funding as research infrastructure. Price per terabyte should be simple and competitive with alternatives, so researchers choose the managed service over unmanaged drives or personal cloud accounts. Show usage regularly to project owners. Data sharing and collaboration S3 makes collaboration easier: Internal sharing through group access to buckets. External collaborators through federated identities or dedicated credentials. Time-limited links for sharing individual datasets. Public datasets with controlled anonymous read access, for open data publication, with safeguards against accidental exposure of sensitive data. Provide guidance on safe sharing, especially for personal or sensitive data. Integration with research computing HPC clusters access S3 for staging data in and out of scratch. GPU and AI platforms read training data from S3. Instrument facilities land data directly in S3 buckets. Data portals and repositories use S3 as their storage back end. Data transfer tools move data between campus S3, other institutions and public clouds. Lifecycle and governance Project end dates trigger reviews: archive, publish, transfer or delete. Lifecycle rules remove temporary data or move it to lower-cost tiers. Versioning and object lock protect important data from accidental deletion or ransomware. Data classification guides which data may be stored and shared, with stricter controls for sensitive data. Audit logging for access to sensitive buckets. Supporting researchers Technical availability is not enough. Successful services provide: Documentation with examples for common tools and languages. Training sessions and office hours. Templates for common setups, such as lab data buckets or public dataset publishing. Help with migrations from external drives, departmental servers or public cloud. Clear service levels for availability and support response. Launching the service A phased launch reduces risk. Start with a pilot group of research teams that already use S3 tools, such as a genomics lab, an imaging facility and a data science group. Use their feedback to refine documentation, quotas and processes. Then open the service more widely with a clear service description, a pricing page and a simple request form or portal. Announce it through research offices, departmental IT contacts and training sessions, and showcase early success stories. Track adoption by department and storage growth, and adjust capacity plans as demand becomes clearer. Common pitfalls Unmanaged access keys shared between people and never rotated. Buckets owned by individuals rather than projects, so data is stranded when someone leaves. Accidental public exposure through overly permissive policies. No end-of-project process, leaving storage full of abandoned data. Pricing that is hard to understand, pushing researchers back to external drives or personal cloud accounts. Insufficient support, especially for researchers new to object storage concepts. Measuring the service Useful indicators include the number of active projects, storage used and growth by department, share of research data held in managed storage versus unmanaged locations, support ticket volumes and themes, and researcher satisfaction from periodic surveys. These metrics help justify investment and guide improvements. Capacity and performance planning Research demand is lumpy. A single new instrument, large grant or AI project can add petabytes. Keep headroom, plan expansions in increments and maintain a pipeline of upcoming projects with expected data volumes, gathered through research offices and facility managers. Monitor throughput as well as capacity, since AI training and large data transfers can stress the platform even when capacity is ample. Regional considerations In Europe, research involving personal data must comply with GDPR, and some projects require data to stay within the EU or country. In the UK, UK GDPR applies. In the US, federally funded research may carry data security requirements, and human subjects data has its own rules. In Japan and elsewhere, national research data policies apply. An on-campus S3 service helps universities meet these requirements while giving researchers cloud-style convenience. Teaching and student use S3 services also support teaching. Courses in data science, bioinformatics and machine learning can give students temporary buckets with small quotas, expiring at the end of term, so they learn cloud-style workflows on institutional infrastructure without personal cloud accounts or unexpected bills. Automate creation and cleanup each term. Checklist: university S3 storage service Deploy scalable, protected S3-compatible object storage with stable endpoints. Integrate with institutional single sign-on and research federations. Organize storage by project with responsible owners and end dates. Set quotas, default allocations and simple pricing. Support safe internal, external and public sharing. Integrate with HPC, AI platforms, instruments and repositories. Apply lifecycle rules, versioning and object lock where needed. Classify data and enforce controls for sensitive data. Provide documentation, training and migration help. Report usage and cost to project owners. Putting it together A university S3 storage service brings cloud-style object storage to campus: scalable, API-driven, easy to share and integrated with research computing, while keeping costs predictable and data under institutional control. Organize it around projects, integrate it with institutional identity, price it simply and support researchers well, and it becomes the default home for research data across instruments, pipelines, collaborations and archives. Frequently asked questions Why would a university run its own S3 service? For predictable costs, data location control, integration with campus identity and research computing, and support for modern tools. How do researchers access a campus S3 service? Through the S3 API with access keys issued via a self-service portal using university login, and through tools and applications that speak S3. Can external collaborators use a university S3 service? Yes, through federated identities, dedicated credentials or time-limited links, subject to data sharing policies. How are university S3 services funded? Through combinations of central funding, free baseline allocations, per-terabyte charges to grants and up-front payments for retention. Can S3 storage be used for sensitive research data? Yes, with appropriate classification, access controls, encryption, audit logging and data location controls. Further reading HPC data tiering Parallel file system tiering to object storage Instrument data storage Research data management storage