11 AI has exposed the limits of siloed storage Enterprise storage was designed around relatively stable categories and access patterns. High-performance data lived on one system. General-purpose file and object data lived on another. Backups had their own target, while archive data moved somewhere cheaper for long-term retention. AI is not a single workload, which makes existing storage models increasingly difficult to sustain. AI pipelines consist of workloads across many phases, typically pre-model stages (data collection and cleansing), model training and fine-tuning, then actual interaction with models during inferencing. Inferencing will increasingly depend on fast KV cache access for extending memory across simultaneous user sessions and growing context windows with high performance, and the use of RAG to incorporate private data. Many organizations will also require long-term data retention for governance and auditability. Each of these phases places different demands on storage throughput, latency, concurrency, protection, and cost. Those requirements will likely change throughout the life of the data. Data requires ultra high-performance storage for training and inferencing, and can eventually age or “cool” sufficiently to be placed on reduced-performance but lower-cost storage, and later onto colder long-term retention media. When each stage relies on a different storage platform, managing the lifecycle becomes a continuous series of data copies and migration projects. AI data storage must satisfy several requirements at once: Keeping GPUs fed. Utilization of these expensive compute resources depends on several factors, and storage is one of the most important. Because inferencing is memory-bound, storage has to deliver enough throughput and low enough latency to act as an extension of the KV cache: a scalable store behind GPU HBM and host DRAM that doesn’t materially degrade performance. Storage media cost and scarcity. AI has created a demand supercycle, and NAND flash shortages in particular have driven up prices and lengthened lead times for high-performance storage, resulting in the delay of critical IT projects. Power and cooling. These have become hard constraints. Organizations can’t keep exabytes of data on hot and warm media for its full life, especially when much of it doesn’t need to be there and belongs on more power-efficient cold storage. Provable recoverability. Ransomware increasingly targets backups and recovery systems, so infrastructure has to demonstrate to boards, regulators, and insurers that the recovery path survives an attack. Sovereign control. Organizations must control where data resides, who can access it, and how the underlying infrastructure operates. Scaling storage capacity with flat IT headcount. As AI fuels further growth of data, all of this falls to infrastructure teams whose headcount isn’t growing at the same rate as their data. They’re being asked to scale capacity and performance, strengthen protection, demonstrate compliance, and manage increasingly diverse workloads without adding proportional operational complexity. It’s clear the demands on enterprise data have changed. The architecture supporting it now has to change with them. A new operating model to serve the full data lifecycle Adding another specialized storage product may solve an immediate need, but each platform adds another namespace, security model, set of policies, and migration path for infrastructure teams to manage. A better model is needed to follow the lifecycle of the data. Infrastructure teams should be able to adjust performance, protection, and economics as requirements change while keeping security, governance, and operations consistent. They should also be able to scale capacity, throughput, and metadata independently, without expanding the entire environment or launching another migration project. Policy-driven automation is also essential at this scale. Data placement and protection should follow customer-defined requirements as its value, use, and access patterns change. This reduces the manual work of tracking those changes and moving data between separate systems. This is the challenge Scality ADI was created to address. Scality ADI (Autonomous Data Infrastructure) is a data infrastructure operating model for enterprise AI, built on Scality’s proven distributed storage foundation. It autonomously aligns the right storage media, performance, and protection to each stage of the data lifecycle, from GPU-speed flash to deep archive. Scality ADI is built on the RING storage engine, which is field-proven for performance and durability in production at 12+ exabytes. With Scality ADI, customer-defined policies govern placement, protection, and lifecycle transitions. AI, sovereign cloud, backup and archive data stay under a single namespace, with consistent security, governance, and operations. The result is infrastructure designed to adapt with the data while addressing the performance, resilience, sovereignty, and economic demands enterprises now face. Four data temperatures, one lifecycle At the center of this model are four data temperatures: Extreme, Hot, Warm, and Cold. Each temperature uses the storage media and architecture best suited to its role. Together, they let data move from GPU-speed access to economical long-term retention as its requirements change, all within the same namespace: Extreme tier: Delivers ultra-low latency for KV cache in AI inferencing workloads, using the native AI API over RDMA with TLC flash SSDs. Hot tier: Combines high performance with resilience for GPU-intensive AI, analytics, and fast backup, using the S3 API with S3 over RDMA or TCP, and can leverage TLC, QLC, or nearline SSDs. Warm tier: Balances throughput, durability, and capacity economics for sovereign cloud and backup at scale, extending S3 access across SSD and HDD configurations. Cold tier: Prioritizes long-term retention economics for archive and compliance data, connecting to tape and cloud archive targets. Scality ADI carries out policy-driven lifecycle transitions outside the application data path. Customer-defined lifecycle rules provide the policies for when data can be moved across tiers. Throughout these transitions, data remains under one security and governance model. Teams can adapt performance, protection, and economics without creating another storage silo or managing another migration project. This approach also improves infrastructure economics and energy efficiency. With Scality ADI, premium media is reserved for the data that benefits from it, while less active data moves to lower-cost, lower-power storage. That becomes increasingly important as power and cooling place practical limits on data center growth. Extreme tier: Flash performance for AI inferencing For ultra-demanding AI inferencing workloads, Scality ADI can serve as an extended KV cache using the streamlined AI object storage API over RDMA. In a 6-server test (1), Scality ADI demonstrated the following performance highlights for Time-to-First-Token (TTFT) and performance maintained while scaling sessions and token context length (tested on Gemma-3 27B with 14K token context baseline): Near-DRAM speed for TTFT, much faster than recompute: RING responds in 166ms, adding only 64ms vs local host DRAM (102ms). Performance is maintained as KV cache grows: Scaling to 850 concurrent sessions, RING is 4.9X faster than recomputing the cache from scratch (TTFT 373ms vs. 1821ms cold), and this grows to 10.3X faster with 128KB token context (TTFT 273ms vs. 2807ms cold). The Extreme tier is deliberately non-resilient. Because KV cache is ephemeral and regenerable, Scality ADI trades replica and parity protection for speed, while the durable copy of the model or dataset lives in the tier beneath. Hot and Warm tiers: Hyperscale performance without making everything all-flash AI infrastructure planning often starts with the assumption that serious performance requires all-flash storage. Some workloads require flash-level latency. Many depend more heavily on sustained aggregate throughput. Scality ADI addresses both requirements. The Hot tier provides GPU-direct access for latency-sensitive workloads over TCP or RDMA, while the Warm tier delivers high aggregate bandwidth on more economical media. Performance testing detailed in the Scality ADI Whitepaper demonstrates what this can achieve. Using 10 MB objects on a Warm S3 configuration with standard HDDs, Scality ADI on a large-scale cluster sustained over 420 GB/s read throughput and 250 GB/s write throughput. During the test, Scality ADI sustained 35,000 concurrent S3 operations across 696 buckets for two hours, with zero storage-layer errors. The same validation also exceeded 173,000 S3 PUT operations per second and 151,000 S3 GET operations per second with 4 KB objects. These results demonstrate that HDD-based object storage can deliver the aggregate throughput required by demanding AI, analytics, and backup workloads. Flash can then be reserved for workloads that genuinely require its latency. Read more about how to feed AI at scale without all-flash. Performance at scale also depends on predictable behavior when individual components slow down. If retrieving part of an object exceeds a defined latency threshold, Scality ADI can fetch an alternate replica or reconstruct the part from erasure-coding parity. Quality-of-service and rate-limiting controls help prevent one account, bucket, or workload from overwhelming shared resources. Scaling capacity, performance, and metadata independently Scality ADI is composed of distributed software services that abstract physical servers and drives into a single storage platform. Stateless connectors provide object and file access to applications. Storage nodes supply capacity and durability. Flash-backed metadata services track objects and namespaces, while the management layer operates outside the data path. Applications can access the same underlying storage through S3 or file protocols, including NFS, SMB, and Linux FUSE. This allows cloud-native and file-based applications to share the platform without maintaining separate object and file environments. The major infrastructure resources scale separately: Add storage nodes to increase capacity. Add connectors to increase throughput and operations per second. Add metadata clusters and RAFT sessions to increase metadata concurrency. Metadata becomes critical at exabyte scale, where an environment may contain hundreds of billions of objects. Scality ADI treats it as an independently scalable, strongly consistent subsystem. RAFT consensus protects consistency, while multiple sessions distribute metadata traffic. Busy buckets can be moved online or sharded by prefix without downtime or application changes. Data protection is similarly workload-aware. Small objects can be replicated for fast access, while larger objects use Reed-Solomon erasure coding to provide durability with lower capacity overhead. When a disk or server fails, Scality ADI rebuilds only the affected objects and distributes the work across the storage pool. Applications remain online while protection is restored. End-to-end checksums detect silent corruption on every read and automatically recover clean data from another replica or parity. Because Scality ADI is software-defined, organizations can use commodity x86 servers and standard Ethernet, adding RDMA-capable networking for the highest performance tiers. Nodes and drives can be added, removed or refreshed while the platform remains available, allowing infrastructure to grow across hardware generations without a forklift upgrade. Cyber resilience built into the storage foundation Ransomware attacks increasingly target backups before production data. If attackers can destroy the recovery path, encryption becomes far more damaging. Scality ADI addresses this threat at the storage-engine level. The engine never overwrites object data in place. A new write receives a new storage location, leaving the previous version untouched. That inherent immutability supports the Scality CORE5 cyber-resilience model: API resilience protects against malicious overwrites and deletes through versioning, lifecycle controls, and S3 Object Lock. Data resilience controls access and protects data through identity management and encryption in transit, at the API and on disk. Storage resilience uses replication, erasure coding, self-healing, and end-to-end checksums to withstand hardware failure and silent corruption. Geographic resilience protects against site-level outages through synchronous and asynchronous multisite configurations. Architecture-level resilience preserves protected data even when a privileged administrator account is compromised. With S3 Object Lock in compliance mode, a protected object can’t be modified or deleted before its retention period expires. Retention can be extended but can’t be shortened, including by a fully privileged administrator. The model also includes short-lived credentials, identity federation, encryption, multisite replication, and open audit logs that can be ingested into security and monitoring systems. Together, these controls provide evidence that the recovery path remains protected through hardware failure, site loss, ransomware, and privileged credential compromise. Connecting storage directly to AI workflows Scality ADI integrates with the access patterns and tools used to build AI and HPC platforms. For latency-sensitive workloads, GPU-direct storage transfers data directly between NVMe storage and GPU memory, bypassing unnecessary CPU and kernel copies. The Extreme tier targets sub-50-microsecond latency for ephemeral data such as KV cache and HPC scratch. For workloads that need resilience alongside high performance, S3 over RDMA combines an S3 control plane with direct data transfers to and from GPU memory. Data remains protected through erasure coding across the cluster. Kubernetes integrations bring storage provisioning into the application workflow: The Container Storage Interface, or CSI, presents Scality ADI as a file volume for applications that expect conventional file paths. The Container Object Storage Interface, or COSI, makes S3 buckets and access credentials declarative Kubernetes resources. S3 bucket notifications can trigger preprocessing, validation, training and restore workflows through open platforms such as Kafka and RabbitMQ. The same data can support GPU-adjacent processing, resilient high-performance access, economical capacity, and long-term retention without leaving the Scality ADI namespace. Operational intelligence at infrastructure scale Policy-driven lifecycle automation handles how data moves through the platform. Scality Guardian applies intelligence to the work of operating that infrastructure as it grows. Scality ADI exposes more than 2,000 metrics through Grafana and Prometheus, with system-wide, component-level, and troubleshooting views. Guardian adds AI-driven monitoring and analysis across system health, capacity, power, and performance. It detects anomalies, provides early warning of potential disk failures, capacity trends, and misconfigurations, plus it guides operators through root-cause analysis. Natural-language access also makes it easier to explore telemetry and technical documentation. Guardian extends these capabilities through the Model Context Protocol, or MCP. Operators can use the built-in agent or connect Scality ADI to their own AI stack. Agents can create, configure, and query Scality ADI within multi-step workflows, translating an operator’s intent into the underlying API calls for review. This allows Scality ADI to participate alongside compute and networking tools in broader infrastructure workflows. Storage can be provisioned as upstream systems require it, while operational tasks that once involved multiple interfaces and documentation lookups can be coordinated through a single instruction. Built for large-scale, multi-workload data estates Scality ADI is designed for large, long-lived data estates where workloads have different performance, protection, and economic requirements. Its strongest use cases include: AI and conventional enterprise workloads sharing infrastructure Petabytes to exabytes of unstructured data Data that changes in value and access requirements over time Sovereignty and regulatory requirements Ransomware-sensitive backup and recovery environments Infrastructure that must scale without a proportional increase in operational overhead Some environments have different needs. A dedicated parallel file system may remain appropriate for a bounded, specialized scratch workload. Public-cloud object storage can be a strong fit for elastic, cloud-native applications without sovereignty, locality, or sustained-scale cost constraints. Transactional databases and low-latency block applications require a different storage class. For smaller, backup-first environments, Scality ARTESCA object storage software provides a simpler entry point into immutable object storage. It’s designed for easy deployment, integrates with leading backup applications, and scales from 20 TB to multiple petabytes. Let the infrastructure follow the data The next operating model for enterprise data infrastructure must support far more than capacity growth. It needs to: Serve multiple AI access patterns and performance requirements Align media, protection, economics, and power with the data lifecycle Keep data under consistent security and governance Preserve recoverability through cyberattack and infrastructure failure Scale capacity, throughput, and metadata independently Reduce the manual work required to operate a growing data estate Scality ADI brings those requirements together through one distributed engine, one namespace, and four data temperatures. From GPU memory to deep archive, organizations can apply the right performance, protection, and economics at every stage of the data lifecycle while maintaining sovereign control and a consistent cyber-resilience model. Read the full Scality ADI technical whitepaper for a deeper look at the architecture, performance validation, and engineering behind the platform. Related content Independent article: Making Object Storage Fast Enough to Feed a GPU: Scality’s ADI, by Brian Booden Blog: Feeding AI at scale doesn’t require all flash Podcast (15 mins.): AI Broke Storage Economics: Not Every Workload Needs All-Flash Podcast (15 mins.): Can Object Storage Keep GPUs Fed? We Put It to the Test Podcast (15 mins.): Is Object Storage the Answer to the KV Cache Memory Bottleneck? Whitepaper: Scality ADI technical whitepaper Footnotes 1. Lab test configuration for Extreme tier performance test: Host GPU server HPE DL380a: 8 x NVIDIA RTX 6000; 2 x Intel Xeon 6737P, 2TB DDR5; 4 x 100G RDMA ports (2 x ConnectX-6-x), one per rail; 4 x 3.84TB local NVMe (baseline KV cache path) RDMA switch: HPE SN3700cM, 32 x 100GbE, single tier, RoCE lossless end to end Scality ADI: 6-node cluster HPE Alletra 4110: 2 x Xeon Gold 6442Y; 512GB DDR5; 20 x 3.84TB NVMe; 4 x 100G RDMA