16 By Giorgio Regni, chief technology officer, Scality. October 8, 2026. Most organizations got started with AI through a cloud API, and for pilots that was the right call. Production is a different story. Once agents move into daily work, token volumes multiply, every request carries proprietary code and documents, and the model behind a validated workflow can change on someone else’s schedule. In IDC’s June 2026 survey of 1,015 infrastructure decision-makers, 85% agreed their AI workloads need hybrid infrastructure, spanning on-premises and public cloud, with interest in turnkey systems growing over data sovereignty concerns. The practical question becomes which workloads belong on-premises first. The answer tends to be the high-volume, always-on work that touches sensitive data, with coding agents near the top of the list. Those are also the hardest workloads to run yourself. What does it take to run AI inference on your own infrastructure, across the entire organization? Every agent session carries context the model has to read before it can answer: the repository, the system prompt, the tool definitions, the conversation so far. The model keeps none of it between requests. Each time a developer returns, a GPU either recomputes that context from scratch or reads it back from wherever it was stored. At scale, that choice decides what inference costs, and it’s why a storage company now ships an inference stack. Because the hardest part of serving a model turned out to be memory. Scality AI Inference Factory, launched October 8, 2026, is an open-code stack built on Scality ADI (AI Data Infrastructure) for running open-weight models on your own infrastructure. It’s available as a software license or a fully managed service. Your coding agents read the same codebase all day Every returning session is a choice between recomputing the context and reading it back. Picture ten thousand developers opening their coding agents on a Monday morning. Each session loads the same repository, the same system prompt, and the same tool definitions. By lunch, every agent has re-read that context dozens of times, and a developer who steps away for a meeting comes back expecting the session to pick up where it stopped. Published measurements show how far this goes. In a study of agentic workloads released this year, 84.6 to 99.5% of input tokens were reused from cache, from one turn to the next. The same study found hit rates falling from about 98% to 16%-26% once too many concurrent agents crowded the cache. The model forgets everything between requests. Something else has to hold that context, and every time the context is missing, a GPU has to recompute work it has already done. Cached tokens: A tenth of the price, for a reason Most of the large providers price a cached input token at about a tenth of an uncached one, which tells you recomputing context is a cost worth avoiding. Anthropic’s published pricing, for example, bills a cache hit at 0.1x the input price, and less on its newest models. Caching saves the provider the recompute. The discount hands you part of that saving, and the provider keeps its margin. When most of your developers’ tokens are repeated, you pay for the same context turn after turn. Providers built that cache because recomputing context is expensive. On your own servers, that cache is hardware already deployed and does not cost you anything more. The cache is large, and we measured it in our own customer lab. One GLM-5.2 session with a 438,764-token context produced 189GB of KV cache, about 430KB per token. With the model loaded across all eight GPUs of our server, that one session did not fit in HBM, so it had to go to external storage. Ten thousand sessions like this would need about 1.9PB. The cache does not fit in GPU memory or in one server’s DRAM. In memory disaggregated GPU architectures, every GPU needs to be able to reach any cached token. Memory at that scale is a storage problem, and storage is what Scality has run in continuous production since 2009. The long-standing assumption is that storage is too slow to sit in the serving path, where every millisecond before the first token counts. We spent the past year testing whether that is still true. We connected GPUs to object storage and measured everything With the cache served from Scality ADI, we measured time to first token (TTFT) of 166 ms on a 14K-token context, only 83 ms behind GPU memory, and unlike GPU memory, capacity grows as you add storage. We wired GPUs directly into Scality ADI over RDMA, with the CPU out of the data path, and measured what came back. The lab was deliberately modest: one server with eight GPUs, five storage servers a few generations old, and four 100 Gbit/s RDMA links. The network was the limit in the lab, and the product ran on a 400 Gbit/s fabric. vLLM cuts the KV cache into fixed-size blocks and gives each one a name, so Scality ADI can store every block as an object. Even with S3 over RDMA, every object read still passes through the S3 control plane, which was built for the open internet — request signing, policy checks, metadata lookups, payload checksums and XML framing — on every request. The Scality AI Connector swaps it for our own stripped-down control plane, with single-pass authentication and no XML or multipart framing, so the control path runs in microseconds and the object model stays the same. For a read, the GPU reserves space in its memory and tells Scality ADI which object to put there. The GPU never drives the transfer itself; Scality ADI does. The TLC NVMe drive places the object in the storage server’s RAM over PCIe, and the network card sends it straight into GPU memory over RDMA. The data bypasses the CPU, the kernel and the software stack, which is how the connector reaches 97% of network line rate between GPUs and storage. Recomputing the 14K-token context took 2.3 seconds, 14 times longer than reading it back. On a bigger model the gap is larger. For GLM-5.2 with a 438,764-token context, recomputing took 529 seconds and reading it back took 7.3 seconds, 72 times faster. In a second run, with 8 vLLM instances (1 per GPU), Scality ADI held 8TB of KV cache — more than 80 times the memory of one GPU and about ten times that of the whole server. The first token stayed between 355 and 373 ms at every size from 69GB to 8TB. Inter-token latency did not change: the fetch completes before the first token, and every token after it runs from GPU memory. Because the cache sits in Scality ADI rather than inside one server, every GPU can reach it, and that’s what makes split serving possible: running the two phases of inference on separate GPU pools. Prefill computes the KV cache for a new prompt, and decode generates the answer, one token at a time. When both run on the same GPUs, every new prefill interrupts the decodes already running. Once split, each pool does one job full time and scales on its own, with both working from the same shared KV cache. The research behind it is public. DistServe (OSDI 2024) measured up to 7.4 times more requests per GPU within the same latency targets, moving the cache directly between GPUs. In our design, prefill writes the KV cache to Scality ADI instead, and decode reads it back, so any decode GPU can pick up any returning context, and the tier puts essentially no limit on the size of this cache. The same path loads models. The Gemma-3 27B model’s 54.86GB of weights loaded into one GPU in 1.9 seconds from Scality ADI, almost ten times faster than from the GPU server’s own local drives, because it reads from many drives across the storage tier at once. Depending on load, a GPU can swap models up to a thousand times a day. Every layer above the storage needed a fix We submitted every fix upstream. The serving engine, the weight loaders, the data movers, and the fabric had all been written for storage that is slow and small. None of them expected a remote tier answering at the speed of the hardware. Several broke in ways that only show up when you measure the storage side of inference. The clearest case was vLLM’s KV offload connector, which fetched about five times the cache blocks each restore needed. Anyone using that connector to offload to an object store was paying that tax. We corrected the connector, left the vLLM core untouched, and committed the fix upstream, so every object store behind that path benefits. We also extended ModelExpress, the open-source model loader, to pull weights straight from object storage. There is no Scality fork of anything in the stack. We package the stack, keep every layer current as models and serving engines change, and can run it for you. You control the model, the data, and the cost You keep your harness, choose from validated open-weight models and decide when anything changes. Your developers keep the tools they already use and change one setting, moving the inference endpoint from the cloud API to yours. Claude Code, Codex, OpenCode, and the other tools your developers use keep working as before. The model you choose stays the model you run. Version, quantization, and retirement change when you decide (not when a provider rebalances its fleet), so a result you validated in March reproduces in September. If you sell AI inside your own product, that reproducibility is part of what you promise your customers. The context also stays put. An external API never receives only a question: it receives the codebase, the retrieved documents, and the output of every tool the agent called. On your own infrastructure none of it leaves your perimeter. It stays under your jurisdiction, and nobody outside your company can switch it off. The whole stack can be securely run. The cost works differently too. A metered API bill grows with every token, so it grows fastest when adoption is working. On your own servers the budget is fixed, and the more your teams use it, the cheaper each token gets. What you deploy Scality AI Inference Factory packages the whole serving stack behind one endpoint. It includes vLLM, Dynamo, the Scality AI Connector, Scality ADI, and more. Validated open-weight model families include Mistral, Gemma, gpt-oss, Qwen, Kimi, GLM and DeepSeek. The stack deploys on standard servers from Dell, HPE, Lenovo, and Supermicro. You receive the source, can audit how inference state is stored and moved, and can submit changes for our engineers to review. The storage tier underneath already runs more than 12EB in production across 70+ countries. Talk to a Scality AI architect about your workload. Bring your traffic and a repository, and we will size it with you. Stop paying to rebuild context you already have. Explore on-premises AI inference with Scality and help your team run AI on their own terms.