6 Scaling mailbox storage from thousands to millions of users is less about buying bigger arrays and more about removing the design choices that stop a platform from growing. Mail platforms at telcos, internet providers and hosting companies grow in two directions at once: more users, and more data per user as quotas rise and attachments get larger. The platforms that scale smoothly share a few traits: users can move between servers freely, mail servers hold little permanent state, caching absorbs most of the read traffic and the storage backend grows by adding nodes rather than replacing systems. This article covers how to size mailbox storage, the architectural patterns that let a platform scale mailbox storage to millions of users, and the operational practices that keep it manageable. For the overall design, see our hub on email storage architecture. Start with a sizing model A useful model separates capacity from activity, since they scale differently. Capacity inputs Number of mailboxes, including inactive accounts that still hold data. Average mailbox size, which usually varies widely between free consumer accounts, paid accounts and business mailboxes. Growth per mailbox per year, driven by incoming mail volume and user deletion habits. Quota policy, which caps growth for some users but not all. Protection overhead from erasure coding or replication, plus copies at a second site. Retention of deleted mail, if deleted messages are kept for a recovery period. An illustrative example: 5 million mailboxes averaging 2 GB each is 10 PB of logical data before compression or protection. If average mailbox size grows 15 percent a year, the platform adds 1.5 PB of logical data in the first year alone. Real numbers vary widely, so build the model from your own account statistics. The general method is in storage capacity planning. Activity inputs Concurrent active users at peak, often a small fraction of total accounts. Messages delivered per second at peak. IMAP operations per user session, such as folder syncs, flag changes and message fetches. Average and median message size, since many messages are small. Delete and expunge rate. Activity drives the number of mail servers, cache sizes and the request rate the storage backend must sustain. A platform can have plenty of capacity and still struggle if the backend cannot keep up with small requests at peak. Remove the islands The first barrier to scale is usually that users are tied to specific storage. If each mail server or NAS volume holds a fixed set of users, every growth step means placing new users carefully, and every rebalance means copying mailboxes. Platforms that scale well decouple users from storage: A proxy or director layer routes each user to a mail server based on a consistent mapping, typically a hash of the user name. Mail servers are interchangeable, able to serve any user because data lives on shared storage. The storage backend is one logical pool, not a set of separate volumes. With this separation, adding mail servers adds processing capacity, and adding storage nodes adds capacity, independently of each other. Make mail servers close to stateless When mail servers keep only cached data, a failed server is a minor event: its users are routed to another server, which fetches their data from the backend. Object storage-based mail designs take this furthest, with messages and index bundles stored as objects and local disks used only for caches. On file-based platforms, the equivalent is shared NAS with no data on local disks, though NFS locking and metadata limits become the next constraint as user numbers grow. Let caching do the heavy lifting Most mail reads are for recent messages and indexes of active users. Local caches on mail servers absorb most of that load: Index caches keep folder state for active users on local SSD. Message caches keep recently read messages in memory or on SSD. User affinity at the proxy layer keeps each user on the same server so caches stay warm. Good caching means the storage backend sees mostly new deliveries, cache misses and periodic index uploads, which is a far more predictable load than raw client traffic. Choose a backend that scales out At millions of users, the storage backend must handle billions of objects or files, high rates of small requests and steady capacity growth. The options broadly are: Scale-up NAS, which performs well until metadata or controller limits are reached, then requires a new array and a migration. Scale-out file systems, which add nodes but still carry file system metadata and locking overhead. Object storage, which scales capacity and request handling by adding nodes, has no shared file system locking and protects data with erasure coding. For large mail platforms, object storage has become a common target because mail messages are immutable once delivered and map naturally to objects. Plan for small objects and high request rates Mail is a small-object workload. Many messages are a few kilobytes, index bundles are modest in size and deletes are constant. When evaluating a backend, test: Sustained small PUT and GET rates with realistic object sizes. Latency percentiles, not just averages, since slow outliers affect user experience. Delete throughput and how quickly capacity is reclaimed. Behavior during drive and node failures, including rebuild impact on latency. The scality.com blog’s article on S3 request performance covers the factors that matter, and object storage rebuilds explains what happens during hardware failures. Scale across sites Large providers usually run mail across at least two data centers. Common patterns include: Active-passive, with users served from one site and data replicated to another for disaster recovery. Active-active by user, where each user is homed to a site and can fail over to the other. Stretched storage, where one object storage platform spans sites so mail servers in either site can read any user’s data. The choice depends on latency between sites, recovery objectives and cost. Whatever the pattern, test site failover with real users. Control growth with policy Storage growth is not only a technical matter. Policies that keep growth manageable include: Quotas by account type, with clear upgrade paths for paid tiers. Inactive account cleanup, removing mailboxes that have not been accessed for a long period after notifying users, in line with terms of service and data protection law. Deleted mail retention limits, so recovery periods do not become indefinite storage. Compression of stored messages, which reduces capacity at modest CPU cost. Data protection rules such as GDPR also require that personal data not be kept longer than necessary, which supports cleanup of closed and abandoned accounts. Operate at scale Operational practices make the difference between a platform that grows smoothly and one that lurches from crisis to crisis: Monitor per layer: proxy routing, mail server load, cache hit rates and storage latency. Automate user moves and server replacement. Add capacity in small steps, well before thresholds are reached. Refresh hardware in place, replacing old storage nodes while the service runs. Test backup and recovery for both single-user restores and larger incidents. Checklist: scaling mailbox storage Build a sizing model separating capacity and activity. Decouple users from storage with a proxy layer and shared backend. Keep mail servers close to stateless, using local disks for caches. Size caches for peak active users and enforce user affinity. Choose a backend that scales out for capacity and request rate. Test small-object performance, latency percentiles and delete throughput. Design multi-site access and test failover. Apply quota, cleanup and retention policies. Automate operations and add capacity in small increments. Putting it together To scale mailbox storage to millions of users, remove the islands that tie users to specific servers, keep mail servers lightweight, let caching absorb read traffic and place durable data on storage that grows by adding nodes. Pair that architecture with growth policies and disciplined operations, and the platform can keep expanding without the periodic forklift migrations that have defined many mail platforms. If your platform is on NAS today, see migrating mail storage from NAS to object storage. Frequently asked questions How much storage do a million mailboxes need? It depends on average mailbox size, growth and protection overhead. At an illustrative 2 GB per mailbox, a million mailboxes is about 2 PB of logical data before compression and protection. What limits mail storage scaling on NAS? Usually metadata performance with billions of small files, NFS locking, controller limits and the cost and disruption of array upgrades. Why make mail servers stateless? So that any server can serve any user and a failed server can be replaced without data recovery, which simplifies scaling and failover. How important is caching for mail platforms? Very. Local index and message caches absorb most read traffic, making storage load more predictable and keeping IMAP responsive. Is object storage suitable for email? Yes. Messages are immutable once delivered and map naturally to objects, and object storage scales to billions of objects with erasure coding protection. Further reading Email storage architecture Email archiving vs email backup Migrating mail from NAS to object storage Storage capacity planning