Monday, October 5, 2026
Home » What Storage Architecture Do Large Email Providers Use?

What Storage Architecture Do Large Email Providers Use?

Email storage architecture is one of the oldest large-scale storage problems in IT, and one of the least discussed. Telcos, internet service providers, hosting companies and webmail platforms store mailboxes for millions of users, often for decades. Every mailbox is a mix of tiny metadata updates and many small immutable messages, with a long tail of attachments. Most mail is written once and rarely read again, yet users expect any message from years ago to open instantly on a phone.

This article explains how large mail platforms are typically built, how storage formats and backends have evolved from local disks to NAS to object storage, and what architects should weigh when they redesign the storage layer. It is written for platform architects and mail engineers at providers, though the same ideas apply to large enterprises running their own mail systems.

The building blocks of a mail platform

A large mail platform usually separates into layers:

  • Edge and proxy servers that accept IMAP, POP3 and webmail connections and route each user to the right backend.
  • Mail transfer agents (MTAs) that receive incoming mail, run spam and malware filtering and hand messages to delivery.
  • Mail delivery and access servers that write messages into mailboxes and serve them to clients.
  • Index and metadata that track folders, flags, message order and search data for each mailbox.
  • Mail storage that holds the messages themselves.
  • Directory and authentication services that map users to backends and credentials.

Storage architecture decisions mostly concern the last three layers: how messages are laid out, where indexes live and what backend holds them.

Mailbox formats

How messages are written to storage affects performance, efficiency and how well the platform scales.

mbox

The oldest format stores a whole folder in one file, appending each new message. It is simple but performs poorly at scale, since deleting or changing one message can mean rewriting a large file.

Maildir

Maildir stores each message as its own file, with flags encoded in the file name. It is robust and easy to reason about, and it became the standard for many large platforms on NAS. The downside is very large numbers of small files, which stress file system metadata and backup tools as mailboxes grow.

Proprietary mailbox formats

Some mail servers use their own proprietary mailbox formats, with single (one message per file) and multi (many messages per file) variants and separate index files holding metadata. They reduce some of the overhead of Maildir.

Object-based layouts

Object storage plugins store each message as an object with a unique identifier, and pack index data into bundles uploaded to the object store. This removes dependence on a shared file system.

Storage backends: from local disk to object storage

Local disks and direct-attached storage

Early large platforms placed users on individual servers with local storage. This is cheap but rigid: a server failure takes its users offline, and rebalancing users means copying mailboxes between servers.

NAS and shared file systems

Most large mail platforms of the last two decades moved to network-attached storage over NFS. Any backend server can serve any user, which makes failover and load balancing simpler. As platforms grew to billions of files, NAS brought its own problems: metadata performance limits, file locking issues over NFS, expensive scale-up arrays, large and slow backups and disruptive migrations at every hardware refresh.

Object storage

Object storage scales out across many standard servers, handles billions of objects and protects data with erasure coding or replication rather than RAID arrays. For mail, it fits the workload well: messages are immutable once delivered, each can be a single object, and there is no need for shared file system locking. Index and metadata activity is handled by local caches and bundling on the mail servers, with the object store holding the durable copy. The general differences are covered in object storage vs traditional storage.

Indexes, caching and metadata

Mail clients constantly ask for folder lists, flags, unread counts and message headers. Serving those from the storage backend for every request would be slow. Large platforms rely on:

  • Mailbox indexes that summarize each folder’s state.
  • Local caches on the mail servers for active users’ indexes and recently accessed messages.
  • Metadata databases for mapping users, folders and objects, especially in object-based designs, where listing objects is avoided for performance.
  • Search indexes for full-text search across mailboxes.

The storage backend must handle frequent small writes from index updates as well as larger message reads. Object-based designs batch index updates into bundles, which turns many small changes into fewer, larger writes.

Durability, availability and multi-site

Mail is a high-visibility service. Users notice even short outages, and lost mail generates complaints, churn and, for business customers, contractual claims. Architectures typically include:

  • Data protection within the storage platform, such as erasure coding that survives multiple drive or node failures.
  • Multiple sites, with users served from one site and data replicated to another, or active-active designs where users can be served from either site.
  • Stateless or near-stateless mail servers, so a failed backend can be replaced without data recovery.
  • Backup and recovery for user-initiated deletions and platform-wide incidents, which is covered in email archiving vs email backup.

Cost and efficiency

Mail storage grows steadily. Users rarely delete mail, mailbox quotas rise over time and attachments get larger. Cost drivers include:

  • Raw capacity and protection overhead.
  • The number of files or objects, which affects metadata performance and backup.
  • Power, rack space and hardware refresh cycles.
  • Operations effort for rebalancing, migrations and backup.

Moving cold mail to lower-cost capacity, compressing messages and deduplicating identical attachments can all reduce cost, though each adds design complexity. The trade-offs between tiers are discussed in hot storage vs cold storage.

Scaling to millions of mailboxes

Growth is where architectural choices are tested. A platform built on scale-up NAS hits ceilings on metadata performance and array size, and each new array creates another island to manage. Scale-out designs, where both mail servers and storage nodes are added incrementally, avoid those ceilings. Our article on how to scale mailbox storage to millions of users covers the patterns in detail.

Security and privacy

Mail contains some of the most sensitive personal data a provider holds. Storage design should cover encryption at rest, encryption in transit between mail servers and storage, strict administrative access with audit logging and clear data location for regulatory purposes such as GDPR. Many providers also need the ability to delete a user’s data completely when an account closes, with evidence that it was removed.

When to redesign the storage layer

Common triggers for re-architecture include a NAS platform reaching end of support, backup windows that no longer fit, metadata performance problems during peak hours, the cost of the next array upgrade, or a consolidation after acquisitions. Moving a live mail platform is a significant project, but it can be done incrementally, user by user.

Checklist: evaluating email storage architecture

  • Which mailbox format do you use, and how many files or objects does it create?
  • Can any mail server serve any user, or are users pinned to storage islands?
  • Where do indexes and caches live, and how are they rebuilt after failure?
  • How does the storage backend scale: up, by adding arrays, or out, by adding nodes?
  • How is data protected within a site and across sites?
  • How long do backups and restores of a large mailbox population take?
  • What is the cost per stored terabyte, including refresh and operations?
  • Can users be migrated to new storage incrementally without downtime?
  • Are encryption, access control and deletion verifiable for compliance?

Putting it together

Email storage architecture has moved from local disks to NAS and, increasingly, to object storage, driven by growth in users, mailbox size and the number of small files. The right design keeps mail servers as stateless as possible, uses local caching and bundled indexes for responsiveness, and places durable data on storage that scales out, protects itself and can be refreshed without moving every mailbox. For providers planning their next platform, the goal is a design that can grow for another decade without another forklift migration.

Frequently asked questions

What storage do large email providers use?

Many large platforms historically used NAS over NFS with Maildir or similar formats. Increasingly, providers use object storage with mail server plugins that store messages as objects and cache indexes locally.

Is Maildir still used?

Yes, it remains common, especially on NAS. Its one-file-per-message design is robust but creates very large file counts at scale.

Why move email to object storage?

Object storage scales out to billions of objects, avoids NFS locking and metadata limits, protects data with erasure coding and allows incremental hardware refresh.

How do mail platforms keep IMAP responsive on object storage?

They cache indexes and recent messages locally on mail servers, bundle index updates and use metadata databases instead of listing objects.

How much storage does a large mail platform need?

It depends on user count, quotas, average mailbox size, retention and protection overhead. Platforms with millions of users commonly reach petabyte scale.

Further reading