5 Yes, you can run a data lakehouse on on-prem object storage, and a growing number of banks, government agencies, utilities, manufacturers and research organizations do. The lakehouse pattern combines the low-cost, scalable storage of a data lake with the reliability and performance features of a data warehouse, using open table formats on top of object storage. Because the storage layer speaks the S3 API, the same architecture that runs in public clouds can run in your own data center, with your own hardware, under your own control. This article explains how an on-prem data lakehouse is built, what it requires from object storage, how query engines and catalogs fit in and when on-premises makes more sense than cloud. It is written for data platform architects. For related background, see data lake vs data warehouse and object storage for data lakes. What a lakehouse is A data lakehouse stores data in open file formats, typically Parquet, on object storage, and adds a table format layer that provides database-like features: ACID transactions, so concurrent writers and readers see consistent data. See ACID transactions. Schema evolution, so tables can change without rewriting all data. Partitioning and pruning, so queries read only relevant files. Time travel, so users can query earlier versions of a table. Snapshot management for consistent reads and rollback. Distributed query and processing engines read and write these tables directly on object storage. A catalog keeps track of tables and their current metadata. The layers of an on-prem lakehouse Storage S3-compatible object storage holds data files, table metadata and manifests. It must scale to petabytes, sustain high throughput and handle large numbers of objects. This is the foundation; every other layer depends on it. Table format Open table formats define how tables are represented on storage. The leading formats have broad support across engines and work well on S3-compatible storage. See do open table formats work on S3-compatible storage on premises. Catalog A catalog maps table names to their current metadata location and coordinates commits. Options include metastore services carried over from older data platforms, REST-based catalog services and database-backed catalogs. The catalog must be highly available, since every query depends on it. Query and processing engines Distributed SQL engines serve interactive analytics and BI. Distributed processing frameworks handle large-scale processing and data engineering. Other engines serve streaming, machine learning and specialized workloads. Because data lives in open formats, multiple engines can share the same tables. See what object storage performance a SQL query engine needs. Governance and security Access control, auditing, data lineage and encryption sit across layers. In regulated environments, these are as important as performance. What the lakehouse needs from object storage S3 API compatibility Engines and table formats use S3 operations extensively: GET with byte ranges, PUT and multipart uploads, LIST, HEAD and DELETE. The storage must implement these correctly and consistently. Strong read-after-write consistency matters for table commits and metadata operations. High aggregate throughput Analytical queries scan large amounts of data in parallel. The storage must deliver high aggregate read throughput across many concurrent connections, scaling as nodes are added. Small object and metadata performance Table metadata, manifests and small data files generate many small requests. Storage must handle high request rates with low latency, not just large sequential reads. Scale Lakehouses grow quickly as new sources are added and history is retained. Storage must scale to many petabytes and billions of objects without disruptive upgrades. Data protection and resilience Erasure coding protects data within a site, and replication or multi-site designs protect against site failure. Versioning and object lock can protect critical data from accidental or malicious deletion. See erasure coding vs replication. Security Encryption at rest and in transit, fine-grained access policies, integration with identity systems and audit logging. The scality.com post on S3 access policies covers policy design. Why run a lakehouse on premises Data sovereignty and regulation Many organizations must keep data within their own country or facilities. In Europe, financial institutions face DORA and national supervisory expectations; public sector bodies and critical infrastructure operators face NIS2 and national security rules. Government and defense agencies in the US, UK, France, Germany, Japan and elsewhere often require data to stay in accredited facilities. An on-prem lakehouse keeps data under direct control. See how regulated firms keep a data lake inside their own country. Cost predictability Analytical workloads read data heavily. In public clouds, storage, request and data transfer charges grow with usage. On premises, costs are largely fixed and predictable, which suits steady, heavy workloads. See storage cost per terabyte. Proximity to data sources and compute Data generated on premises, from trading systems, sensors, manufacturing lines or scientific instruments, can be analyzed where it lands without moving it to the cloud. On-premises GPU and CPU clusters can read directly from local object storage. Avoiding lock-in Open table formats on S3-compatible storage keep data portable between engines and between on-premises and cloud. See how to avoid cloud lock-in. When cloud may fit better Cloud lakehouses can make sense for organizations with highly variable workloads, small teams, cloud-native data sources or no on-premises infrastructure. Many organizations end up hybrid, with sensitive or heavy workloads on premises and others in the cloud, sharing open formats. Migrating from Hadoop Many on-prem data platforms began on Hadoop with HDFS. Moving to a lakehouse on object storage separates compute from storage, reduces operational complexity and lets organizations adopt modern engines. See how to migrate from HDFS to object storage. AI and machine learning on the lakehouse Lakehouses increasingly feed AI and machine learning as well as BI. Feature engineering, model training and retrieval pipelines read large volumes of data from the same tables analysts query. Keeping these workloads on one governed copy of the data, rather than exporting extracts to separate systems, reduces duplication and keeps lineage clear. It does raise the bar for storage throughput, since training jobs read data in large parallel bursts. Planning the object storage layer for both analytical and AI access patterns from the start avoids a second, disconnected data platform later. See storage for generative AI for related considerations. A typical on-prem reference design A common on-premises design looks like this: an S3-compatible object storage cluster spread across racks, or across two sites for resilience; a highly available catalog service backed by a replicated database; a distributed SQL engine for interactive queries and BI tools; a processing cluster, often container-based, for data engineering; streaming ingestion writing to open table format tables; and a governance layer that enforces access policies and records lineage. Compute clusters scale independently of storage, so adding analysts does not require adding capacity, and adding years of history does not require adding compute. Sizing an on-prem lakehouse Capacity: current data, growth rate, retention and table history from snapshots, plus protection overhead. Throughput: peak concurrent query and processing load, measured in aggregate GB per second. Request rates: metadata and small-file activity, especially for frequently updated tables. Network: bandwidth between compute clusters and storage; analytical engines can saturate links. Growth: plan expansion in small increments. See storage capacity planning. Operating the lakehouse Lakehouses need regular maintenance at the table format level: compacting small files, expiring old snapshots and removing orphaned files. These tasks keep query performance high and storage consumption under control. Monitor storage throughput, latency and capacity alongside engine performance, since slow storage shows up as slow queries. Checklist: on-prem data lakehouse Choose an open table format with broad engine support. Select S3-compatible object storage with strong consistency and high throughput. Deploy a highly available catalog. Size storage for capacity, throughput, request rates and growth. Provide high-bandwidth networking between compute and storage. Implement access control, encryption and audit logging. Plan table maintenance: compaction and snapshot expiry. Design multi-site protection for critical data. Plan migration from legacy platforms such as HDFS. Putting it together An on-prem data lakehouse is not only possible but increasingly common. Open table formats, a reliable catalog and modern query engines run on S3-compatible object storage in your own data center, giving you cloud-style architecture with on-premises control, predictable cost and data that stays where regulation requires. The storage layer must deliver S3 compatibility, strong consistency, high throughput and scale; get that right, and the rest of the stack can evolve freely on top. Frequently asked questions Can a data lakehouse run on premises? Yes. Open table formats and query engines run on S3-compatible object storage in your own data center, the same way they run in public clouds. What table format works best on premises? Choose a format with broad support across the engines your teams use. From the storage side, the leading formats have almost identical requirements. What does object storage need for a lakehouse? S3 API compatibility, strong read-after-write consistency, high aggregate throughput, good small-object performance and scale. Why choose on-premises over cloud for a lakehouse? For data sovereignty, regulatory requirements, cost predictability for heavy workloads, proximity to data sources and portability. Do I still need Hadoop for an on-prem data lake? No. Many organizations move from HDFS to object storage with modern engines, separating compute and storage. Further reading See open table formats on on-prem S3, SQL query engine object storage performance, HDFS to object storage migration, sovereign data lakes and object storage for data lakes.