Monday, October 5, 2026
Home » What Is a Security Data Lake and Where Should It Live?

What Is a Security Data Lake and Where Should It Live?

A security data lake is a central repository that stores large volumes of security telemetry, such as logs, endpoint events, network flows and cloud audit trails, in open formats on scalable storage so that many tools can query it. It has emerged because traditional SIEM platforms struggle with the cost of ingesting and retaining everything security teams now collect. Instead of forcing every log through the SIEM and keeping it there for years, organizations keep the full history in the lake and use the SIEM for real-time detection on the most important data.

This article explains what a security data lake is, how it relates to SIEM, the formats and architecture it typically uses and the factors that decide where it should live: in a public cloud, on premises or both. For log retention requirements that often drive the decision, see how long security logs should be retained.

Why security data lakes exist

Security telemetry has grown much faster than security budgets. Endpoint detection and response, cloud audit logs, identity providers, SaaS applications and network sensors all produce large volumes of data. SIEM platforms are often priced by ingest volume or compute, and keeping a year or more of high-volume data searchable can be very expensive.

The result is a familiar compromise: teams drop data sources, shorten retention or filter aggressively, and then lack the data they need during investigations. A security data lake separates storage from analytics so that data can be kept affordably and queried when needed.

Security data lake vs SIEM

A SIEM and a security data lake do different jobs:

  • SIEM focuses on real-time correlation, detection rules, alerting, case management and fast search of recent data.
  • Security data lake focuses on cost-efficient long-term storage of large volumes of telemetry, flexible querying, threat hunting across long periods and feeding analytics or machine learning.

Many organizations use both: high-value data flows into the SIEM for detection, while all data, including the high-value sources, lands in the lake for retention and hunting. Some SIEM platforms now query lake data directly. See what is SIEM for background, and data lake vs data warehouse for the general distinction.

How a security data lake is built

Storage layer

At the bottom is object storage, which provides low cost per terabyte, scale to petabytes and the S3 API that analytics engines expect. Object storage also supports immutability, which matters for security evidence. See object storage for your data lake.

Open file and table formats

Data is typically stored in columnar formats such as Apache Parquet, which compress well and allow queries to read only the columns they need. Open table formats add features like schema evolution, partitioning and time travel on top of files in object storage, making the lake behave more like a database while keeping data in open formats.

Common schema

Security data comes from many sources with different field names. Normalizing to a common schema makes cross-source queries possible. The Open Cybersecurity Schema Framework (OCSF), an open-source project backed by a number of security vendors, is one widely discussed option, and some organizations use schemas defined by their SIEM or their own standards.

Ingestion pipeline

Pipelines collect data from sources, parse and normalize it, enrich it with context such as asset and identity information and write it to the lake in the chosen formats. Pipelines may also route subsets of data to the SIEM.

Query and analytics layer

SQL engines, notebooks, threat hunting tools and machine learning platforms query the lake. Because the data is in open formats, multiple engines can read the same data without copying it.

What a security data lake enables

  • Longer retention at lower cost, supporting compliance and investigations into long-running intrusions.
  • Threat hunting across months or years of data.
  • Incident investigation that reaches back to the initial compromise.
  • Detection engineering, testing new detection logic against historical data.
  • Analytics and machine learning on large datasets.
  • Tool flexibility, avoiding lock-in to a single analytics vendor.

Where should a security data lake live?

Public cloud

Cloud object storage and managed query services make it quick to start, and they suit organizations whose telemetry is mostly cloud-native. Considerations include ongoing storage cost as data grows, query and data transfer charges, data residency and the effort of moving large volumes later. See how to avoid cloud lock-in.

On premises

On-premises object storage keeps security data within the organization’s own facilities, which matters for regulated industries, government and organizations with strict sovereignty requirements. Costs are predictable, and there are no per-query or egress charges when analysts run large hunts. It requires capacity planning and operations. The general design principles are covered in big data analytics and object storage.

Hybrid

Many organizations combine both: cloud telemetry lands in a cloud lake, on-premises telemetry in an on-premises lake, or one acts as the long-term archive for the other. Open formats make hybrid designs workable because data can be queried in place or moved without conversion.

Factors that decide placement

  • Where most telemetry originates and the cost of moving it.
  • Data residency and sovereignty requirements.
  • Retention period and volume, which drive long-term storage cost.
  • Query patterns: frequent large hunts favor storage without per-query fees.
  • Existing skills and infrastructure.
  • Security of the lake itself, including who can access it and how it is isolated.

Securing the security data lake

A lake full of security telemetry is a high-value target. It contains information about the organization’s infrastructure, users and defenses. Protections include:

  • Strict access control with least privilege and separation from production administration.
  • Immutability for raw data so attackers cannot erase evidence. See S3 object lock: immutability and WORM.
  • Encryption at rest and in transit, with careful key management. See encryption key management.
  • Audit logging of access to the lake itself.
  • Data minimization and retention limits for personal data.

Relationship to SIEM storage tiers

Many SIEM platforms can already place warm data on object storage and keep only a cache on their indexers. A security data lake can complement this, holding raw or normalized data in open formats for other tools and longer retention. Some organizations use the same object storage platform for both, with separate buckets and policies. See security log retention.

Cost drivers to model

The economics of a security data lake depend on a few variables. Storage capacity grows with ingest and retention, though columnar compression often reduces volume substantially compared with raw logs. Compute for queries is used in bursts, during hunts and investigations, rather than continuously. Pipeline processing adds cost proportional to ingest. In public cloud, request charges and data transfer can be significant for large scans; on premises, the main costs are hardware, power and operations, spread over the platform’s life.

A useful comparison is the cost per terabyte retained per year in the lake versus in the SIEM, for the same data. That number usually makes the case for moving high-volume, lower-value sources to the lake while keeping detection-critical sources in the SIEM.

Getting started

A pragmatic path:

  • Identify high-volume sources that are expensive to keep in the SIEM.
  • Define retention and schema requirements.
  • Stand up object storage and an open table format.
  • Build pipelines for a few sources and validate queries.
  • Route detection-critical data to the SIEM and everything to the lake.
  • Expand sources, analytics and retention over time.

Checklist: planning a security data lake

  • Define the jobs the lake will do alongside the SIEM.
  • Choose open formats and a common schema.
  • Select object storage that scales and supports immutability.
  • Decide on cloud, on-premises or hybrid placement using cost, residency and query patterns.
  • Build ingestion pipelines with normalization and enrichment.
  • Secure the lake with access control, encryption and audit logging.
  • Set retention by data type and automate deletion.
  • Test hunts and investigations over long time ranges.

Putting it together

A security data lake gives security teams affordable, long-term access to all their telemetry in open formats, complementing a SIEM focused on real-time detection. It lives on object storage, either in the cloud, on premises or both, and the right placement depends on where data originates, residency requirements, retention volume and how often analysts query history. Build it with open formats, protect it as carefully as any production system and let it carry the retention load the SIEM cannot.

Frequently asked questions

What is a security data lake?

A central repository of security telemetry stored in open formats on scalable storage, queried by multiple tools for hunting, investigation and analytics.

Does a security data lake replace a SIEM?

Usually not. Most organizations use the SIEM for real-time detection and the lake for long-term retention and broad analysis.

What formats do security data lakes use?

Commonly columnar file formats such as Parquet, open table formats and a common schema such as OCSF.

Should a security data lake be in the cloud or on premises?

It depends on where telemetry originates, residency requirements, data volumes and query patterns. Many organizations use a hybrid approach.

How do you protect a security data lake?

With strict access control, immutability for raw data, encryption, audit logging and data minimization.

Further reading

See security log retention and data lake vs data warehouse.