Monday, October 5, 2026
Home » How Do You Migrate from Hadoop HDFS to Object Storage?

How Do You Migrate from Hadoop HDFS to Object Storage?

HDFS to object storage migration is one of the most common data platform projects of recent years. Organizations that built Hadoop clusters a decade ago now face aging hardware, tightly coupled compute and storage, complex operations and, in some cases, changes in vendor support. Moving data from the Hadoop Distributed File System to S3-compatible object storage separates storage from compute, opens the door to modern engines and table formats and simplifies operations. It is also a large, delicate project, often involving petabytes of data and hundreds of jobs that assume HDFS behavior.

This article explains why organizations move from HDFS to object storage, the differences that matter, migration approaches and tools, how to handle tables, security and validation and how to phase the project. For the target architecture, see our hub on running a data lakehouse on on-prem object storage.

Why move off HDFS

  • Coupled compute and storage: HDFS stores data on the same nodes that run compute, so growing storage means adding compute, and vice versa.
  • Replication overhead: HDFS traditionally keeps three replicas of data, tripling raw capacity. Erasure coding in newer HDFS versions helps but adds complexity. Object storage typically uses erasure coding natively. See erasure coding vs replication.
  • Operational complexity: NameNode high availability, balancing, upgrades and ecosystem version management require specialized skills.
  • NameNode limits: HDFS metadata lives in NameNode memory, which constrains file counts and makes small files expensive.
  • Modern engines: distributed query and processing engines, often running on container platforms, work naturally with S3-compatible storage and open table formats.
  • Hardware refresh: aging Hadoop clusters are an opportunity to change architecture rather than replace like for like.

Key differences between HDFS and object storage

Renames

HDFS renames are fast, atomic metadata operations, and many Hadoop jobs rely on writing to a temporary directory and renaming it at the end. Object storage does not offer atomic directory renames; renaming means copying and deleting objects. Jobs must use committers designed for object storage, such as the S3A committers in Hadoop, or open table formats that avoid renames altogether.

Directory listings

HDFS lists directories quickly. Object storage listings can be slower for very large prefixes. Table formats that track files in metadata, rather than by listing, avoid this issue.

Permissions

HDFS uses POSIX-like permissions and often Ranger or Sentry policies. Object storage uses bucket and object policies and IAM-style identities. Access models must be mapped and redesigned.

Data locality

Hadoop moved compute to data on the same nodes. With object storage, compute reads data over the network, so network bandwidth between compute and storage becomes critical.

Migration approaches

Lift and shift with DistCp

Hadoop’s DistCp tool copies data in parallel from HDFS to S3-compatible storage through the S3A connector. It is the standard way to move large volumes and can be run incrementally to catch changes. Plan bandwidth and run copies in batches by dataset.

Table-by-table modernization

Rather than copying files blindly, migrate tables to a modern format. Legacy directory-based tables can be converted to an open table format, either by registering existing files in place or by creating new tables and copying data. Most open table formats provide procedures to create tables from existing data files. See open table formats on on-prem S3.

Hybrid period

Most migrations run HDFS and object storage side by side for a period. Jobs are moved to read from and write to object storage dataset by dataset, while HDFS remains available until all workloads have moved.

Moving workloads, not just data

Data is only half the migration. Each job, pipeline and query must be updated:

  • Paths: change hdfs:// paths to s3a:// or the equivalent for each engine.
  • Committers: enable object storage committers for processing jobs that write output.
  • Configuration: endpoints, path-style access, credentials and TLS for each engine.
  • Engine upgrades: many teams move from older batch engines to modern distributed SQL and processing engines at the same time.
  • Scheduling: update orchestration to point jobs at new locations.

Build an inventory of jobs and their dependencies early; it usually takes longer than moving the data.

Security and governance

Map HDFS users, groups and policies to object storage identities and bucket policies. Decide on bucket layout by domain or sensitivity, apply encryption and audit logging and integrate with existing identity providers. If Ranger or similar tools are used for fine-grained access, check how they integrate with the new engines and storage. The scality.com post on S3 access policies covers policy design.

Validating the migration

Checksums computed by HDFS are not directly comparable with object storage ETags, especially for multipart uploads. Validation approaches include:

  • File counts and sizes per dataset and partition.
  • Content hashes computed independently on source and target files for critical datasets.
  • Row counts and aggregates for tables, comparing query results on source and target.
  • Job output comparison during parallel runs.

Record validation results for audit, especially in regulated environments.

Handling small files and cold data

Hadoop estates often contain huge numbers of small files from streaming jobs, logs and poorly configured pipelines. Copying them as-is to object storage carries over the problem, with more requests and slower queries. A migration is a good moment to compact small files into larger ones, either during copy or as part of converting tables to an open table format. Cold data that has not been read for years deserves scrutiny too: some can move to a lower-cost tier or bucket with lifecycle rules, and some can be deleted under retention policy instead of migrated at all. Reducing what moves shortens the project and lowers the cost of the target platform.

Bandwidth and copy time

Plan copy time realistically. Moving 5 PB at a sustained 10 GB per second takes about six days of continuous transfer, and real-world rates are usually lower once production workloads, validation and retries are included. Many teams schedule copies for nights and weekends over several months, prioritizing hot datasets so that the most-used workloads can move early. Incremental DistCp runs catch up on changes before each cutover.

Skills and teams

Moving from Hadoop to a lakehouse changes skills as well as technology. Hadoop administrators become platform engineers running object storage, container platforms and modern engines. Data engineers adopt open table formats and new tools. Plan training early, and involve the operations team in the pilot so they are ready to run the new platform before the old one is retired.

Phasing the project

  • Assess: inventory data, tables, jobs, users and dependencies; identify cold data that may not need to move.
  • Build: deploy object storage, network, engines and catalog.
  • Pilot: migrate a non-critical dataset and its jobs end to end.
  • Migrate in waves: move datasets and jobs by domain, validating each wave.
  • Cut over: switch remaining workloads and freeze HDFS writes.
  • Decommission: retire HDFS after a defined safety period.

Sizing the target

Size object storage for migrated data plus growth, with erasure coding overhead instead of HDFS triple replication. Many organizations find they need significantly less raw capacity after moving from triple-replicated HDFS. Size network bandwidth for analytical scan rates, since compute no longer reads local disks. See what object storage performance a SQL query engine needs and storage capacity planning.

Regional and regulatory considerations

Many Hadoop estates belong to banks, telcos, government agencies and utilities with strict data location requirements. Migrating to on-premises object storage keeps data in-country while modernizing the platform, avoiding the regulatory review that moving to a foreign public cloud might trigger. See how regulated firms keep a data lake inside their own country.

Checklist: HDFS to object storage migration

  • Inventory data, tables, jobs, users and dependencies.
  • Decide which data to migrate, archive or delete.
  • Choose between lift-and-shift and table modernization, often both.
  • Enable object storage committers or move to an open table format.
  • Map permissions to bucket policies and identities.
  • Size network bandwidth between compute and storage.
  • Copy data with DistCp or table-level tools in waves.
  • Validate with counts, hashes and query comparisons.
  • Run HDFS and object storage in parallel during transition.
  • Decommission HDFS after a safety period.

Putting it together

HDFS to object storage migration separates storage from compute, reduces replication overhead and lets organizations adopt modern engines and table formats. The main challenges are not copying bytes but adapting workloads to object storage semantics, redesigning security and validating results. Inventory carefully, migrate in waves, use committers or open table formats to avoid rename problems, size the network properly and keep HDFS available until every workload has moved and been verified.

Frequently asked questions

Why migrate from HDFS to object storage?

To separate compute from storage, reduce replication overhead, simplify operations and adopt modern engines and table formats.

What tool copies data from HDFS to S3-compatible storage?

Hadoop DistCp with the S3A connector is the standard tool, run in parallel and incrementally.

Do Hadoop jobs work unchanged on object storage?

Often not. Paths, committers, configuration and sometimes engines must change, especially for jobs that rely on renames.

How do you validate an HDFS migration?

Compare file counts and sizes, compute independent content hashes for critical data and compare table row counts and query results.

Will I need less capacity after migrating?

Often yes, because object storage typically uses erasure coding instead of HDFS triple replication.

Further reading

See on-prem data lakehouse, open table formats on on-prem S3, SQL query engine object storage performance, sovereign data lakes and object storage for data lakes.