Phase 2: Data Storage

Delta Lake, Iceberg & Hudi table formats

Intermediate ~3 min read
Think of it this way A friendly analogy. Read this if the technical version feels dense. Show Hide

Imagine you have a gigantic playroom, bigger than any store, full of all sorts of LEGO bricks. This playroom is like a "Data Lake" – a huge place where you can store tons and tons of information, just like keeping all your LEGO bricks. It’s amazing because you can keep everything there! But here's the tricky part: if you just dump all your bricks in random piles, it becomes really hard to build big, complicated things like a super detailed castle or a huge city. If someone accidentally knocks down a small piece of your castle, how do you fix just that part without the whole thing falling apart? Or what if you want to know exactly what your castle looked like last week? It's chaos!

This is where our special "building managers" come in: Delta Lake, Iceberg, and Hudi. Think of them not as new bricks, but as clever systems that help you manage your existing LEGO bricks and builds. They're like adding a special rulebook and a magical undo button to your playroom. They help make sure everyone builds neatly (like making sure all your red bricks are used for walls and not roofs by mistake), and they keep track of every change. So, if your little brother accidentally takes a brick from your castle, these managers can actually put it back exactly how it was, or even show you how your castle looked before he touched it. They create an organized "instruction manual" for your builds, even when they're made of millions of bricks.

So, what does this mean you can do? With these building managers, you can make sure your giant LEGO castle is always strong and correct. If you need to add a new tower, the manager makes sure you don't accidentally mess up the existing ones. If a piece breaks, you can fix just that piece without rebuilding the whole thing. You can also "time travel" with your castle – going back to see how it was built last Tuesday, or even the day before! Delta Lake, Iceberg, and Hudi are just three different kinds of these "building managers," each with their own clever ways of helping you manage your enormous LEGO constructions. This means you can build incredible, sturdy, and always-correct structures, even with billions of bricks, letting you focus on the fun of building rather than the mess.

As a Data Engineer working with Data Lakes, you'll quickly encounter the need for more than just raw file storage. While Data Lakes offer vast scalability and cost efficiency, they traditionally lack key features found in data warehouses, such as ACID (Atomicity, Consistency, Isolation, Durability) transactions, schema enforcement, and efficient data modification. This gap is precisely what modern table formats like Delta Lake, Apache Iceberg, and Apache Hudi aim to bridge. They operate as a transactional layer on top of your data lake storage (e.g., S3, ADLS) and underlying file formats (like Parquet or ORC), bringing data warehouse-like capabilities to your lakehouse architecture.

Each of these formats approaches this problem with slightly different strengths and design philosophies. Delta Lake, developed by Databricks, is highly integrated with Apache Spark and offers robust ACID transactions, schema evolution and enforcement, time travel capabilities (to query historical versions of data), and unified batch and streaming data processing. It's often favored in Spark-centric environments for general-purpose data warehousing on a lake. Apache Iceberg, originating from Netflix, focuses on correctness and performance for massive tables, offering capabilities like hidden partitioning, schema evolution without rewriting data, and powerful metadata management that allows multiple compute engines (Spark, Flink, Presto, Trino) to safely work with the same tables. Its design emphasizes reliable table semantics and efficient query planning.

Apache Hudi (short for Hadoop Upserts Deletes and Incrementals), developed by Uber, excels in scenarios requiring very efficient upserts and deletes, making it ideal for real-time data ingestion, Change Data Capture (CDC) workloads, and managing frequently changing datasets. Hudi offers different storage types, such as Copy-on-Write (optimized for read performance) and Merge-on-Read (optimized for write performance and low-latency ingestion). While all three aim to provide transactional guarantees and enhance data lake functionality, your choice often comes down to your primary compute engine, workload patterns (e.g., real-time updates vs. large batch analytics), and desired ecosystem integration.

Key Takeaways

  • Table formats (Delta, Iceberg, Hudi) add ACID transactions and data warehouse features to Data Lakes.
  • They sit on top of object storage and file formats like Parquet/ORC, managing metadata for transactional capabilities.
  • Delta Lake excels with Spark, offering unified batch/streaming and schema enforcement.
  • Apache Iceberg focuses on large tables, multi-engine support, and robust schema/partition evolution.
  • Apache Hudi is optimized for efficient upserts/deletes, ideal for real-time and CDC workloads.

Code Example

python
from pyspark.sql import SparkSession
from pyspark.sql.functions import lit

spark = SparkSession.builder \
    .appName("DeltaLakeWriteRead") \
    .config("spark.sql.extensions", "io.delta.sql.DeltaSparkSessionExtension") \
    .config("spark.sql.catalog.spark_catalog", "org.apache.spark.sql.delta.catalog.DeltaCatalog") \
    .getOrCreate()

# Create a simple DataFrame
data = [("Laptop", 1200), ("Keyboard", 75), ("Mouse", 25)]
cols = ["product", "price"]
df = spark.createDataFrame(data, cols)

# Define the path for the Delta table
delta_table_path = "/tmp/products_delta_table"

# Write the DataFrame as a Delta Lake table
df.write.format("delta").mode("overwrite").save(delta_table_path)
print("Data written to Delta Lake successfully.")

# Read data from the Delta Lake table
read_df = spark.read.format("delta").load(delta_table_path)
print("Data read from Delta Lake:")
read_df.show()

# Perform a simple update (append a new row)
new_product_df = spark.createDataFrame([("Monitor", 300)], cols)
new_product_df.write.format("delta").mode("append").save(delta_table_path)

# Read again to see the updated table
print("Data after append:")
spark.read.format("delta").load(delta_table_path).show()

spark.stop()

How this code works

This code demonstrates how to create, write, read, and update a Delta Lake table using PySpark, illustrating core capabilities of these modern table formats within a Data Lakehouse. It starts by configuring a SparkSession specifically for Delta Lake, enabling Spark to understand and interact with Delta tables through extensions like io.delta.sql.DeltaSparkSessionExtension and org.apache.spark.sql.delta.catalog.DeltaCatalog. A simple DataFrame is then created with product data.

The initial data is written to a specified delta_table_path using df.write.format("delta").mode("overwrite").save(). The format("delta") explicitly tells Spark to use the Delta Lake format, while mode("overwrite") ensures a fresh table, replacing any existing data at that path – a crucial choice for initial setup or ensuring idempotence. The table is then read back using spark.read.format("delta").load() to verify the write. Next, a new product DataFrame is created and appended to the existing Delta table using mode("append"), showcasing how Delta Lake allows efficient, transactional updates to add new records without rewriting the entire dataset. A final read confirms the successful append operation.