Download the free Kindle app and start reading Kindle books instantly on your smartphone, tablet, or computer - no Kindle device required.
Read instantly on your browser with Kindle for Web.
Using your mobile phone camera - scan the code below and download the Kindle app.
Architecting an Apache Iceberg Lakehouse: A scalable, open-source data platform
Purchase options and add-ons
Design an Apache Iceberg lakehouse from scratch!
The “lakehouse” data architecture is a powerful way to combine the flexibility of data lakes with the management features of data warehouses. The open source Apache Iceberg framework delivers the scalability, reliability, and performance you want from a lakehouse without the expense and vendor lock-in of platforms like Snowflake, BigQuery, and Redshift.
In Architecting an Apache Iceberg Data Lakehouse, data guru Alex Merced shows you:
• How to create a modular, scalable Iceberg lakehouse architecture
• Where Spark, Flink, Dremio, Polaris fit into your design
• Reliable batch and streaming ingestion pipelines
• Strategies for governance, security, and performance at scale
Apache Iceberg is an open source table format perfect for massive analytic datasets. Iceberg enables ACID transactions, schema evolution, and high-performance queries on data lakes using multiple compute engines like Spark, Trino, Flink, Presto, and Hive. An Iceberg data lakehouse enables fast, reliable analytics at scale while retaining the observability you need for compliance audits, governance, and provable data security.
Foreword by Tim Berglund. Afterword by Adi Polak.
About the technology
Apache Iceberg is an open data format that lets data lake files work like database tables. It helps turn a data lake into a more reliable and capable lakehouse.
About the book
Architecting an Apache Iceberg Lakehouse shows you how to design an open, scalable, and cost-effective lakehouse platform with Apache Iceberg. More than a set of blueprints, the book explains the reasoning behind the architecture. You’ll build a mini lakehouse by ingesting sales and marketing data from PostgreSQL into Iceberg tables with Apache Spark and then create interactive dashboards in Apache Superset. You’ll appreciate expert Alex Merced’s real-world insights about operating an Iceberg lakehouse.
What's inside
• Create a modular, scalable Iceberg lakehouse architecture
• Fit Spark, Flink, Dremio, Polaris and more into your design
• Batch and streaming ingestion pipelines
• Governance, security, and performance at scale
About the reader
For data architects familiar with the basics of a data lakehouse.
About the author
Alex Merced is Head of Developer Relations at Dremio. He shares his expertise through videos, podcasts, and articles, and leads the DataLakehouseHub.com community.
Table of Contents
Part 1
1 The world of the data lakehouse
2 Apache Iceberg and the lakehouse
3 Hands-on with Apache Iceberg
Part 2
4 Preparing for your move to Apache Iceberg
5 Selecting the storage layer
6 Architecting the ingestion layer
7 Implementing the catalog layer
8 Designing the federation layer
9 Understanding the consumption layer
Part 3 Operating your Apache Iceberg lakehouse
10 Maintaining an Iceberg lakehouse
11 Operationalizing Apache Iceberg
A The metadata tables
B Python for Apache Iceberg
C The Apache Iceberg specification
- ISBN-101633435105
- ISBN-13978-1633435100
- Publication dateMay 19, 2026
- LanguageEnglish
- Dimensions7.38 x 1.02 x 9.25 inches
- Print length408 pages
Frequently bought together

Customers who viewed this item also viewed
- Practical Lakehouse Architecture: Designing and Implementing Modern Data Platforms at ScalePaperbackFREE Shipping by AmazonGet it as soon as Sunday, Sep 20Only 8 left in stock (more on the way).
- The Claude Code Operating Model: Build scalable AI coding systems with Skills, MCP, Hooks, agent orchestration, and SDK patternsPaperbackFREE Shipping by AmazonGet it as soon as Monday, Sep 21
- Engineering Lakehouses with Open Table Formats: Build scalable and efficient lakehouses with Apache Iceberg, Apache Hudi, and Delta LakePaperbackFREE Shipping by AmazonGet it as soon as Sunday, Sep 20
- Building Medallion Architectures: Designing with Delta Lake and SparkPaperbackFREE Shipping by AmazonGet it as soon as Sunday, Sep 20Only 20 left in stock (more on the way).
- Designing Data-Intensive Applications: The Big Ideas Behind Reliable, Scalable, and Maintainable SystemsPaperbackFREE Shipping by AmazonGet it as soon as Sunday, Sep 20
- Building Claude Skills: Turn repeatable work into reusable Claude SkillsPaperbackFREE Shipping on orders over $35 shipped by AmazonGet it as soon as Monday, Sep 21
Customers also bought or read
- Kafka: The Definitive Guide: Real-Time Data and Stream Processing at Scale
Paperback$43.99$43.99FREE delivery Sun, Sep 20 - Data Pipelines with Apache Airflow, Second Edition: Orchestration for data and AI
Paperback$56.99$56.99FREE delivery Mon, Sep 21 - Apache Polaris: The Definitive Guide: Enriching Apache Iceberg Data Lakehouses with an Open Source Catalog
Paperback$70.87$70.87FREE delivery Sun, Sep 20 - Kubernetes in Action, Second Edition: Deploying and managing containers and cloud-native applications
Paperback$54.07$54.07FREE delivery Mon, Sep 21 - Terraform in Depth: Infrastructure as Code with Terraform and OpenTofu
Paperback$49.31$49.31FREE delivery Mon, Sep 21 - Think Distributed Systems: Mental models of reliable and scalable software
Paperback$47.99$47.99FREE delivery Sun, Sep 20 - Database Internals: A Deep Dive into How Distributed Data Systems Work
Paperback$36.33$36.33FREE delivery Sun, Sep 20 - Learning Go: An Idiomatic Approach to Real-World Go Programming
Paperback$43.64$43.64FREE delivery Sun, Sep 20
From the Publisher
“Gives you the practical grounding to build with confidence, and maybe even enjoy the process.”
Matt Topol Apache Iceberg PMC Member
“Building a lakehouse without this book is like building a house without a foundation.”
Roy Hasson, Microsoft
“Th e author’s passion and competence shine through in every chapter of this book.”
Joe Reis, co-author of Fundamentals of Data Engineerin
about the book
Architecting an Apache Iceberg Lakehouse shows you how to design a complete lakehouse architecture around Apache Iceberg, so you can make solid decisions about storage, ingestion, querying, and governance instead of just stitching tools together.
You’ll get hands-on experience building and running an Iceberg-based platform, which helps you move from theory to a working system you can adapt for production.
By the end, you’ll understand the trade-offs behind real-world data architectures and how to build a scalable, maintainable platform that supports reliable analytics at large scale.
about the authors
Manning helps developers and tech professionals stay ahead in a fast-moving industry with expert-led books, videos, and projects. Learning never stops, but it’s hard to keep up, so we focus on content that’s practical, clear, and trusted. As an independent publisher, we adapt quickly, from pioneering early-access books to offering DRM-free eBooks. Our series, like "In Action" and "In a Month of Lunches", reflect a commitment to making complex topics accessible.
Apache Kafka in Action: From basics to production
|
Data Pipelines with Apache Airflow, Second Edition: Orchestration for data an...
|
Data Pipelines with Apache Airflow
|
Kafka Streams in Action, Second Edition: Event-driven applications and micros...
|
Grokking Data Structures
|
Grokking Algorithms, Second Edition
|
|
|---|---|---|---|---|---|---|
|
Add to Cart
|
Add to Cart
|
Add to Cart
|
Add to Cart
|
Add to Cart
|
Add to Cart
|
|
| Customer Reviews |
4.7 out of 5 stars 5
|
4.6 out of 5 stars 6
|
4.4 out of 5 stars 77
|
3.8 out of 5 stars 7
|
4.5 out of 5 stars 18
|
4.7 out of 5 stars 282
|
| User experience level | Intermediate | Intermediate | Intermediate | Intermediate | Beginner | Beginner |
| About the reader | For IT operators, software architects and developers. | For data engineers, machine learning engineers, DevOps, and sysadmins with intermediate Python skills. | For DevOps, data engineers, machine learning engineers, and sysadmins | For Java developers. | For readers who know the basics of Python. | No advanced math or programming skills required. |
| Special features | Includes liveBook with out built-in AI assistant. | Includes liveBook with out built-in AI assistant. | Includes liveBook with out built-in AI assistant. | Includes liveBook with out built-in AI assistant. | Includes liveBook with out built-in AI assistant. | Includes liveBook with out built-in AI assistant. |
| Page count | 368 | 512 | 480 | 504 | 280 | 320 |
Editorial Reviews
About the Author
Product details
- Publisher : Manning Publications
- Publication date : May 19, 2026
- Language : English
- Print length : 408 pages
- ISBN-10 : 1633435105
- ISBN-13 : 978-1633435100
- Item Weight : 1.62 pounds
- Dimensions : 7.38 x 1.02 x 9.25 inches
- Best Sellers Rank: #1,403,146 in Books (See Top 100 in Books)
- #166 in Data Warehousing (Books)
- #313 in Data Modeling & Design (Books)
- #897 in Database Storage & Design
- Customer Reviews:

















