Download the free Kindle app and start reading Kindle books instantly on your smartphone, tablet, or computer - no Kindle device required.
Read instantly on your browser with Kindle for Web.
Using your mobile phone camera - scan the code below and download the Kindle app.
Follow the author
OK
Data Engineering Design Patterns: Recipes for Solving the Most Common Data Engineering Problems
Purchase options and add-ons
Data projects are an intrinsic part of an organization's technical ecosystem, but data engineers in many companies continue to work on problems that others have already solved. This hands-on guide shows you how to provide valuable data by focusing on various aspects of data engineering, including data ingestion, data quality, idempotency, and more.
Author Bartosz Konieczny guides you through the process of building reliable end-to-end data engineering projects, from data ingestion to data observability, focusing on data engineering design patterns that solve common business problems in a secure and storage-optimized manner. Each pattern includes a user-facing description of the problem, solutions, and consequences that place the pattern into the context of real-life scenarios.
Throughout this journey, you'll use open source data tools and public cloud services to apply each pattern. You'll learn:
- Challenges data engineers face and their impact on data systems
- How these challenges relate to data system components
- Useful applications of data engineering patterns
- How to identify and fix issues with your current data components
- Technology-agnostic solutions to new and existing data projects, with open source implementation examples
Bartosz Konieczny is a freelance data engineer who's been coding since 2010. He's held various senior hands-on positions that allowed him to work on many data engineering problems in batch and stream processing.
- ISBN-101098165810
- ISBN-13978-1098165819
- Edition1st
- PublisherO'Reilly Media
- Publication dateMay 20, 2025
- LanguageEnglish
- Dimensions7 x 1 x 9.25 inches
- Print length372 pages
Frequently bought together

Customers who viewed this item also viewed
- Fundamentals of Data Engineering: Plan and Build Robust Data SystemsPaperbackGet it as soon as Tuesday, Sep 29
- Deciphering Data Architectures: Choosing Between a Modern Data Warehouse, Data Fabric, Data Lakehouse, and Data MeshPaperbackFREE Shipping by AmazonGet it as soon as Sunday, Sep 20Only 13 left in stock (more on the way).
- Data Pipelines Pocket Reference: Moving and Processing Data for AnalyticsPaperbackFREE Shipping on orders over $35 shipped by AmazonGet it as soon as Sunday, Sep 20
- Designing Data-Intensive Applications: The Big Ideas Behind Reliable, Scalable, and Maintainable SystemsPaperbackFREE Shipping by AmazonGet it as soon as Sunday, Sep 20
- The Data Warehouse Toolkit: The Definitive Guide to Dimensional ModelingPaperbackFREE Shipping by AmazonGet it as soon as Sunday, Sep 20
- Database Internals: A Deep Dive into How Distributed Data Systems WorkPaperbackFREE Shipping by AmazonGet it as soon as Sunday, Sep 20
Customers also bought or read
- Deciphering Data Architectures: Choosing Between a Modern Data Warehouse, Data Fabric, Data Lakehouse, and Data Mesh
Paperback$50.99$50.99FREE delivery Sun, Sep 20 - Fundamentals of Data Engineering: Plan and Build Robust Data Systems
Paperback$40.00$40.00FREE delivery Sep 29 - Oct 14 - Building Medallion Architectures: Designing with Delta Lake and Spark
Paperback$39.49$39.49FREE delivery Sun, Sep 20 - Data Pipelines Pocket Reference: Moving and Processing Data for Analytics
Paperback$16.93$16.93Delivery Sun, Sep 20 - The Data Warehouse Toolkit: The Definitive Guide to Dimensional Modeling
Paperback$38.44$38.44FREE delivery Sun, Sep 20 - Database Internals: A Deep Dive into How Distributed Data Systems Work
Paperback$36.33$36.33FREE delivery Sun, Sep 20 - Practical Lakehouse Architecture: Designing and Implementing Modern Data Platforms at Scale
Paperback$45.39$45.39FREE delivery Sun, Sep 20 - Data Quality Fundamentals: A Practitioner's Guide to Building Trustworthy Data Pipelines
Paperback$39.63$39.63FREE delivery Sun, Sep 20 - Databricks Certified Data Engineer Associate Study Guide: In-Depth Guidance and Practice
Paperback$58.60$58.60FREE delivery Mon, Sep 21 - Delta Lake: The Definitive Guide: Modern Data Lakehouse Architectures with Data Lakes
Paperback$49.00$49.00FREE delivery Sun, Sep 20 - Kafka: The Definitive Guide: Real-Time Data and Stream Processing at Scale
Paperback$43.99$43.99FREE delivery Sun, Sep 20 - Data Engineering with Databricks Cookbook: Build effective data and AI solutions using Apache Spark, Databricks, and Delta Lake
Paperback$35.99$35.99FREE delivery Mon, Sep 21 - Spark: The Definitive Guide: Big Data Processing Made Simple
Paperback$43.70$43.70FREE delivery Sep 25 - 30 - Data Modeling with Snowflake: A practical guide to accelerating Snowflake development using universal modeling techniques
Paperback$44.99$44.99FREE delivery Mon, Sep 21 - Data Management at Scale: Modern Data Architecture with Data Mesh and Data Fabric
Paperback$43.80$43.80FREE delivery Sun, Sep 20 - Data Governance: The Definitive Guide: People, Processes, and Tools to Operationalize Data Trustworthiness
Paperback$45.99$45.99FREE delivery Sun, Sep 20 - Designing Machine Learning Systems: An Iterative Process for Production-Ready Applications
Paperback$40.00$40.00FREE delivery Sun, Sep 20 - The Enterprise Data Catalog: Improve Data Discovery, Ensure Data Governance, and Enable Innovation
Paperback$41.00$41.00FREE delivery Sun, Sep 20 - Streaming Systems: The What, Where, When, and How of Large-Scale Data Processing
Paperback$59.80$59.80$3.99 delivery Sep 29 - Oct 2 - Managing Data as a Product: Design and build data-product-centered socio-technical architectures
Paperback$28.87$28.87Delivery Mon, Sep 21 - Ace the Data Engineering Interview: Questions and Answers for Python, SQL, Data Modeling and More
Paperback$19.99$19.99Delivery Mon, Sep 21 - Hands-On Large Language Models: Language Understanding and Generation
Paperback$37.68$37.68FREE delivery Mon, Sep 21 - Semantic Modeling for Data: Avoiding Pitfalls and Breaking Dilemmas
Paperback$49.99$49.99FREE delivery Sun, Sep 20
From the brand
-
Databases, data science & more
-
Data Science
-
Visit the Store
-
Data Visualization
-
Databases
-
Streaming
-
Sharing the knowledge of experts
O'Reilly's mission is to change the world by sharing the knowledge of innovators. For over 40 years, we've inspired companies and individuals to do new things (and do them better) by providing the skills and understanding that are necessary for success.
Our customers are hungry to build the innovations that propel the world forward. And we help them do just that.
From the Publisher
From the Preface
The Structure of This Book
This book follows the workflow of a classical data engineering project that starts with data ingestion and ends with day-to-day monitoring. The steps of the project correspond to the main chapters, so you can easily identify the stage to which each pattern in a chapter applies.
Additionally, each chapter has a two-level structure, with the levels being design pattern categories and the design patterns themselves. Why this two-level organization? First, a given data engineering problem can have at least two possible solutions, and it wouldn’t be possible to logically group them without having this first level of design pattern categories. Second, data engineering design patterns have their own names that sometimes sound mysterious, and design pattern categories provide extra application context that helps you know where to apply a particular pattern without requiring you to delve into details.
Finally, for each pattern, you’ll find the following subsections:
Problem: This subsection provides a real-world example of when you can use the pattern.
Solution: This subsection describes the pattern in more technical detail. Usually, it starts with a high-level explanation followed by the technical implementation model.
Consequences: Patterns have their trade-offs, and this section explains what you should look for before implementing them. Whenever possible, each gotcha is completed with a mitigation solution.
Examples: In this final part, you’ll find code snippets explaining how to use the pattern within the modern data engineering tools. Unfortunately, it’s not technically possible to share the pattern’s implementation in all existing data tools, so this book uses popular open source projects (Apache Spark, Apache Flink, Apache Airflow, PostgreSQL, and Delta Lake). Occasionally, the implementation extends the scope to the managed services in the public cloud. The code snippets are written in Python, SQL, and sometimes Scala or Java if the Python implementation is not available.
At the end of this book, you will find a table summarizing all the described patterns. Also, the book has a GitHub repository that includes a glossary of terms that should give you the definitions of the most frequently used acronyms in the book.
What Should I Know Prior to Reading This Book?
This book will not be a great resource if you have just started in data engineering and don’t have any commercial experience. In my opinion, a minimum of six months of commercial experience with data engineering should help you grasp the ideas more easily. Other than that, the minimum technical knowledge you’ll need to get the most out of this book is as follows:
- Familiarity with data engineering concepts, such as extract, transform, load (ETL), extract, load, transform (ELT), data warehousing, data ingestion, and data orchestration.
- Cloud awareness. Even though this book tends to favor open source technologies, there are places where cloud technology is more appropriate (e.g., data security). You don’t have to be a cloud expert, but you should at least be able to understand the basics, such as what a managed service is.
- Hands-on experience with data processing logic in Java, Scala, Python, or SQL. Ideally, you have already deployed this logic in production.
If you feel like there are gaps in your knowledge of the required topics, you should be able to easily fill in the gaps by reading 'Fundamentals of Data Engineering' by Joe Reis and Matt Housley (O’Reilly, 2022). The book provides a comprehensive overview of the data engineering space that will not only help you understand the content in this book but also better prepare you to deal with the challenges you will face in your day-to-day work.
Generative AI Design Patterns
|
Machine Learning Design Patterns
|
Data Engineering Design Patterns
|
Cloud Application Architecture Patterns
|
Design Patterns for Cloud Native Applications
|
Head First Design Patterns
|
|
|---|---|---|---|---|---|---|
|
Add to Cart
|
Add to Cart
|
Add to Cart
|
Add to Cart
|
Add to Cart
|
Add to Cart
|
|
| Customer Reviews |
4.6 out of 5 stars 26
|
4.6 out of 5 stars 417
|
4.0 out of 5 stars 20
|
4.4 out of 5 stars 34
|
4.2 out of 5 stars 52
|
4.8 out of 5 stars 1,427
|
| Design Patterns from O'Reilly Media | no data | no data | no data | no data | no data | no data |
Editorial Reviews
About the Author
Besides that, you can read his blog posts at waitingforcode.com, or improve your data engineering skills with one of his courses or training. Bartosz is also an occasional speaker at conferences and meetups, including Data+AI Summit, Big Data Technology Warsaw Summit, or NDC Porto.
Product details
- Publisher : O'Reilly Media
- Publication date : May 20, 2025
- Edition : 1st
- Language : English
- Print length : 372 pages
- ISBN-10 : 1098165810
- ISBN-13 : 978-1098165819
- Item Weight : 1.35 pounds
- Dimensions : 7 x 1 x 9.25 inches
- Best Sellers Rank: #114,718 in Books (See Top 100 in Books)
- #10 in Data Warehousing (Books)
- #25 in Data Modeling & Design (Books)
- #31 in Database Storage & Design
- Customer Reviews:
About the author

Bartosz Konieczny is a freelance data engineering enthusiast who has been coding since 2010. Throughout his career, he has leveraged major public cloud services and
open source technologies—like Apache Spark, Apache Kafka, Apache Airflow, and Delta Lake—to tackle various data engineering problems, including sessionization, data ingestion, data cleansing, ordered data processing, and data migration.
In addition to helping companies bring their data projects to life, Bartosz is deeply engaged with the data engineering community. He provides a comprehensive set of resources to support data engineers in their learning journey, including online and in-person training, data-related blog posts on waitingforcode, and conference talks at industry events like the Spark+AI Summit, the Data+AI Summit, and the Big Data Technology Warsaw Summit.


















