Buy New
-
To see product details, add this item to your cart.
Ships from: Amazon.com Sold by: Amazon.com
Used - Good
-
To see product details, add this item to your cart.
Ships from: World of Books (previously glenthebookseller) Sold by: World of Books (previously glenthebookseller)
Download the free Kindle app and start reading Kindle books instantly on your smartphone, tablet, or computer - no Kindle device required.
Read instantly on your browser with Kindle for Web.
Using your mobile phone camera - scan the code below and download the Kindle app.
Follow the author
OK
Data Pipelines Pocket Reference: Moving and Processing Data for Analytics
Purchase options and add-ons
Data pipelines are the foundation for success in data analytics. Moving data from numerous diverse sources and transforming it to provide context is the difference between having data and actually gaining value from it. This pocket reference defines data pipelines and explains how they work in today's modern data stack.
You'll learn common considerations and key decision points when implementing pipelines, such as batch versus streaming data ingestion and build versus buy. This book addresses the most common decisions made by data professionals and discusses foundational concepts that apply to open source frameworks, commercial products, and homegrown solutions.
You'll learn:
- What a data pipeline is and how it works
- How data is moved and processed on modern data infrastructure, including cloud platforms
- Common tools and products used by data engineers to build pipelines
- How pipelines support analytics and reporting needs
- Considerations for pipeline maintenance, testing, and alerting
- ISBN-101492087831
- ISBN-13978-1492087830
- Edition1st
- PublisherO'Reilly Media
- Publication dateMarch 16, 2021
- LanguageEnglish
- Dimensions4 x 0.75 x 7 inches
- Print length274 pages
Frequently bought together

Customers who viewed this item also viewed
- Fundamentals of Data Engineering: Plan and Build Robust Data SystemsPaperbackGet it as soon as Sunday, Sep 27
- The Data Warehouse Toolkit: The Definitive Guide to Dimensional ModelingPaperbackFREE Shipping by AmazonGet it as soon as Saturday, Sep 19
- Designing Data-Intensive Applications: The Big Ideas Behind Reliable, Scalable, and Maintainable SystemsPaperbackFREE Shipping by AmazonGet it as soon as Saturday, Sep 19
- Data Engineering Design Patterns: Recipes for Solving the Most Common Data Engineering ProblemsPaperbackFREE Shipping by AmazonGet it as soon as Saturday, Sep 19Only 3 left in stock (more on the way).
- AI Engineering: Building Applications with Foundation ModelsPaperbackFREE Shipping by AmazonGet it as soon as Saturday, Sep 19
Customers also bought or read
- Fundamentals of Data Engineering: Plan and Build Robust Data Systems
Paperback$40.00$40.00FREE delivery Sep 27 - Oct 14 - The Data Warehouse Toolkit: The Definitive Guide to Dimensional Modeling
Paperback$38.44$38.44FREE delivery Sat, Sep 19 - Deciphering Data Architectures: Choosing Between a Modern Data Warehouse, Data Fabric, Data Lakehouse, and Data Mesh
Paperback$50.99$50.99FREE delivery Sat, Sep 19 - Data Engineering Design Patterns: Recipes for Solving the Most Common Data Engineering Problems
Paperback$35.46$35.46FREE delivery Sat, Sep 19 - Data Quality Fundamentals: A Practitioner's Guide to Building Trustworthy Data Pipelines
Paperback$39.63$39.63FREE delivery Sat, Sep 19 - Ace the Data Engineering Interview: Questions and Answers for Python, SQL, Data Modeling and More
Paperback$19.99$19.99Delivery Mon, Sep 21 - Data Engineering with AWS: Acquire the skills to design and build AWS-based data transformation pipelines like a pro
Paperback$29.41$29.41Delivery Mon, Sep 21 - Data Governance: The Definitive Guide: People, Processes, and Tools to Operationalize Data Trustworthiness
Paperback$45.99$45.99FREE delivery Sat, Sep 19 - SQL for Data Analysis: Advanced Techniques for Transforming Data into Insights
Paperback$36.49$36.49FREE delivery Sat, Sep 19 - Machine Learning Pocket Reference: Working with Structured Data in Python
Paperback$16.49$16.49Delivery Sat, Sep 19 - Database Internals: A Deep Dive into How Distributed Data Systems Work
Paperback$36.33$36.33FREE delivery Sat, Sep 19 - Building Medallion Architectures: Designing with Delta Lake and Spark
Paperback$39.49$39.49FREE delivery Sat, Sep 19 - Python Pocket Reference: Python In Your Pocket (Pocket Reference (O'Reilly))
Paperback$13.09$13.09Delivery Sat, Sep 19 - Data Engineering with Databricks Cookbook: Build effective data and AI solutions using Apache Spark, Databricks, and Delta Lake
Paperback$35.99$35.99FREE delivery Mon, Sep 21 - Practical Lakehouse Architecture: Designing and Implementing Modern Data Platforms at Scale
Paperback$45.39$45.39FREE delivery Sat, Sep 19 - Data Engineering with Python: Work with massive datasets to design data models and automate data pipelines using Python
Paperback$41.99$41.99FREE delivery Mon, Sep 21 - The Data Warehouse ETL Toolkit: Practical Techniques for Extracting, Cleaning, Conforming, and Delivering Data
Paperback$22.37$22.37Delivery Sat, Sep 19 - Spark: The Definitive Guide: Big Data Processing Made Simple
Paperback$43.70$43.70FREE delivery Sep 25 - 30 - Data Science from Scratch: First Principles with Python
Paperback$35.54$35.54FREE delivery Sat, Sep 19 - Designing Machine Learning Systems: An Iterative Process for Production-Ready Applications
Paperback$40.00$40.00FREE delivery Sat, Sep 19 - Essential Math for Data Science: Take Control of Your Data with Fundamental Linear Algebra, Probability, and Statistics
Paperback$36.22$36.22FREE delivery Sat, Sep 19 - Kafka: The Definitive Guide: Real-Time Data and Stream Processing at Scale
Paperback$43.99$43.99FREE delivery Sat, Sep 19 - Managing Data as a Product: Design and build data-product-centered socio-technical architectures
Paperback$28.87$28.87Delivery Mon, Sep 21 - Python for Data Analysis: Data Wrangling with pandas, NumPy, and Jupyter
Paperback$43.99$43.99FREE delivery Sep 30 - Oct 16
From the brand
-
Databases, data science & more
-
Data Science
-
Visit the Store
-
Data Visualization
-
Databases
-
Streaming
-
Sharing the knowledge of experts
O'Reilly's mission is to change the world by sharing the knowledge of innovators. For over 40 years, we've inspired companies and individuals to do new things (and do them better) by providing the skills and understanding that are necessary for success.
Our customers are hungry to build the innovations that propel the world forward. And we help them do just that.
From the Publisher
From the Preface
Data pipelines are the foundation for success in data analytics and machine learning. Moving data from numerous, diverse sources and processing it to provide context is the difference between having data and getting value from it.
I’ve worked as a data analyst, data engineer, and leader in the data analytics field for more than 10 years. In that time, I’ve seen rapid change and growth in the field. The emergence of cloud infrastructure, and cloud data warehouses in particular, has created an opportunity to rethink the way data pipelines are designed and implemented.
This book describes what I believe are the foundations and best practices of building data pipelines in the modern era. I base my opinions and observations on my own experience as well as those of industry leaders who I know and follow.
My goal is for this book to serve as a blueprint as well as a reference. While your needs are specific to your organization and the problems you’ve set out to solve, I’ve found success with variations of these foundations many times over. I hope you find it a valuable resource in your journey to building and maintaining data pipelines that power your data organization.
Who This Book Is For
This book’s primary audience is current and aspiring data engineers as well as analytics team members who want to understand what data pipelines are and how they are implemented. Their job titles include data engineers, technical leads, data warehouse engineers, analytics engineers, business intelligence engineers, and director/VP-level analytics leaders.
I assume that you have a basic understanding of data warehousing concepts. To implement the examples discussed, you should be comfortable with SQL databases, REST APIs, and JSON. You should be proficient in a scripting language, such as Python. Basic knowledge of the Linux command line and at least one cloud computing platform is ideal as well.
All code samples are written in Python and SQL and make use of many open source libraries. I use Amazon Web Services (AWS) to demonstrate the techniques described in the book, and AWS services are used in many of the code samples. When possible, I note similar services on other major cloud providers such as Microsoft Azure and Google Cloud Platform (GCP). All code samples can be modified for the cloud provider of your choice, as well as for on-premises use.
Editorial Reviews
About the Author
Product details
- Publisher : O'Reilly Media
- Publication date : March 16, 2021
- Edition : 1st
- Language : English
- Print length : 274 pages
- ISBN-10 : 1492087831
- ISBN-13 : 978-1492087830
- Item Weight : 2.31 pounds
- Dimensions : 4 x 0.75 x 7 inches
- Best Sellers Rank: #217,141 in Books (See Top 100 in Books)
- #24 in Data Warehousing (Books)
- #56 in Data Modeling & Design (Books)
- #66 in Database Storage & Design
- Customer Reviews:
About the author

James Densmore is a Director of Engineering at HubSpot as well as the Founder and Principal Consultant at Data Liftoff. He has more than 10 years of experience leading data teams and building data infrastructure at Wayfair, O'Reilly Media, HubSpot, and Degreed. James has a BS in Computer Science from Northeastern University and an MBA from Boston College.


















