Buy New
-
To see product details, add this item to your cart.
Ships from: Amazon.com Sold by: Amazon.com
Used - Good
-
To see product details, add this item to your cart.
FREE Returns
Return this item for free
We offer easy, convenient returns with at least one free return option: no shipping charges. All returns must comply with our returns policy.
Learn more about free returns. How to return the item? - Go to your orders and start the return
- Select your preferred free shipping option
- Drop off and leave!
Ships from: Amazon Sold by: L&O Goods
Return this item for free
We offer easy, convenient returns with at least one free return option: no shipping charges. All returns must comply with our returns policy.
Learn more about free returns.- Go to your orders and start the return
- Select your preferred free shipping option
- Drop off and leave!
Download the free Kindle app and start reading Kindle books instantly on your smartphone, tablet, or computer - no Kindle device required.
Read instantly on your browser with Kindle for Web.
Using your mobile phone camera - scan the code below and download the Kindle app.
Follow the authors
OK
Feature Engineering for Machine Learning: Principles and Techniques for Data Scientists
Purchase options and add-ons
Feature engineering is a crucial step in the machine-learning pipeline, yet this topic is rarely examined on its own. With this practical book, you’ll learn techniques for extracting and transforming features―the numeric representations of raw data―into formats for machine-learning models. Each chapter guides you through a single data problem, such as how to represent text or image data. Together, these examples illustrate the main principles of feature engineering.
Rather than simply teach these principles, authors Alice Zheng and Amanda Casari focus on practical application with exercises throughout the book. The closing chapter brings everything together by tackling a real-world, structured dataset with several feature-engineering techniques. Python packages including numpy, Pandas, Scikit-learn, and Matplotlib are used in code examples.
You’ll examine:
- Feature engineering for numeric data: filtering, binning, scaling, log transforms, and power transforms
- Natural text techniques: bag-of-words, n-grams, and phrase detection
- Frequency-based filtering and feature scaling for eliminating uninformative features
- Encoding techniques of categorical variables, including feature hashing and bin-counting
- Model-based feature engineering with principal component analysis
- The concept of model stacking, using k-means as a featurization technique
- Image feature extraction with manual and deep-learning techniques
- ISBN-101491953241
- ISBN-13978-1491953242
- Edition1st
- PublisherO'Reilly Media
- Publication dateMay 8, 2018
- LanguageEnglish
- Dimensions7 x 0.5 x 9 inches
- Print length215 pages
Frequently bought together

Customers who viewed this item also viewed
- Designing Machine Learning Systems: An Iterative Process for Production-Ready ApplicationsPaperbackFREE Shipping by AmazonGet it as soon as Sunday, Sep 20
- Feature Engineering and Selection: A Practical Approach for Predictive Models (Chapman & Hall/CRC Data Science Series)HardcoverFREE Shipping by AmazonGet it as soon as Sunday, Sep 20Only 11 left in stock - order soon.
- Hands-On Machine Learning with Scikit-Learn and PyTorch: Concepts, Tools, and Techniques to Build Intelligent SystemsPaperbackFREE Shipping by AmazonGet it as soon as Monday, Sep 21
- Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent SystemsPaperbackFREE Shipping by AmazonGet it as soon as Monday, Sep 21
- AI Engineering: Building Applications with Foundation ModelsPaperbackFREE Shipping by AmazonGet it as soon as Sunday, Sep 20
- AI Systems Performance Engineering: Optimizing Model Training and Inference Workloads with GPUs, CUDA, and PyTorchPaperbackFREE Shipping by AmazonGet it as soon as Monday, Sep 21
Customers also bought or read
- Designing Machine Learning Systems: An Iterative Process for Production-Ready Applications
Paperback$40.00$40.00FREE delivery Sun, Sep 20 - Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems
Paperback$49.50$49.50FREE delivery Mon, Sep 21 - Feature Engineering and Selection (Chapman & Hall/CRC Data Science Series)
Paperback$44.79$44.79FREE delivery Mon, Sep 21 - The Hundred-Page Machine Learning Book (The Hundred-Page Books)
Paperback$26.14$26.14Delivery Mon, Sep 21 - Practical Statistics for Data Scientists: 50+ Essential Concepts Using R and Python
Paperback$69.08$69.08FREE delivery Mon, Sep 21 - Python for Data Analysis: Data Wrangling with pandas, NumPy, and Jupyter
Paperback$43.99$43.99FREE delivery Oct 1 - 16 - Fundamentals of Data Engineering: Plan and Build Robust Data Systems
Paperback$40.00$40.00FREE delivery Sep 29 - Oct 14 - AI Engineering: Building Applications with Foundation Models#1 Best SellerEnterprise Applications
Paperback$52.40$52.40FREE delivery Sun, Sep 20 - Machine Learning Design Patterns: Solutions to Common Challenges in Data Preparation, Model Building, and MLOps
Paperback$36.99$36.99FREE delivery Sun, Sep 20 - Natural Language Processing with Transformers, Revised Edition
Paperback$41.60$41.60FREE delivery Sun, Sep 20 - Data Science from Scratch: First Principles with Python
Paperback$35.54$35.54FREE delivery Mon, Sep 21 - Hands-On Large Language Models: Language Understanding and Generation
Paperback$37.68$37.68FREE delivery Mon, Sep 21 - An Introduction to Statistical Learning: with Applications in Python (Springer Texts in Statistics)
Hardcover$56.21$56.21FREE delivery Fri, Oct 2 - Build a Large Language Model (From Scratch)#1 Best SellerComputer Neural Networks
Paperback$49.24$49.24FREE delivery Sun, Sep 20 - Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems
Paperback$78.13$78.13FREE delivery Thu, Oct 15 - Python Feature Engineering Cookbook: A complete guide to crafting powerful features for your machine learning models
Paperback$46.99$46.99FREE delivery Mon, Sep 21 - Python Data Science Handbook: Essential Tools for Working with Data
Paperback$44.18$44.18FREE delivery Sun, Sep 20 - Essential Math for Data Science: Take Control of Your Data with Fundamental Linear Algebra, Probability, and Statistics
Paperback$36.22$36.22FREE delivery Mon, Sep 21 - Probabilistic Machine Learning: An Introduction (Adaptive Computation and Machine Learning series)
Hardcover$105.22$105.22FREE delivery Tue, Sep 22 - Introducing MLOps: How to Scale Machine Learning in the Enterprise
Paperback$36.49$36.49FREE delivery Sun, Sep 20 - Fluent Python: Clear, Concise, and Effective Programming
Paperback$43.99$43.99FREE delivery Sun, Sep 20 - Deep Learning (Adaptive Computation and Machine Learning series)
Hardcover$51.51$51.51FREE delivery Sun, Sep 20 - Reinforcement Learning, second edition: An Introduction (Adaptive Computation and Machine Learning series)
Hardcover$80.48$80.48FREE delivery Sun, Sep 20 - Hands-On Machine Learning with Scikit-Learn and PyTorch: Concepts, Tools, and Techniques to Build Intelligent Systems
Paperback$77.72$77.72FREE delivery Mon, Sep 21 - Hadoop: The Definitive Guide: Storage and Analysis at Internet Scale
Paperback$23.28$23.28Delivery Sep 30 - Oct 12 - Generative Deep Learning: Teaching Machines To Paint, Write, Compose, and Play
Paperback$47.37$47.37FREE delivery Sun, Sep 20
From the brand
-
Machine Learning, AI & more
-
Machine Learning
-
Artificial Intelligence
-
Deep Learning
-
Language Processing (NLP, LLM)
-
Sharing the knowledge of experts
O'Reilly's mission is to change the world by sharing the knowledge of innovators. For over 40 years, we've inspired companies and individuals to do new things (and do them better) by providing the skills and understanding that are necessary for success.
Our customers are hungry to build the innovations that propel the world forward. And we help them do just that.
From the Publisher
The place of feature engineering in the machine learning workflow.
The book assumes knowledge of basic machine learning concepts, such as 'what is a model,' and 'what is a vector.' It does not assume mastery of mathematics or statistics. Experience with linear algebra, probability distributions, and optimization are helpful, but not necessary.
Code examples in this book are given in Python, using a variety of free and open-source packages. The numpy library provides numeric vector and matrix operations. Pandas is a powerful dataframe that is the building block of data science in Python. Scikit-learn is a general purpose machine learning package with extensive coverage of models and feature transformers. Matplotlib and the styling library Seaborn provide plotting and visualizations.
From the Preface
Machine learning fits mathematical models to data in order to derive insights or make predictions. These models take features as input. A feature is a numeric representation of an aspect of raw data. Features sit between data and models in the machine learning pipeline. Feature engineering is the act of extracting features from raw data, and transforming them into formats that is suitable for the machine learning model. It is a crucial step in the machine learning pipeline, because the right features can ease the difficulty of modeling, and therefore enable the pipeline to output results of higher quality. Practitioners agree that the vast majority of time in building a machine learning pipeline is spent on feature engineering and data cleaning. Yet, despite its importance, the topic is rarely discussed on its own. Perhaps it’s because the right features can only be defined in the context of both the model and the data. Since data and models are so diverse, it’s difficult to generalize the practice of feature engineering across projects.
Nevertheless, feature engineering is not just an ad hoc practice. There are deeper principles at work, and they are best illustrated in situ. Each chapter of this book addresses one data problem: how to represent text data or image data, how to reduce dimensionality of auto-generated features, when and how to normalize, etc. Think of this as a collection of inter-connected short stories, as opposed to a single long novel. Each chapter provides a vignette into in the vast array of existing feature engineering techniques. Together, they illustrate the overarching principles.
Mastering a subject is not just about knowing the definitions and being able to derive the formulas. It is not enough to know how the mechanism works and what it can do. It must also involve understanding why it is designed that way, how it relates to other techniques, and what are the pros and cons of each approach. Mastery is about knowing precisely how something is done, having an intuition for the underlying principles, and integrating it into the knowledge web of what we already know. One does not become a master of something by simply reading a book, though a good book can open new doors. It has to involve practice—putting the ideas to use, which is an iterative process. With every iteration, we know the ideas better and become increasingly more adept and creative at applying them. The goal of this book is to facilitate the application of its ideas.
This book tries to teach the intuition first, and the mathematics second. Instead of only discussing how something is done, we try to teach the why. Our goal is to provide the intuition behind the ideas, so that the reader may understand how and when to apply them. There are tons of descriptions and pictures for folks who learn in different ways. Mathematical formulas are presented in order to make the intuitions precise, and also to bridge this book with other existing offerings of knowledge.
Editorial Reviews
About the Author
Product details
- Publisher : O'Reilly Media
- Publication date : May 8, 2018
- Edition : 1st
- Language : English
- Print length : 215 pages
- ISBN-10 : 1491953241
- ISBN-13 : 978-1491953242
- Item Weight : 12.6 ounces
- Dimensions : 7 x 0.5 x 9 inches
- Best Sellers Rank: #799,776 in Books (See Top 100 in Books)
- #255 in Data Modeling & Design (Books)
- #276 in Data Mining (Books)
- #365 in Data Processing
- Customer Reviews:
About the authors

Discover more of the author’s books, see similar authors, read book recommendations and more.

Amanda is endlessly fascinated (and sometimes horrified) by the difference between the systems we aim to create and the ones that emerge. She has worked in a breadth of cross-functional roles and engineering disciplines for the last 18 years, including developer relations, data science, data engineering, complexity science, and robotics. She creates projects and programs to make the data world more responsible and approachable, including co-authoring the O'Reilly book, Feature Engineering for Machine Learning Principles and Techniques for Data Scientists. She is currently leading research and data engineering to better understand risk and resilience in open source ecosystems.
















