Data Preparation for Data Mining Using SAS

Refaat, Mamdouh

In stock
Regular price 32.250 KD inc. VAT
License
Table of contents
  • Copyright Pageiv
  • Contentsv
  • List of Figuresxv
  • List of Tablesxvii
  • Prefacexxi
  • CHAPTER 1. INTRODUCTION1
  • 1.1 The Data Mining Process1
  • 1.2 Methodologies of Data Mining1
  • 1.3 The Mining View3
  • 1.4 The Scoring View4
  • 1.5 Notes on Data Mining Software4
  • CHAPTER 2. TASKS AND DATA FLOW7
  • 2.1 Data Mining Tasks7
  • 2.2 Data Mining Competencies9
  • 2.3 The Data Flow10
  • 2.4 Types of Variables11
  • 2.5 The Mining View and the Scoring View12
  • 2.6 Steps of Data Preparation13
  • CHAPTER 3. REVIEW OF DATA MINING MODELING TECHNIQUES15
  • 3.1 Introduction15
  • 3.2 Regression Models15
  • 3.3 Decision Trees21
  • 3.4 Neural Networks22
  • 3.5 Cluster Analysis25
  • 3.6 Association Rules26
  • 3.7 Time Series Analysis26
  • 3.8 Support Vector Machines26
  • CHAPTER 4. SAS MACROS: A QUICK START29
  • 4.1 Introduction:Why Macros?29
  • 4.2 The Basics: The Macro and Its Variables30
  • 4.3 Doing Calculations32
  • 4.4 Programming Logic33
  • 4.5 Working with Strings35
  • 4.6 Macros That Call Other Macros36
  • 4.7 Common Macro Patterns and Caveats37
  • 4.8 Where to Go From Here41
  • CHAPTER 5. DATA ACQUISITION AND INTEGRATION43
  • 5.1 Introduction43
  • 5.2 Sources of Data43
  • 5.3 Variable Types45
  • 5.4 Data Rollup47
  • 5.5 Rollup with Sums, Averages, and Counts54
  • 5.6 Calculation of the Mode55
  • 5.7 Data Integration56
  • CHAPTER 6. INTEGRITY CHECKS63
  • 6.1 Introduction63
  • 6.2 Comparing Datasets66
  • 6.3 Dataset Schema Checks66
  • 6.4 Nominal Variables70
  • 6.5 Continuous Variables76
  • CHAPTER 7. EXPLORATORY DATA ANALYSIS83
  • 7.1 Introduction83
  • 7.2 Common EDA Procedures83
  • 7.3 Univariate Statistics84
  • 7.4 Variable Distribution86
  • 7.5 Detection of Outliers86
  • 7.6 Testing Normality96
  • 7.7 Cross-tabulation97
  • 7.8 Investigating Data Structures97
  • CHAPTER 8. SAMPLING AND PARTITIONING99
  • 8.1 Introduction99
  • 8.2 Contents of Samples100
  • 8.3 Random Sampling101
  • 8.4 Balanced Sampling104
  • 8.5 Minimum Sample Size110
  • 8.6 Checking Validity of Sample113
  • CHAPTER 9. DATA TRANSFORMATIONS115
  • 9.1 Raw and Analytical Variables115
  • 9.2 Scope of Data Transformations116
  • 9.3 Creation of New Variables119
  • 9.4 Mapping of Nominal Variables126
  • 9.5 Normalization of Continuous Variables130
  • 9.6 Changing the Variable Distribution131
  • CHAPTER 10. BINNING AND REDUCTION OF CARDINALITY141
  • 10.1 Introduction141
  • 10.2 Cardinality Reduction142
  • 10.3 Binning of Continuous Variables157
  • CHAPTER 11. TREATMENT OF MISSING VALUES171
  • 11.1 Introduction171
  • 11.2 Simple Replacement174
  • 11.3 Imputing Missing Values179
  • 11.4 Imputation Methods and Strategy181
  • 11.5 SAS Macros for Multiple Imputation185
  • 11.6 Predicting Missing Values204
  • CHAPTER 12. PREDICTIVE POWER AND VARIABLE REDUCTION I207
  • 12.1 Introduction207
  • 12.2 Metrics of Predictive Power208
  • 12.3 Methods of Variable Reduction209
  • 12.4 Variable Reduction: Before or During Modeling210
  • CHAPTER 13. ANALYSIS OF NOMINAL AND ORDINAL VARIABLES211
  • 13.1 Introduction211
  • 13.2 Contingency Tables211
  • 13.3 Notation and Definitions212
  • 13.4 Contingency Tables for Binary Variables214
  • 13.5 Contingency Tables for Multicategory Variables225
  • 13.6 Analysis of Ordinal Variables227
  • 13.7 Implementation Scenarios231
  • CHAPTER 14. ANALYSIS OF CONTINUOUS VARIABLES233
  • 14.1 Introduction233
  • 14.2 When Is Binning Necessary?233
  • 14.3 Measures of Association234
  • 14.4 Correlation Coefficients239
  • CHAPTER 15. PRINCIPAL COMPONENT ANALYSIS247
  • 15.1 Introduction247
  • 15.2 Mathematical Formulations248
  • 15.3 Implementing and Using PCA249
  • 15.4 Comments on Using PCA254
  • CHAPTER 16. FACTOR ANALYSIS257
  • 16.1 Introduction257
  • 16.2 Relationship Between PCA and FA263
  • 16.3 Implementation of Factor Analysis263
  • CHAPTER 17. PREDICTIVE POWER AND VARIABLE REDUCTION II267
  • 17.1 Introduction267
  • 17.2 Data with Binary Dependent Variables267
  • 17.3 Data with Continuous Dependent Variables275
  • 17.4 Variable Reduction Strategies275
  • CHAPTER 18. PUTTING IT ALL TOGETHER279
  • 18.1 Introduction279
  • 18.2 The Process of Data Preparation279
  • 18.3 Case Study: The Bookstore281
  • APPENDIX. LISTING OF SAS MACROS297
  • A.1 Copyright and Software License297
  • A.2 Dependencies between Macros298
  • A.3 Data Acquisition and Integration299
  • A.4 Integrity Checks304
  • A.5 Exploratory Data Analysis310
  • A.6 Sampling and Partitioning313
  • A.7 Data Transformations318
  • A.8 Binning and Reduction of Cardinality325
  • A.9 Treatment of Missing Values341
  • A.10 Analysis of Nominal and Ordinal Variables352
  • A.11 Analysis of Continuous Variables358
  • A.12 Principal Component Analysis360
  • A.13 Factor Analysis362
  • A.14 Predictive Power and Variable Reduction II363
  • A.15 Other Macros372
  • Bibliography373
  • Index375
  • About the Author393
Book details
  • Vendor Elsevier S & T
  • SKU 9780123735775
  • ISBN-13 9780080491004
  • Author Refaat, Mamdouh
  • Category Computers
  • Subject Data Warehousing

Do you have questions about this book?

Ask an expert!

Are you a data mining analyst, who spends up to 80% of your time assuring data quality, then preparing that data for developing and deploying predictive models? And do you find lots of literature on data mining theory and concepts, but when it comes to practical advice on developing good mining views find little “how to” information? And are you, like most analysts, preparing the data in SAS?

This book is intended to fill this gap as your source of practical recipes. It introduces a framework for the process of data preparation for data mining, and presents the detailed implementation of each step in SAS. In addition, business applications of data mining modeling require you to deal with a large number of variables, typically hundreds if not thousands. Therefore, the book devotes several chapters to the methods of data transformation and variable selection.

FEATURES
* A complete framework for the data preparation process, including implementation details for each step.
* The complete SAS implementation code, which is readily usable by professional analysts and data miners.
* A unique and comprehensive approach for the treatment of missing values, optimal binning, and cardinality reduction.
* Assumes minimal proficiency in SAS and includes a quick-start chapter on writing SAS macros.
* CD includes dozens of SAS macros plus the sample data and the program for the book's case study.