Problem Statement¶
Business Context¶
Renewable energy sources play an increasingly important role in the global energy mix, as the effort to reduce the environmental impact of energy production increases.
Out of all the renewable energy alternatives, wind energy is one of the most developed technologies worldwide. The U.S Department of Energy has put together a guide to achieving operational efficiency using predictive maintenance practices.
Predictive maintenance uses sensor information and analysis methods to measure and predict degradation and future component capability. The idea behind predictive maintenance is that failure patterns are predictable and if component failure can be predicted accurately and the component is replaced before it fails, the costs of operation and maintenance will be much lower.
The sensors fitted across different machines involved in the process of energy generation collect data related to various environmental factors (temperature, humidity, wind speed, etc.) and additional features related to various parts of the wind turbine (gearbox, tower, blades, break, etc.).
Objective¶
“ReneWind” is a company working on improving the machinery/processes involved in the production of wind energy using machine learning and has collected data of generator failure of wind turbines using sensors. They have shared a ciphered version of the data, as the data collected through sensors is confidential (the type of data collected varies with companies). Data has 40 predictors, 20000 observations in the training set and 5000 in the test set.
The objective is to build various classification models, tune them, and find the best one that will help identify failures so that the generators could be repaired before failing/breaking to reduce the overall maintenance cost. The nature of predictions made by the classification model will translate as follows:
- True positives (TP) are failures correctly predicted by the model. These will result in repairing costs.
- False negatives (FN) are real failures where there is no detection by the model. These will result in replacement costs.
- False positives (FP) are detections where there is no failure. These will result in inspection costs.
It is given that the cost of repairing a generator is much less than the cost of replacing it, and the cost of inspection is less than the cost of repair.
“1” in the target variables should be considered as “failure” and “0” represents “No failure”.
Data Description¶
- The data provided is a transformed version of original data which was collected using sensors.
- Train.csv - To be used for training and tuning of models.
- Test.csv - To be used only for testing the performance of the final best model.
- Both the datasets consist of 40 predictor variables and 1 target variable
Importing necessary libraries¶
# Libraries to help with reading and manipulating data
import pandas as pd
import numpy as np
# Libaries to help with data visualization
import matplotlib.pyplot as plt
import seaborn as sns
# To tune model, get different metric scores, and split data
from sklearn.metrics import (
f1_score,
accuracy_score,
recall_score,
precision_score,
confusion_matrix,
roc_auc_score,
ConfusionMatrixDisplay,
)
from sklearn import metrics
from sklearn.model_selection import train_test_split, StratifiedKFold, cross_val_score
# To be used for data scaling and one hot encoding
from sklearn.preprocessing import StandardScaler, MinMaxScaler, OneHotEncoder
# To impute missing values
from sklearn.impute import SimpleImputer
# To oversample and undersample data
from imblearn.over_sampling import SMOTE
from imblearn.under_sampling import RandomUnderSampler
# To do hyperparameter tuning
from sklearn.model_selection import RandomizedSearchCV
# To be used for creating pipelines and personalizing them
from sklearn.pipeline import Pipeline
from sklearn.compose import ColumnTransformer
# To define maximum number of columns to be displayed in a dataframe
pd.set_option("display.max_columns", None)
pd.set_option("display.max_rows", None)
# To supress scientific notations for a dataframe
pd.set_option("display.float_format", lambda x: "%.3f" % x)
# To help with model building
from sklearn.linear_model import LogisticRegression
from sklearn.tree import DecisionTreeClassifier
from sklearn.ensemble import (
AdaBoostClassifier,
GradientBoostingClassifier,
RandomForestClassifier,
BaggingClassifier,
)
from xgboost import XGBClassifier
# To suppress scientific notations
pd.set_option("display.float_format", lambda x: "%.3f" % x)
# To suppress warnings
import warnings
warnings.filterwarnings("ignore")
Loading the dataset¶
df = pd.read_csv('train.csv')
df_test = pd.read_csv('test.csv')
Data Overview¶
- Observations
- Sanity checks
df.shape
(20000, 41)
df_test.shape
(5000, 41)
data = df.copy()
data_test = df_test.copy()
data.head()
| V1 | V2 | V3 | V4 | V5 | V6 | V7 | V8 | V9 | V10 | V11 | V12 | V13 | V14 | V15 | V16 | V17 | V18 | V19 | V20 | V21 | V22 | V23 | V24 | V25 | V26 | V27 | V28 | V29 | V30 | V31 | V32 | V33 | V34 | V35 | V36 | V37 | V38 | V39 | V40 | Target | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 0 | -4.465 | -4.679 | 3.102 | 0.506 | -0.221 | -2.033 | -2.911 | 0.051 | -1.522 | 3.762 | -5.715 | 0.736 | 0.981 | 1.418 | -3.376 | -3.047 | 0.306 | 2.914 | 2.270 | 4.395 | -2.388 | 0.646 | -1.191 | 3.133 | 0.665 | -2.511 | -0.037 | 0.726 | -3.982 | -1.073 | 1.667 | 3.060 | -1.690 | 2.846 | 2.235 | 6.667 | 0.444 | -2.369 | 2.951 | -3.480 | 0 |
| 1 | 3.366 | 3.653 | 0.910 | -1.368 | 0.332 | 2.359 | 0.733 | -4.332 | 0.566 | -0.101 | 1.914 | -0.951 | -1.255 | -2.707 | 0.193 | -4.769 | -2.205 | 0.908 | 0.757 | -5.834 | -3.065 | 1.597 | -1.757 | 1.766 | -0.267 | 3.625 | 1.500 | -0.586 | 0.783 | -0.201 | 0.025 | -1.795 | 3.033 | -2.468 | 1.895 | -2.298 | -1.731 | 5.909 | -0.386 | 0.616 | 0 |
| 2 | -3.832 | -5.824 | 0.634 | -2.419 | -1.774 | 1.017 | -2.099 | -3.173 | -2.082 | 5.393 | -0.771 | 1.107 | 1.144 | 0.943 | -3.164 | -4.248 | -4.039 | 3.689 | 3.311 | 1.059 | -2.143 | 1.650 | -1.661 | 1.680 | -0.451 | -4.551 | 3.739 | 1.134 | -2.034 | 0.841 | -1.600 | -0.257 | 0.804 | 4.086 | 2.292 | 5.361 | 0.352 | 2.940 | 3.839 | -4.309 | 0 |
| 3 | 1.618 | 1.888 | 7.046 | -1.147 | 0.083 | -1.530 | 0.207 | -2.494 | 0.345 | 2.119 | -3.053 | 0.460 | 2.705 | -0.636 | -0.454 | -3.174 | -3.404 | -1.282 | 1.582 | -1.952 | -3.517 | -1.206 | -5.628 | -1.818 | 2.124 | 5.295 | 4.748 | -2.309 | -3.963 | -6.029 | 4.949 | -3.584 | -2.577 | 1.364 | 0.623 | 5.550 | -1.527 | 0.139 | 3.101 | -1.277 | 0 |
| 4 | -0.111 | 3.872 | -3.758 | -2.983 | 3.793 | 0.545 | 0.205 | 4.849 | -1.855 | -6.220 | 1.998 | 4.724 | 0.709 | -1.989 | -2.633 | 4.184 | 2.245 | 3.734 | -6.313 | -5.380 | -0.887 | 2.062 | 9.446 | 4.490 | -3.945 | 4.582 | -8.780 | -3.383 | 5.107 | 6.788 | 2.044 | 8.266 | 6.629 | -10.069 | 1.223 | -3.230 | 1.687 | -2.164 | -3.645 | 6.510 | 0 |
data.sample(5)
| V1 | V2 | V3 | V4 | V5 | V6 | V7 | V8 | V9 | V10 | V11 | V12 | V13 | V14 | V15 | V16 | V17 | V18 | V19 | V20 | V21 | V22 | V23 | V24 | V25 | V26 | V27 | V28 | V29 | V30 | V31 | V32 | V33 | V34 | V35 | V36 | V37 | V38 | V39 | V40 | Target | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 12338 | -1.535 | -1.734 | -3.303 | -2.111 | -0.904 | -0.166 | 0.378 | 2.944 | -1.823 | -1.676 | 2.035 | 5.845 | 1.113 | -0.306 | -1.230 | 2.112 | 1.129 | 1.965 | 0.183 | 0.988 | -0.095 | 1.848 | 5.339 | 0.173 | -1.662 | -3.129 | -2.100 | 0.187 | 2.560 | 4.607 | -3.900 | -0.078 | 0.231 | 0.116 | 0.336 | -0.803 | 2.744 | -0.387 | -0.225 | 1.268 | 0 |
| 6861 | 3.354 | 0.765 | 11.293 | 4.605 | -5.913 | -1.075 | -4.636 | -9.821 | 6.242 | 1.500 | -0.515 | -4.854 | 7.709 | -3.849 | -9.656 | -17.272 | -6.138 | -1.355 | 9.069 | 1.912 | -16.670 | 1.441 | -12.101 | -4.114 | 0.970 | 6.103 | 3.265 | -2.477 | -0.550 | 1.640 | -4.021 | -6.431 | 3.720 | 1.756 | 12.765 | -2.533 | -4.037 | 0.523 | 2.150 | -10.119 | 0 |
| 10172 | 1.229 | -3.084 | 4.597 | -1.302 | -3.546 | -1.403 | -1.686 | -1.573 | 0.632 | 2.423 | 0.341 | 1.293 | 5.984 | 0.412 | -4.308 | -4.128 | -5.455 | -0.421 | 3.501 | 1.051 | -6.896 | 0.997 | -2.978 | -3.247 | 0.401 | 0.328 | 4.355 | -2.268 | -1.221 | 0.961 | 0.133 | -1.827 | 0.270 | 1.896 | 5.619 | 2.953 | -0.882 | -2.762 | 2.103 | -4.408 | 0 |
| 17344 | 3.723 | 4.795 | 8.260 | 5.925 | -1.802 | -2.128 | -2.184 | -5.727 | 4.375 | -0.510 | -4.019 | -2.947 | 1.977 | -4.014 | -4.676 | -12.724 | 0.610 | -1.411 | 6.735 | 0.560 | -11.515 | 1.786 | -6.687 | 1.945 | 1.364 | 7.092 | -1.059 | -0.551 | -2.514 | -1.494 | -1.135 | -1.810 | 0.342 | -0.253 | 8.278 | -3.288 | -3.498 | 1.560 | -0.117 | -5.568 | 0 |
| 17949 | 0.999 | -2.325 | -0.219 | 0.865 | -1.828 | -0.997 | -2.084 | 1.375 | 0.011 | -0.437 | -0.669 | 0.963 | 0.931 | -0.233 | -3.270 | -3.223 | 1.191 | 1.962 | 1.895 | 2.151 | -5.378 | 3.289 | 3.040 | 3.056 | -0.715 | -2.208 | -2.762 | 0.066 | 0.859 | 5.127 | -1.672 | 4.347 | 2.029 | -1.891 | 6.287 | -1.894 | -0.088 | -2.955 | -1.923 | -1.566 | 0 |
data.tail(5)
| V1 | V2 | V3 | V4 | V5 | V6 | V7 | V8 | V9 | V10 | V11 | V12 | V13 | V14 | V15 | V16 | V17 | V18 | V19 | V20 | V21 | V22 | V23 | V24 | V25 | V26 | V27 | V28 | V29 | V30 | V31 | V32 | V33 | V34 | V35 | V36 | V37 | V38 | V39 | V40 | Target | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 19995 | -2.071 | -1.088 | -0.796 | -3.012 | -2.288 | 2.807 | 0.481 | 0.105 | -0.587 | -2.899 | 8.868 | 1.717 | 1.358 | -1.777 | 0.710 | 4.945 | -3.100 | -1.199 | -1.085 | -0.365 | 3.131 | -3.948 | -3.578 | -8.139 | -1.937 | -1.328 | -0.403 | -1.735 | 9.996 | 6.955 | -3.938 | -8.274 | 5.745 | 0.589 | -0.650 | -3.043 | 2.216 | 0.609 | 0.178 | 2.928 | 1 |
| 19996 | 2.890 | 2.483 | 5.644 | 0.937 | -1.381 | 0.412 | -1.593 | -5.762 | 2.150 | 0.272 | -2.095 | -1.526 | 0.072 | -3.540 | -2.762 | -10.632 | -0.495 | 1.720 | 3.872 | -1.210 | -8.222 | 2.121 | -5.492 | 1.452 | 1.450 | 3.685 | 1.077 | -0.384 | -0.839 | -0.748 | -1.089 | -4.159 | 1.181 | -0.742 | 5.369 | -0.693 | -1.669 | 3.660 | 0.820 | -1.987 | 0 |
| 19997 | -3.897 | -3.942 | -0.351 | -2.417 | 1.108 | -1.528 | -3.520 | 2.055 | -0.234 | -0.358 | -3.782 | 2.180 | 6.112 | 1.985 | -8.330 | -1.639 | -0.915 | 5.672 | -3.924 | 2.133 | -4.502 | 2.777 | 5.728 | 1.620 | -1.700 | -0.042 | -2.923 | -2.760 | -2.254 | 2.552 | 0.982 | 7.112 | 1.476 | -3.954 | 1.856 | 5.029 | 2.083 | -6.409 | 1.477 | -0.874 | 0 |
| 19998 | -3.187 | -10.052 | 5.696 | -4.370 | -5.355 | -1.873 | -3.947 | 0.679 | -2.389 | 5.457 | 1.583 | 3.571 | 9.227 | 2.554 | -7.039 | -0.994 | -9.665 | 1.155 | 3.877 | 3.524 | -7.015 | -0.132 | -3.446 | -4.801 | -0.876 | -3.812 | 5.422 | -3.732 | 0.609 | 5.256 | 1.915 | 0.403 | 3.164 | 3.752 | 8.530 | 8.451 | 0.204 | -7.130 | 4.249 | -6.112 | 0 |
| 19999 | -2.687 | 1.961 | 6.137 | 2.600 | 2.657 | -4.291 | -2.344 | 0.974 | -1.027 | 0.497 | -9.589 | 3.177 | 1.055 | -1.416 | -4.669 | -5.405 | 3.720 | 2.893 | 2.329 | 1.458 | -6.429 | 1.818 | 0.806 | 7.786 | 0.331 | 5.257 | -4.867 | -0.819 | -5.667 | -2.861 | 4.674 | 6.621 | -1.989 | -1.349 | 3.952 | 5.450 | -0.455 | -2.202 | 1.678 | -1.974 | 0 |
data_test.head()
| V1 | V2 | V3 | V4 | V5 | V6 | V7 | V8 | V9 | V10 | V11 | V12 | V13 | V14 | V15 | V16 | V17 | V18 | V19 | V20 | V21 | V22 | V23 | V24 | V25 | V26 | V27 | V28 | V29 | V30 | V31 | V32 | V33 | V34 | V35 | V36 | V37 | V38 | V39 | V40 | Target | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 0 | -0.613 | -3.820 | 2.202 | 1.300 | -1.185 | -4.496 | -1.836 | 4.723 | 1.206 | -0.342 | -5.123 | 1.017 | 4.819 | 3.269 | -2.984 | 1.387 | 2.032 | -0.512 | -1.023 | 7.339 | -2.242 | 0.155 | 2.054 | -2.772 | 1.851 | -1.789 | -0.277 | -1.255 | -3.833 | -1.505 | 1.587 | 2.291 | -5.411 | 0.870 | 0.574 | 4.157 | 1.428 | -10.511 | 0.455 | -1.448 | 0 |
| 1 | 0.390 | -0.512 | 0.527 | -2.577 | -1.017 | 2.235 | -0.441 | -4.406 | -0.333 | 1.967 | 1.797 | 0.410 | 0.638 | -1.390 | -1.883 | -5.018 | -3.827 | 2.418 | 1.762 | -3.242 | -3.193 | 1.857 | -1.708 | 0.633 | -0.588 | 0.084 | 3.014 | -0.182 | 0.224 | 0.865 | -1.782 | -2.475 | 2.494 | 0.315 | 2.059 | 0.684 | -0.485 | 5.128 | 1.721 | -1.488 | 0 |
| 2 | -0.875 | -0.641 | 4.084 | -1.590 | 0.526 | -1.958 | -0.695 | 1.347 | -1.732 | 0.466 | -4.928 | 3.565 | -0.449 | -0.656 | -0.167 | -1.630 | 2.292 | 2.396 | 0.601 | 1.794 | -2.120 | 0.482 | -0.841 | 1.790 | 1.874 | 0.364 | -0.169 | -0.484 | -2.119 | -2.157 | 2.907 | -1.319 | -2.997 | 0.460 | 0.620 | 5.632 | 1.324 | -1.752 | 1.808 | 1.676 | 0 |
| 3 | 0.238 | 1.459 | 4.015 | 2.534 | 1.197 | -3.117 | -0.924 | 0.269 | 1.322 | 0.702 | -5.578 | -0.851 | 2.591 | 0.767 | -2.391 | -2.342 | 0.572 | -0.934 | 0.509 | 1.211 | -3.260 | 0.105 | -0.659 | 1.498 | 1.100 | 4.143 | -0.248 | -1.137 | -5.356 | -4.546 | 3.809 | 3.518 | -3.074 | -0.284 | 0.955 | 3.029 | -1.367 | -3.412 | 0.906 | -2.451 | 0 |
| 4 | 5.828 | 2.768 | -1.235 | 2.809 | -1.642 | -1.407 | 0.569 | 0.965 | 1.918 | -2.775 | -0.530 | 1.375 | -0.651 | -1.679 | -0.379 | -4.443 | 3.894 | -0.608 | 2.945 | 0.367 | -5.789 | 4.598 | 4.450 | 3.225 | 0.397 | 0.248 | -2.362 | 1.079 | -0.473 | 2.243 | -3.591 | 1.774 | -1.502 | -2.227 | 4.777 | -6.560 | -0.806 | -0.276 | -3.858 | -0.538 | 0 |
data_test.sample(5)
| V1 | V2 | V3 | V4 | V5 | V6 | V7 | V8 | V9 | V10 | V11 | V12 | V13 | V14 | V15 | V16 | V17 | V18 | V19 | V20 | V21 | V22 | V23 | V24 | V25 | V26 | V27 | V28 | V29 | V30 | V31 | V32 | V33 | V34 | V35 | V36 | V37 | V38 | V39 | V40 | Target | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 3755 | -3.111 | -4.198 | 3.243 | 2.446 | -0.960 | -6.274 | -1.849 | 4.942 | -1.634 | 1.914 | -9.737 | 6.287 | 2.006 | 1.868 | -2.965 | -2.087 | 5.542 | 1.729 | 4.447 | 8.866 | -4.126 | 2.628 | 3.626 | 4.040 | 2.063 | -3.632 | -1.745 | 1.644 | -7.131 | -2.544 | 0.214 | 3.403 | -8.752 | 4.689 | 2.270 | 6.903 | 2.160 | -6.531 | 2.218 | -3.451 | 0 |
| 1617 | 1.123 | 3.759 | 7.521 | 2.782 | -0.660 | -0.293 | -1.557 | -7.559 | 3.208 | 0.948 | -4.199 | -2.527 | 0.678 | -3.635 | -3.209 | -11.914 | -0.288 | 0.686 | 4.772 | -0.199 | -7.553 | 0.593 | -8.347 | 0.437 | 2.065 | 5.635 | 1.779 | 0.039 | -3.203 | -4.493 | -1.236 | -6.284 | -1.118 | 1.662 | 3.080 | 0.780 | -1.870 | 4.905 | 2.764 | -3.755 | 0 |
| 1522 | -1.817 | 2.683 | 3.408 | 0.017 | 1.528 | -0.965 | -0.279 | 0.278 | -2.126 | 0.032 | -0.061 | 3.216 | 1.391 | -2.184 | -2.170 | 0.059 | -2.022 | -0.118 | 1.822 | -3.959 | -2.774 | -0.706 | -0.485 | 3.557 | -2.341 | 5.595 | -3.007 | -2.142 | 1.109 | 1.739 | 3.682 | 4.109 | 4.195 | -1.665 | 3.772 | 1.193 | -1.271 | 1.329 | 0.658 | -0.801 | 0 |
| 1064 | -0.754 | 2.413 | 1.505 | -5.091 | 2.987 | 1.801 | -2.129 | -0.762 | 0.928 | -4.410 | 1.320 | 0.023 | 5.378 | -1.602 | -6.680 | -0.974 | -3.418 | 4.976 | -8.182 | -5.399 | -4.067 | 0.190 | 1.687 | -1.615 | -2.444 | 7.471 | -3.623 | -6.248 | 3.766 | 3.533 | 3.660 | 2.630 | 7.808 | -10.004 | 0.912 | 1.522 | 1.081 | -3.112 | 0.081 | 4.676 | 0 |
| 4816 | -3.225 | 2.331 | -2.313 | 2.851 | 1.508 | -1.172 | 2.272 | 1.842 | -4.185 | 1.020 | -1.984 | 6.542 | -6.318 | -2.288 | 3.745 | 1.114 | 5.602 | -0.076 | 6.172 | -0.330 | 2.945 | 1.212 | 3.575 | 8.726 | -1.347 | -1.801 | -4.206 | 4.643 | -0.971 | 0.034 | -2.558 | 2.244 | -2.616 | 4.572 | -0.547 | -1.259 | 0.621 | 7.473 | 0.107 | -0.735 | 0 |
data_test.tail(5)
| V1 | V2 | V3 | V4 | V5 | V6 | V7 | V8 | V9 | V10 | V11 | V12 | V13 | V14 | V15 | V16 | V17 | V18 | V19 | V20 | V21 | V22 | V23 | V24 | V25 | V26 | V27 | V28 | V29 | V30 | V31 | V32 | V33 | V34 | V35 | V36 | V37 | V38 | V39 | V40 | Target | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 4995 | -5.120 | 1.635 | 1.251 | 4.036 | 3.291 | -2.932 | -1.329 | 1.754 | -2.985 | 1.249 | -6.878 | 3.715 | -2.512 | -1.395 | -2.554 | -2.197 | 4.772 | 2.403 | 3.792 | 0.487 | -2.028 | 1.778 | 3.668 | 11.375 | -1.977 | 2.252 | -7.319 | 1.907 | -3.734 | -0.012 | 2.120 | 9.979 | 0.063 | 0.217 | 3.036 | 2.109 | -0.557 | 1.939 | 0.513 | -2.694 | 0 |
| 4996 | -5.172 | 1.172 | 1.579 | 1.220 | 2.530 | -0.669 | -2.618 | -2.001 | 0.634 | -0.579 | -3.671 | 0.460 | 3.321 | -1.075 | -7.113 | -4.356 | -0.001 | 3.698 | -0.846 | -0.222 | -3.645 | 0.736 | 0.926 | 3.278 | -2.277 | 4.458 | -4.543 | -1.348 | -1.779 | 0.352 | -0.214 | 4.424 | 2.604 | -2.152 | 0.917 | 2.157 | 0.467 | 0.470 | 2.197 | -2.377 | 0 |
| 4997 | -1.114 | -0.404 | -1.765 | -5.879 | 3.572 | 3.711 | -2.483 | -0.308 | -0.922 | -2.999 | -0.112 | -1.977 | -1.623 | -0.945 | -2.735 | -0.813 | 0.610 | 8.149 | -9.199 | -3.872 | -0.296 | 1.468 | 2.884 | 2.792 | -1.136 | 1.198 | -4.342 | -2.869 | 4.124 | 4.197 | 3.471 | 3.792 | 7.482 | -10.061 | -0.387 | 1.849 | 1.818 | -1.246 | -1.261 | 7.475 | 0 |
| 4998 | -1.703 | 0.615 | 6.221 | -0.104 | 0.956 | -3.279 | -1.634 | -0.104 | 1.388 | -1.066 | -7.970 | 2.262 | 3.134 | -0.486 | -3.498 | -4.562 | 3.136 | 2.536 | -0.792 | 4.398 | -4.073 | -0.038 | -2.371 | -1.542 | 2.908 | 3.215 | -0.169 | -1.541 | -4.724 | -5.525 | 1.668 | -4.100 | -5.949 | 0.550 | -1.574 | 6.824 | 2.139 | -4.036 | 3.436 | 0.579 | 0 |
| 4999 | -0.604 | 0.960 | -0.721 | 8.230 | -1.816 | -2.276 | -2.575 | -1.041 | 4.130 | -2.731 | -3.292 | -1.674 | 0.465 | -1.646 | -5.263 | -7.988 | 6.480 | 0.226 | 4.963 | 6.752 | -6.306 | 3.271 | 1.897 | 3.271 | -0.637 | -0.925 | -6.759 | 2.990 | -0.814 | 3.499 | -8.435 | 2.370 | -1.062 | 0.791 | 4.952 | -7.441 | -0.070 | -0.918 | -2.291 | -5.363 | 0 |
data.info()
<class 'pandas.core.frame.DataFrame'> RangeIndex: 20000 entries, 0 to 19999 Data columns (total 41 columns): # Column Non-Null Count Dtype --- ------ -------------- ----- 0 V1 19982 non-null float64 1 V2 19982 non-null float64 2 V3 20000 non-null float64 3 V4 20000 non-null float64 4 V5 20000 non-null float64 5 V6 20000 non-null float64 6 V7 20000 non-null float64 7 V8 20000 non-null float64 8 V9 20000 non-null float64 9 V10 20000 non-null float64 10 V11 20000 non-null float64 11 V12 20000 non-null float64 12 V13 20000 non-null float64 13 V14 20000 non-null float64 14 V15 20000 non-null float64 15 V16 20000 non-null float64 16 V17 20000 non-null float64 17 V18 20000 non-null float64 18 V19 20000 non-null float64 19 V20 20000 non-null float64 20 V21 20000 non-null float64 21 V22 20000 non-null float64 22 V23 20000 non-null float64 23 V24 20000 non-null float64 24 V25 20000 non-null float64 25 V26 20000 non-null float64 26 V27 20000 non-null float64 27 V28 20000 non-null float64 28 V29 20000 non-null float64 29 V30 20000 non-null float64 30 V31 20000 non-null float64 31 V32 20000 non-null float64 32 V33 20000 non-null float64 33 V34 20000 non-null float64 34 V35 20000 non-null float64 35 V36 20000 non-null float64 36 V37 20000 non-null float64 37 V38 20000 non-null float64 38 V39 20000 non-null float64 39 V40 20000 non-null float64 40 Target 20000 non-null int64 dtypes: float64(40), int64(1) memory usage: 6.3 MB
data.duplicated().sum()
0
data.nunique()
V1 19982 V2 19982 V3 20000 V4 20000 V5 20000 V6 20000 V7 20000 V8 20000 V9 20000 V10 20000 V11 20000 V12 20000 V13 20000 V14 20000 V15 20000 V16 20000 V17 20000 V18 20000 V19 20000 V20 20000 V21 20000 V22 20000 V23 20000 V24 20000 V25 20000 V26 20000 V27 20000 V28 20000 V29 20000 V30 20000 V31 20000 V32 20000 V33 20000 V34 20000 V35 20000 V36 20000 V37 20000 V38 20000 V39 20000 V40 20000 Target 2 dtype: int64
data.isnull().sum()
V1 18 V2 18 V3 0 V4 0 V5 0 V6 0 V7 0 V8 0 V9 0 V10 0 V11 0 V12 0 V13 0 V14 0 V15 0 V16 0 V17 0 V18 0 V19 0 V20 0 V21 0 V22 0 V23 0 V24 0 V25 0 V26 0 V27 0 V28 0 V29 0 V30 0 V31 0 V32 0 V33 0 V34 0 V35 0 V36 0 V37 0 V38 0 V39 0 V40 0 Target 0 dtype: int64
data.describe().T
| count | mean | std | min | 25% | 50% | 75% | max | |
|---|---|---|---|---|---|---|---|---|
| V1 | 19982.000 | -0.272 | 3.442 | -11.876 | -2.737 | -0.748 | 1.840 | 15.493 |
| V2 | 19982.000 | 0.440 | 3.151 | -12.320 | -1.641 | 0.472 | 2.544 | 13.089 |
| V3 | 20000.000 | 2.485 | 3.389 | -10.708 | 0.207 | 2.256 | 4.566 | 17.091 |
| V4 | 20000.000 | -0.083 | 3.432 | -15.082 | -2.348 | -0.135 | 2.131 | 13.236 |
| V5 | 20000.000 | -0.054 | 2.105 | -8.603 | -1.536 | -0.102 | 1.340 | 8.134 |
| V6 | 20000.000 | -0.995 | 2.041 | -10.227 | -2.347 | -1.001 | 0.380 | 6.976 |
| V7 | 20000.000 | -0.879 | 1.762 | -7.950 | -2.031 | -0.917 | 0.224 | 8.006 |
| V8 | 20000.000 | -0.548 | 3.296 | -15.658 | -2.643 | -0.389 | 1.723 | 11.679 |
| V9 | 20000.000 | -0.017 | 2.161 | -8.596 | -1.495 | -0.068 | 1.409 | 8.138 |
| V10 | 20000.000 | -0.013 | 2.193 | -9.854 | -1.411 | 0.101 | 1.477 | 8.108 |
| V11 | 20000.000 | -1.895 | 3.124 | -14.832 | -3.922 | -1.921 | 0.119 | 11.826 |
| V12 | 20000.000 | 1.605 | 2.930 | -12.948 | -0.397 | 1.508 | 3.571 | 15.081 |
| V13 | 20000.000 | 1.580 | 2.875 | -13.228 | -0.224 | 1.637 | 3.460 | 15.420 |
| V14 | 20000.000 | -0.951 | 1.790 | -7.739 | -2.171 | -0.957 | 0.271 | 5.671 |
| V15 | 20000.000 | -2.415 | 3.355 | -16.417 | -4.415 | -2.383 | -0.359 | 12.246 |
| V16 | 20000.000 | -2.925 | 4.222 | -20.374 | -5.634 | -2.683 | -0.095 | 13.583 |
| V17 | 20000.000 | -0.134 | 3.345 | -14.091 | -2.216 | -0.015 | 2.069 | 16.756 |
| V18 | 20000.000 | 1.189 | 2.592 | -11.644 | -0.404 | 0.883 | 2.572 | 13.180 |
| V19 | 20000.000 | 1.182 | 3.397 | -13.492 | -1.050 | 1.279 | 3.493 | 13.238 |
| V20 | 20000.000 | 0.024 | 3.669 | -13.923 | -2.433 | 0.033 | 2.512 | 16.052 |
| V21 | 20000.000 | -3.611 | 3.568 | -17.956 | -5.930 | -3.533 | -1.266 | 13.840 |
| V22 | 20000.000 | 0.952 | 1.652 | -10.122 | -0.118 | 0.975 | 2.026 | 7.410 |
| V23 | 20000.000 | -0.366 | 4.032 | -14.866 | -3.099 | -0.262 | 2.452 | 14.459 |
| V24 | 20000.000 | 1.134 | 3.912 | -16.387 | -1.468 | 0.969 | 3.546 | 17.163 |
| V25 | 20000.000 | -0.002 | 2.017 | -8.228 | -1.365 | 0.025 | 1.397 | 8.223 |
| V26 | 20000.000 | 1.874 | 3.435 | -11.834 | -0.338 | 1.951 | 4.130 | 16.836 |
| V27 | 20000.000 | -0.612 | 4.369 | -14.905 | -3.652 | -0.885 | 2.189 | 17.560 |
| V28 | 20000.000 | -0.883 | 1.918 | -9.269 | -2.171 | -0.891 | 0.376 | 6.528 |
| V29 | 20000.000 | -0.986 | 2.684 | -12.579 | -2.787 | -1.176 | 0.630 | 10.722 |
| V30 | 20000.000 | -0.016 | 3.005 | -14.796 | -1.867 | 0.184 | 2.036 | 12.506 |
| V31 | 20000.000 | 0.487 | 3.461 | -13.723 | -1.818 | 0.490 | 2.731 | 17.255 |
| V32 | 20000.000 | 0.304 | 5.500 | -19.877 | -3.420 | 0.052 | 3.762 | 23.633 |
| V33 | 20000.000 | 0.050 | 3.575 | -16.898 | -2.243 | -0.066 | 2.255 | 16.692 |
| V34 | 20000.000 | -0.463 | 3.184 | -17.985 | -2.137 | -0.255 | 1.437 | 14.358 |
| V35 | 20000.000 | 2.230 | 2.937 | -15.350 | 0.336 | 2.099 | 4.064 | 15.291 |
| V36 | 20000.000 | 1.515 | 3.801 | -14.833 | -0.944 | 1.567 | 3.984 | 19.330 |
| V37 | 20000.000 | 0.011 | 1.788 | -5.478 | -1.256 | -0.128 | 1.176 | 7.467 |
| V38 | 20000.000 | -0.344 | 3.948 | -17.375 | -2.988 | -0.317 | 2.279 | 15.290 |
| V39 | 20000.000 | 0.891 | 1.753 | -6.439 | -0.272 | 0.919 | 2.058 | 7.760 |
| V40 | 20000.000 | -0.876 | 3.012 | -11.024 | -2.940 | -0.921 | 1.120 | 10.654 |
| Target | 20000.000 | 0.056 | 0.229 | 0.000 | 0.000 | 0.000 | 0.000 | 1.000 |
Exploratory Data Analysis (EDA)¶
Plotting histograms and boxplots for all the variables¶
# function to plot a boxplot and a histogram along the same scale.
def histogram_boxplot(data, feature, figsize=(12, 7), kde=False, bins=None):
"""
Boxplot and histogram combined
data: dataframe
feature: dataframe column
figsize: size of figure (default (12,7))
kde: whether to the show density curve (default False)
bins: number of bins for histogram (default None)
"""
f2, (ax_box2, ax_hist2) = plt.subplots(
nrows=2, # Number of rows of the subplot grid= 2
sharex=True, # x-axis will be shared among all subplots
gridspec_kw={"height_ratios": (0.25, 0.75)},
figsize=figsize,
) # creating the 2 subplots
sns.boxplot(
data=data, x=feature, ax=ax_box2, showmeans=True, color="violet"
) # boxplot will be created and a star will indicate the mean value of the column
sns.histplot(
data=data, x=feature, kde=kde, ax=ax_hist2, bins=bins, palette="winter"
) if bins else sns.histplot(
data=data, x=feature, kde=kde, ax=ax_hist2
) # For histogram
ax_hist2.axvline(
data[feature].mean(), color="green", linestyle="--"
) # Add mean to the histogram
ax_hist2.axvline(
data[feature].median(), color="black", linestyle="-"
) # Add median to the histogram
Plotting all the features at one go¶
for feature in df.columns:
histogram_boxplot(df, feature, figsize=(12, 7), kde=False, bins=None) ## Please change the dataframe name as you define while reading the data
- All the columns seem to be about normally distributed.
- The target is unbalanced.
Data Pre-processing¶
data["Target"].value_counts(1)
Target 0 0.945 1 0.056 Name: proportion, dtype: float64
data_test["Target"].value_counts(1)
Target 0 0.944 1 0.056 Name: proportion, dtype: float64
There is about 5% failure rate in both train and test.
Data Pre-Processing¶
# Dividing train data into X and y
X = data.drop(["Target"], axis=1)
y = data["Target"]
# Dividing train data into X and y
X_test = data_test.drop(["Target"], axis=1)
y_test = data_test["Target"]
# Splitting train dataset into training and validation set in the ratio 70:30
X_train, X_val, y_train, y_val = train_test_split(X, y, test_size=0.30, random_state=1, stratify=y)
X_train.shape
(14000, 40)
X_val.shape
(6000, 40)
X_test.shape
(5000, 40)
Missing value imputation¶
# creating an instace of the imputer to be used
imputer = SimpleImputer(strategy="median")
# Fit and transform the train data
X_train = pd.DataFrame(imputer.fit_transform(X_train), columns=X_train.columns)
# Transform the validation data
X_val = pd.DataFrame(imputer.transform(X_val), columns=X_train.columns)
# Transform the test data
X_test = pd.DataFrame(imputer.transform(X_test), columns=X_train.columns)
# Checking that no column has missing values in train, val or test sets
print(X_train.isna().sum())
print("-" * 50)
print(X_val.isna().sum())
print("-" * 50)
V1 0 V2 0 V3 0 V4 0 V5 0 V6 0 V7 0 V8 0 V9 0 V10 0 V11 0 V12 0 V13 0 V14 0 V15 0 V16 0 V17 0 V18 0 V19 0 V20 0 V21 0 V22 0 V23 0 V24 0 V25 0 V26 0 V27 0 V28 0 V29 0 V30 0 V31 0 V32 0 V33 0 V34 0 V35 0 V36 0 V37 0 V38 0 V39 0 V40 0 dtype: int64 -------------------------------------------------- V1 0 V2 0 V3 0 V4 0 V5 0 V6 0 V7 0 V8 0 V9 0 V10 0 V11 0 V12 0 V13 0 V14 0 V15 0 V16 0 V17 0 V18 0 V19 0 V20 0 V21 0 V22 0 V23 0 V24 0 V25 0 V26 0 V27 0 V28 0 V29 0 V30 0 V31 0 V32 0 V33 0 V34 0 V35 0 V36 0 V37 0 V38 0 V39 0 V40 0 dtype: int64 --------------------------------------------------
Model Building¶
Model evaluation criterion¶
The nature of predictions made by the classification model will translate as follows:
- True positives (TP) are failures correctly predicted by the model.
- False negatives (FN) are real failures in a generator where there is no detection by model.
- False positives (FP) are failure detections in a generator where there is no failure.
Which metric to optimize?
- We need to choose the metric which will ensure that the maximum number of generator failures are predicted correctly by the model.
- We would want Recall to be maximized as greater the Recall, the higher the chances of minimizing false negatives.
- We want to minimize false negatives because if a model predicts that a machine will have no failure when there will be a failure, it will increase the maintenance cost.
Let's define a function to output different metrics (including recall) on the train and test set and a function to show confusion matrix so that we do not have to use the same code repetitively while evaluating models.
# defining a function to compute different metrics to check performance of a classification model built using sklearn
def model_performance_classification_sklearn(model, predictors, target):
"""
Function to compute different metrics to check classification model performance
model: classifier
predictors: independent variables
target: dependent variable
"""
# predicting using the independent variables
pred = model.predict(predictors)
acc = accuracy_score(target, pred) # to compute Accuracy
recall = recall_score(target, pred) # to compute Recall
precision = precision_score(target, pred) # to compute Precision
f1 = f1_score(target, pred) # to compute F1-score
# creating a dataframe of metrics
df_perf = pd.DataFrame(
{
"Accuracy": acc,
"Recall": recall,
"Precision": precision,
"F1": f1
},
index=[0],
)
return df_perf
Defining scorer to be used for cross-validation and hyperparameter tuning¶
- We want to reduce false negatives and will try to maximize "Recall".
- To maximize Recall, we can use Recall as a scorer in cross-validation and hyperparameter tuning.
# Type of scoring used to compare parameter combinations
scorer = metrics.make_scorer(metrics.recall_score)
Model Building with original data¶
Sample Decision Tree model building with original data
models = [] # Empty list to store all the models
# Appending models into the list
models.append(("Logistic Regression", LogisticRegression(random_state=1, n_jobs=-1)))
models.append(("Bagging", BaggingClassifier(random_state=1, n_jobs=-1)))
models.append(("Random forest", RandomForestClassifier(random_state=1, n_jobs=-1)))
models.append(("GBM", GradientBoostingClassifier(random_state=1)))
models.append(("Adaboost", AdaBoostClassifier(random_state=1, algorithm="SAMME")))
models.append(("Xgboost", XGBClassifier(random_state=1, eval_metric="logloss", n_jobs=-1, tree_method='hist')))
models.append(("dtree", DecisionTreeClassifier(random_state=1)))
results1 = [] # Empty list to store all model's CV scores
names = [] # Empty list to store name of the models
# loop through all models to get the mean cross validated score
print("\n" "Cross-Validation performance on training dataset:" "\n")
for name, model in models:
kfold = StratifiedKFold(n_splits=5,
shuffle=True,
random_state=1
)
cv_result = cross_val_score(estimator=model,
X=X_train,
y=y_train,
scoring = scorer,
cv=kfold,
n_jobs=-1
)
results1.append(cv_result)
names.append(name)
print("{}: {}".format(name, cv_result.mean()))
print("\n" "Validation Performance:" "\n")
for name, model in models:
model.fit(X_train, y_train)
scores = recall_score(y_val, model.predict(X_val))
print("{}: {}".format(name, scores))
Cross-Validation performance on training dataset: Logistic Regression: 0.4902481389578163 Bagging: 0.707808105872622 Random forest: 0.7194127377998345 GBM: 0.7220016542597187 Adaboost: 0.5109181141439206 Xgboost: 0.8095368072787427 dtree: 0.7052605459057072 Validation Performance: Logistic Regression: 0.5015015015015015 Bagging: 0.7267267267267268 Random forest: 0.7357357357357357 GBM: 0.7357357357357357 Adaboost: 0.5555555555555556 Xgboost: 0.8288288288288288 dtree: 0.7057057057057057
# Plotting boxplots for CV scores of all models defined above
fig = plt.figure(figsize=(10, 7))
fig.suptitle("Algorithm Comparison")
ax = fig.add_subplot(111)
plt.boxplot(results1)
ax.set_xticklabels(names)
plt.show()
Xgboost is giving the highest recall score followed by GBM, Random Forest, Boosting.
Model Building with Oversampled data¶
print("Before OverSampling, counts of label '1': {}".format(sum(y_train == 1)))
print("Before OverSampling, counts of label '0': {} \n".format(sum(y_train == 0)))
# Synthetic Minority Over Sampling Technique
sm = SMOTE(sampling_strategy=1, k_neighbors=5, random_state=1)
X_train_over, y_train_over = sm.fit_resample(X_train, y_train)
print("After OverSampling, counts of label '1': {}".format(sum(y_train_over == 1)))
print("After OverSampling, counts of label '0': {} \n".format(sum(y_train_over == 0)))
print("After OverSampling, the shape of train_X: {}".format(X_train_over.shape))
print("After OverSampling, the shape of train_y: {} \n".format(y_train_over.shape))
Before OverSampling, counts of label '1': 777 Before OverSampling, counts of label '0': 13223 After OverSampling, counts of label '1': 13223 After OverSampling, counts of label '0': 13223 After OverSampling, the shape of train_X: (26446, 40) After OverSampling, the shape of train_y: (26446,)
models = [] # Empty list to store all the models
# Appending models into the list
models.append(("Logistic Regression", LogisticRegression(random_state=1, n_jobs=-1)))
models.append(("Bagging", BaggingClassifier(random_state=1, n_jobs=-1)))
models.append(("Random forest", RandomForestClassifier(random_state=1, n_jobs=-1)))
models.append(("GBM", GradientBoostingClassifier(random_state=1)))
models.append(("Adaboost", AdaBoostClassifier(random_state=1, algorithm="SAMME")))
models.append(("Xgboost", XGBClassifier(random_state=1, eval_metric="logloss", n_jobs=-1, tree_method='hist')))
models.append(("dtree", DecisionTreeClassifier(random_state=1)))
results_over = [] # Empty list to store all model's CV scores
names = [] # Empty list to store name of the models
# loop through all models to get the mean cross validated score
print("\n" "Cross-Validation performance on over sampled training dataset:" "\n")
for name, model in models:
kfold = StratifiedKFold(n_splits=5,
shuffle=True,
random_state=1
)
cv_result = cross_val_score(estimator=model,
X=X_train_over,
y=y_train_over,
scoring = scorer,
cv=kfold,
n_jobs=-1
)
results_over.append(cv_result)
names.append(name)
print("{}: {}".format(name, cv_result.mean()))
print("\n" "Validation Performance on over sampled training dataset:" "\n")
for name, model in models:
model.fit(X_train_over, y_train_over)
scores = recall_score(y_val, model.predict(X_val))
print("{}: {}".format(name, scores))
Cross-Validation performance on over sampled training dataset: Logistic Regression: 0.8917044404851445 Bagging: 0.9749681555985804 Random forest: 0.9827577795000415 GBM: 0.9329201902370526 Adaboost: 0.8966949600908288 Xgboost: 0.9904713314591799 dtree: 0.970128321355339 Validation Performance on over sampled training dataset: Logistic Regression: 0.8498498498498499 Bagging: 0.8228228228228228 Random forest: 0.8558558558558559 GBM: 0.8768768768768769 Adaboost: 0.8588588588588588 Xgboost: 0.8558558558558559 dtree: 0.7837837837837838
# Plotting boxplots for CV scores of all models defined above
fig = plt.figure(figsize=(10, 7))
fig.suptitle("Algorithm Comparison over sampled training dataset")
ax = fig.add_subplot(111)
plt.boxplot(results_over)
ax.set_xticklabels(names)
plt.show()
- SMOTE improves cross-validation and validation.
- XGBoost, Random Forest and GBM has the most improved recall.
Model Building with Undersampled data¶
# Random undersampler for under sampling the data
rus = RandomUnderSampler(random_state=1, sampling_strategy=1)
X_train_under, y_train_under= rus.fit_resample(X_train, y_train)
print("Before UnderSampling, counts of label '1': {}".format(sum(y_train == 1)))
print("Before UnderSampling, counts of label '0': {} \n".format(sum(y_train == 0)))
print("After UnderSampling, counts of label '1': {}".format(sum(y_train_under == 1)))
print("After UnderSampling, counts of label '0': {} \n".format(sum(y_train_under == 0)))
print("After UnderSampling, the shape of train_X: {}".format(X_train_under.shape))
print("After UnderSampling, the shape of train_y: {} \n".format(y_train_under.shape))
Before UnderSampling, counts of label '1': 777 Before UnderSampling, counts of label '0': 13223 After UnderSampling, counts of label '1': 777 After UnderSampling, counts of label '0': 777 After UnderSampling, the shape of train_X: (1554, 40) After UnderSampling, the shape of train_y: (1554,)
models = [] # Empty list to store all the models
# Appending models into the list
models.append(("Logistic Regression", LogisticRegression(random_state=1, n_jobs=-1)))
models.append(("Bagging", BaggingClassifier(random_state=1, n_jobs=-1)))
models.append(("Random forest", RandomForestClassifier(random_state=1, n_jobs=-1)))
models.append(("GBM", GradientBoostingClassifier(random_state=1)))
models.append(("Adaboost", AdaBoostClassifier(random_state=1, algorithm="SAMME")))
models.append(("Xgboost", XGBClassifier(random_state=1, eval_metric="logloss", n_jobs=-1, tree_method='hist')))
models.append(("dtree", DecisionTreeClassifier(random_state=1)))
results_under = [] # Empty list to store all model's CV scores
names = [] # Empty list to store name of the models
# loop through all models to get the mean cross validated score
print("\n" "Cross-Validation performance on under sampled training dataset:" "\n")
for name, model in models:
kfold = StratifiedKFold(n_splits=5,
shuffle=True,
random_state=1
)
cv_result = cross_val_score(estimator=model,
X=X_train_under,
y=y_train_under,
scoring=scorer,
cv=kfold,
n_jobs=-1
)
results_under.append(cv_result)
names.append(name)
print("{}: {}".format(name, cv_result.mean()))
print("\n" "Validation Performance on under sampled training dataset:" "\n")
for name, model in models:
model.fit(X_train_under, y_train_under)
scores = recall_score(y_val, model.predict(X_val))
print("{}: {}".format(name, scores))
Cross-Validation performance on under sampled training dataset: Logistic Regression: 0.8726220016542598 Bagging: 0.880339123242349 Random forest: 0.9034822167080232 GBM: 0.8932009925558313 Adaboost: 0.8687427626137303 Xgboost: 0.8983457402812242 dtree: 0.8622167080231596 Validation Performance on under sampled training dataset: Logistic Regression: 0.8468468468468469 Bagging: 0.8708708708708709 Random forest: 0.8828828828828829 GBM: 0.8828828828828829 Adaboost: 0.8468468468468469 Xgboost: 0.8828828828828829 dtree: 0.8408408408408409
# Plotting boxplots for CV scores of all models defined above
fig = plt.figure(figsize=(10, 7))
fig.suptitle("Algorithm Comparison under sampled training dataset")
ax = fig.add_subplot(111)
plt.boxplot(results_under)
ax.set_xticklabels(names)
plt.show()
- Under sampling shows improvement but less than SMOTE
XGBoost, Bagging, GBM and Random Forest perform well across all with SMOTE being the best recall score.
HyperparameterTuning¶
Sample Parameter Grids¶
Hyperparameter tuning can take a long time to run, so to avoid that time complexity - you can use the following grids, wherever required.
- For Gradient Boosting:
param_grid = { "n_estimators": np.arange(100,150,25), "learning_rate": [0.2, 0.05, 1], "subsample":[0.5,0.7], "max_features":[0.5,0.7] }
- For Adaboost:
param_grid = { "n_estimators": [100, 150, 200], "learning_rate": [0.2, 0.05], "base_estimator": [DecisionTreeClassifier(max_depth=1, random_state=1), DecisionTreeClassifier(max_depth=2, random_state=1), DecisionTreeClassifier(max_depth=3, random_state=1), ] }
- For Bagging Classifier:
param_grid = { 'max_samples': [0.8,0.9,1], 'max_features': [0.7,0.8,0.9], 'n_estimators' : [30,50,70], }
- For Random Forest:
param_grid = { "n_estimators": [200,250,300], "min_samples_leaf": np.arange(1, 4), "max_features": [np.arange(0.3, 0.6, 0.1),'sqrt'], "max_samples": np.arange(0.4, 0.7, 0.1) }
- For Decision Trees:
param_grid = { 'max_depth': np.arange(2,6), 'min_samples_leaf': [1, 4, 7], 'max_leaf_nodes' : [10, 15], 'min_impurity_decrease': [0.0001,0.001] }
- For Logistic Regression:
param_grid = {'C': np.arange(0.1,1.1,0.1)}
- For XGBoost:
param_grid={ 'n_estimators': [150, 200, 250], 'scale_pos_weight': [5,10], 'learning_rate': [0.1,0.2], 'gamma': [0,3,5], 'subsample': [0.8,0.9] }
Tuning AdaBoost using oversampled data¶
# defining model
Model = AdaBoostClassifier(random_state=1, algorithm="SAMME")
# Parameter grid to pass in RandomSearchCV
param_grid = {
"n_estimators": [100, 150, 200],
"learning_rate": [0.2, 0.05],
"estimator": [DecisionTreeClassifier(max_depth=1, random_state=1),
DecisionTreeClassifier(max_depth=2, random_state=1),
DecisionTreeClassifier(max_depth=3, random_state=1),
]
}
#Calling RandomizedSearchCV
randomized_cv = RandomizedSearchCV(estimator=Model,
param_distributions=param_grid,
n_iter=50, n_jobs = -1,
scoring=scorer,
cv=5,
random_state=1
)
#Fitting parameters in RandomizedSearchCV
randomized_cv.fit(X_train_over,y_train_over)
print("Best parameters are {} with CV score={}:" .format(randomized_cv.best_params_,randomized_cv.best_score_))
Best parameters are {'n_estimators': 200, 'learning_rate': 0.2, 'estimator': DecisionTreeClassifier(max_depth=3, random_state=1)} with CV score=0.9215760047359074:
# Creating new pipeline with best parameters
tuned_ada = AdaBoostClassifier(n_estimators= 200,
learning_rate= 0.2,
estimator= DecisionTreeClassifier(max_depth=3, random_state=1),
random_state=1
)
tuned_ada.fit(X_train_over, y_train_over)
AdaBoostClassifier(estimator=DecisionTreeClassifier(max_depth=3,
random_state=1),
learning_rate=0.2, n_estimators=200, random_state=1)In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook. On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
AdaBoostClassifier(estimator=DecisionTreeClassifier(max_depth=3,
random_state=1),
learning_rate=0.2, n_estimators=200, random_state=1)DecisionTreeClassifier(max_depth=3, random_state=1)
DecisionTreeClassifier(max_depth=3, random_state=1)
ada_train_perf = model_performance_classification_sklearn(tuned_ada, X_train_over, y_train_over)
ada_train_perf
| Accuracy | Recall | Precision | F1 | |
|---|---|---|---|---|
| 0 | 0.994 | 0.991 | 0.997 | 0.994 |
ada_val_perf = model_performance_classification_sklearn(tuned_ada, X_val, y_val)
ada_val_perf
| Accuracy | Recall | Precision | F1 | |
|---|---|---|---|---|
| 0 | 0.980 | 0.847 | 0.803 | 0.825 |
Tuning Random forest using under sampled data¶
# defining model
Model = RandomForestClassifier(random_state=1, n_jobs=-1)
# Parameter grid to pass in RandomSearchCV
param_grid = {
"n_estimators": [200,250,300],
"min_samples_leaf": np.arange(1, 4),
"max_features": [np.arange(0.3, 0.6, 0.1),'sqrt'],
"max_samples": np.arange(0.4, 0.7, 0.1)}
#Calling RandomizedSearchCV
randomized_cv = RandomizedSearchCV(estimator=Model,
param_distributions=param_grid,
n_iter=50, n_jobs=-1,
scoring=scorer,
cv=5,
random_state=1
)
#Fitting parameters in RandomizedSearchCV
randomized_cv.fit(X_train_under, y_train_under)
print("Best parameters are {} with CV score={}:" .format(randomized_cv.best_params_,randomized_cv.best_score_))
Best parameters are {'n_estimators': 300, 'min_samples_leaf': 1, 'max_samples': 0.6, 'max_features': 'sqrt'} with CV score=0.9047477253928868:
# Creating new pipeline with best parameters
tuned_rf2 = RandomForestClassifier(max_features='sqrt',
random_state=1,
max_samples=0.6,
n_estimators=300,
min_samples_leaf=1
)
tuned_rf2.fit(X_train_under, y_train_under)
RandomForestClassifier(max_samples=0.6, n_estimators=300, random_state=1)In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
RandomForestClassifier(max_samples=0.6, n_estimators=300, random_state=1)
rf2_train_perf = model_performance_classification_sklearn(tuned_rf2, X_train_under, y_train_under)
rf2_train_perf
| Accuracy | Recall | Precision | F1 | |
|---|---|---|---|---|
| 0 | 0.988 | 0.978 | 0.999 | 0.988 |
rf2_val_perf = model_performance_classification_sklearn(tuned_rf2, X_val, y_val)
rf2_val_perf
| Accuracy | Recall | Precision | F1 | |
|---|---|---|---|---|
| 0 | 0.934 | 0.880 | 0.451 | 0.596 |
Tuning Gradient Boosting using oversampled data¶
# defining model
Model = GradientBoostingClassifier(random_state=1)
#Parameter grid to pass in RandomSearchCV
param_grid={"n_estimators": np.arange(100,150,25), "learning_rate": [0.2, 0.05, 1], "subsample":[0.5,0.7], "max_features":[0.5,0.7]}
#Calling RandomizedSearchCV
randomized_cv = RandomizedSearchCV(estimator=Model,
param_distributions=param_grid,
scoring=scorer,
n_iter=50,
n_jobs=-1,
cv=5,
random_state=1
)
#Fitting parameters in RandomizedSearchCV
randomized_cv.fit(X_train_over, y_train_over)
print("Best parameters are {} with CV score={}:" .format(randomized_cv.best_params_,randomized_cv.best_score_))
Best parameters are {'subsample': 0.7, 'n_estimators': 125, 'max_features': 0.5, 'learning_rate': 1} with CV score=0.9720937229208193:
# Creating new pipeline with best parameters
tuned_gbm = GradientBoostingClassifier(max_features=0.5,
random_state=1,
learning_rate=1,
n_estimators=125,
subsample=0.7
)
tuned_gbm.fit(X_train_over, y_train_over)
GradientBoostingClassifier(learning_rate=1, max_features=0.5, n_estimators=125,
random_state=1, subsample=0.7)In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook. On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
GradientBoostingClassifier(learning_rate=1, max_features=0.5, n_estimators=125,
random_state=1, subsample=0.7)gbm_train_perf = model_performance_classification_sklearn(tuned_gbm, X_train_over, y_train_over)
gbm_train_perf
| Accuracy | Recall | Precision | F1 | |
|---|---|---|---|---|
| 0 | 0.986 | 0.985 | 0.988 | 0.986 |
gbm_val_perf = model_performance_classification_sklearn(tuned_gbm, X_val, y_val)
gbm_val_perf
| Accuracy | Recall | Precision | F1 | |
|---|---|---|---|---|
| 0 | 0.958 | 0.826 | 0.588 | 0.687 |
Tuning XGBoost using over sampled data¶
# defining model
Model = XGBClassifier(random_state=1,eval_metric='logloss')
#Parameter grid to pass in RandomSearchCV
param_grid={'n_estimators':[150,200,250],
'scale_pos_weight':[5,10],
'learning_rate':[0.1,0.2],
'gamma':[0,3,5],
'subsample':[0.8,0.9]}
#Calling RandomizedSearchCV
randomized_cv = RandomizedSearchCV(estimator=Model,
param_distributions=param_grid,
n_iter=50,
n_jobs=-1,
scoring=scorer,
cv=5,
random_state=1
)
#Fitting parameters in RandomizedSearchCV
randomized_cv.fit(X_train_over, y_train_over)
print("Best parameters are {} with CV score={}:" .format(randomized_cv.best_params_,randomized_cv.best_score_))
Best parameters are {'subsample': 0.8, 'scale_pos_weight': 10, 'n_estimators': 250, 'learning_rate': 0.1, 'gamma': 0} with CV score=0.9966722243035557:
xgb2 = XGBClassifier(random_state=1,
eval_metric='logloss',
subsample=0.8,
scale_pos_weight=10,
n_estimators=250,
learning_rate=0.1,
gamma=0,
n_jobs=-1,
)
xgb2.fit(X_train_over, y_train_over)
XGBClassifier(base_score=None, booster=None, callbacks=None,
colsample_bylevel=None, colsample_bynode=None,
colsample_bytree=None, device=None, early_stopping_rounds=None,
enable_categorical=False, eval_metric='logloss',
feature_types=None, gamma=0, grow_policy=None,
importance_type=None, interaction_constraints=None,
learning_rate=0.1, max_bin=None, max_cat_threshold=None,
max_cat_to_onehot=None, max_delta_step=None, max_depth=None,
max_leaves=None, min_child_weight=None, missing=nan,
monotone_constraints=None, multi_strategy=None, n_estimators=250,
n_jobs=-1, num_parallel_tree=None, random_state=1, ...)In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook. On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
XGBClassifier(base_score=None, booster=None, callbacks=None,
colsample_bylevel=None, colsample_bynode=None,
colsample_bytree=None, device=None, early_stopping_rounds=None,
enable_categorical=False, eval_metric='logloss',
feature_types=None, gamma=0, grow_policy=None,
importance_type=None, interaction_constraints=None,
learning_rate=0.1, max_bin=None, max_cat_threshold=None,
max_cat_to_onehot=None, max_delta_step=None, max_depth=None,
max_leaves=None, min_child_weight=None, missing=nan,
monotone_constraints=None, multi_strategy=None, n_estimators=250,
n_jobs=-1, num_parallel_tree=None, random_state=1, ...)xgb2_train_perf = model_performance_classification_sklearn(xgb2, X_train_over, y_train_over)
xgb2_train_perf
| Accuracy | Recall | Precision | F1 | |
|---|---|---|---|---|
| 0 | 1.000 | 1.000 | 0.999 | 1.000 |
xgb2_val_perf = model_performance_classification_sklearn(xgb2, X_val, y_val)
xgb2_val_perf
| Accuracy | Recall | Precision | F1 | |
|---|---|---|---|---|
| 0 | 0.982 | 0.880 | 0.807 | 0.842 |
Model performance comparison and choosing the final model¶
# training performance comparison
models_train_comp_df = pd.concat(
[
gbm_train_perf.T,
ada_train_perf.T,
rf2_train_perf.T,
xgb2_train_perf.T,
],
axis=1,
)
models_train_comp_df.columns = [
"Gradient Boosting tuned with oversampled data",
"AdaBoost classifier tuned with oversampled data",
"Random forest tuned with undersampled data",
"XGBoost tuned with oversampled data",
]
print("Training performance comparison:")
models_train_comp_df
Training performance comparison:
| Gradient Boosting tuned with oversampled data | AdaBoost classifier tuned with oversampled data | Random forest tuned with undersampled data | XGBoost tuned with oversampled data | |
|---|---|---|---|---|
| Accuracy | 0.986 | 0.994 | 0.988 | 1.000 |
| Recall | 0.985 | 0.991 | 0.978 | 1.000 |
| Precision | 0.988 | 0.997 | 0.999 | 0.999 |
| F1 | 0.986 | 0.994 | 0.988 | 1.000 |
- XGBoost has near perfect matching but that suggests it might be over fitting and will likely not generalize well
- AdaBoost with oversampled data has near perfect results with high accuracy, precision, recall and f1-score.
- AdaBoost is less likely to be overfitting given the slight differences in the metrics.
ada_test = model_performance_classification_sklearn(tuned_ada, X_test, y_test)
ada_test
| Accuracy | Recall | Precision | F1 | |
|---|---|---|---|---|
| 0 | 0.976 | 0.830 | 0.765 | 0.796 |
Feature Importances¶
feature_names = X_train.columns
importances = tuned_ada.feature_importances_
indices = np.argsort(importances)
plt.figure(figsize=(12, 12))
plt.title("Feature Importances")
plt.barh(range(len(indices)), importances[indices], color="violet", align="center")
plt.yticks(range(len(indices)), [feature_names[i] for i in indices])
plt.xlabel("Relative Importance")
plt.show()
Pipelines to build the final model¶
Pipeline_model = Pipeline(
steps=[("imputer", SimpleImputer(strategy="median")),
("AdaBoost Classifier", AdaBoostClassifier(random_state=1,
n_estimators= 200,
learning_rate= 0.2,
estimator= DecisionTreeClassifier(max_depth=3, random_state=1),
))
]
)
# Separating target variable and other variables
X1 = data.drop(columns="Target")
Y1 = data["Target"]
# Since we already have a separate test set, we don't need to divide data into train and test
X_test1 = df_test.drop(['Target'], axis=1)
y_test1 = df_test['Target']
imputer = SimpleImputer(strategy="median")
X1 = imputer.fit_transform(X1)
sm = SMOTE(sampling_strategy=1, k_neighbors=5, random_state=1)
X_over1, y_over1 = sm.fit_resample(X1, Y1)
Pipeline_model.fit(X_over1, y_over1)
Pipeline(steps=[('imputer', SimpleImputer(strategy='median')),
('AdaBoost Classifier',
AdaBoostClassifier(estimator=DecisionTreeClassifier(max_depth=3,
random_state=1),
learning_rate=0.2, n_estimators=200,
random_state=1))])In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook. On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
Pipeline(steps=[('imputer', SimpleImputer(strategy='median')),
('AdaBoost Classifier',
AdaBoostClassifier(estimator=DecisionTreeClassifier(max_depth=3,
random_state=1),
learning_rate=0.2, n_estimators=200,
random_state=1))])SimpleImputer(strategy='median')
AdaBoostClassifier(estimator=DecisionTreeClassifier(max_depth=3,
random_state=1),
learning_rate=0.2, n_estimators=200, random_state=1)DecisionTreeClassifier(max_depth=3, random_state=1)
DecisionTreeClassifier(max_depth=3, random_state=1)
Pipeline_model_test = model_performance_classification_sklearn(Pipeline_model,X_test1, y_test1)
Pipeline_model_test
| Accuracy | Recall | Precision | F1 | |
|---|---|---|---|---|
| 0 | 0.978 | 0.851 | 0.774 | 0.811 |
Business Insights and Conclusions¶
- The AdaBoost classifier with over sampling has the best performance.
- V30, V18, V12 are the most important features.
- The model does a good job predicting if a wind turbine will fail.
- Accuracy - 0.978:
- The model correctly predicted 97.8% in test set.
- The high accuracy indicates the model is performing well overall.
- Recall - 0.851
- The model correctly identified 85% of the true positives in the test set.
- The model is good at capturing the majority of the positives but missed 15%
- Precision - 0.774:
- Of all the actual predicted positives, 77% were true positives.
- Precision is slightly lower than recall so it is occasionally has false positives.
- F1 - 0.811:
- The model strikes a good balance between recall and precision but leans a little toward recall meaning it prioritizes identifying positive cases over avoiding false positives.
- Accuracy - 0.978:
Overall Conclusion:
- The model performs well at predicting wind turbine failures and shows a slight bias towards recall, meaning it prioritizes finding as many potential failures as possible, even if it occasionally flags turbines that won’t fail.
- Recommendation: In this case, prioritizing recall is more cost-effective, as missing a potential turbine failure can lead to significant downtime and repair costs. The occasional false positive (unnecessary maintenance) is likely less costly than missing a failure.