Completed UT Austin postgraduate coursework using a supplied scenario, dataset, and starter notebook. The charts below come from my completed notebook.
ReneWind
Imbalanced classification and hyperparameter tuning
pandas, NumPy, scikit-learn, imbalanced-learn, XGBoost, Seaborn, Matplotlib
The supplied data
The supplied predictive-maintenance records contained 40 anonymized sensor-derived predictors and a rare generator-failure label, with separate training and test data. The scenario gave different costs for missed failures, unnecessary inspections, and planned repairs, so accuracy alone was an incomplete measure.
What I did
- Checked the class imbalance, prepared missing values, and compared original, oversampled, and undersampled training data.
- Compared logistic regression, decision trees, and ensemble models with cross-validation.
- Tuned candidate models with randomized search, assembled a final pipeline, and evaluated the selection on the supplied held-out test set.


Finding on the course test split
I selected AdaBoost trained with oversampling. It reached 0.851 recall and 0.774 precision on the held-out course test set, emphasizing detection of failures while accepting some false alerts.
Learning reinforced
This project made model selection depend on the cost of each error, not accuracy alone. It also applied resampling, cross-validation, hyperparameter search, and pipeline construction to an imbalanced problem.
Original notebook export
View the complete HTML report, including code, outputs, and the original written analysis.