SOFTWARE ENGINEER & SDET

Jennifer Montgomery

Backend · full-stack · quality engineering

COURSEWORK / EASYVISA

Ensemble techniques for classification

How do tree ensembles compare when classifying a historical case outcome?

← Coursework on resume

Completed UT Austin postgraduate coursework using a supplied scenario, dataset, and starter notebook. The charts below come from my completed notebook.

ENSEMBLE TECHNIQUES

EasyVisa

COURSE FOCUS

Bagging, boosting, model comparison, and tuning

LIBRARIES USED

pandas, NumPy, scikit-learn, XGBoost, Seaborn, Matplotlib

The supplied data

The course provided historical labor-certification cases with a certified or denied label. Features described the employer and position, including company size, establishment year, prevailing wage and wage unit, plus applicant education, experience, and geographic fields. These are historical labels in a classroom dataset, not a basis for deciding an individual case.

What I did

  • Checked the outcome distribution and explored relationships between case attributes and the recorded status.
  • Compared decision trees, random forests, bagging, AdaBoost, gradient boosting, XGBoost, and stacking using classification measures.
  • Tuned promising ensembles and compared training with held-out performance to look for overfitting.
Bar chart of certified and denied labels in the supplied EasyVisa case data.
From the notebook: certified cases outnumbered denied cases, making label balance relevant to model evaluation. Open chart ↗
Feature-importance bars from the tuned XGBoost classroom model, including education, wage unit, experience, and geographic variables.
From the notebook: model importance indicates what the fitted model used, not which factors should determine an immigration outcome. Open chart ↗

Finding in the course comparison

Tuned gradient boosting and XGBoost were among the stronger held-out models in the notebook, with F1 scores around 0.82. Some tree models fit the training data much better than the test data, underscoring the need to compare both.

Learning reinforced

The project made ensemble methods and overfitting tangible. Because the features include education and geography and the outcome affects people, a classroom score and feature-importance chart cannot justify real eligibility decisions without legal, fairness, and data-quality review.