EasyVisa Project¶

Context:¶

Business communities in the United States are facing high demand for human resources, but one of the constant challenges is identifying and attracting the right talent, which is perhaps the most important element in remaining competitive. Companies in the United States look for hard-working, talented, and qualified individuals both locally as well as abroad.

The Immigration and Nationality Act (INA) of the US permits foreign workers to come to the United States to work on either a temporary or permanent basis. The act also protects US workers against adverse impacts on their wages or working conditions by ensuring US employers' compliance with statutory requirements when they hire foreign workers to fill workforce shortages. The immigration programs are administered by the Office of Foreign Labor Certification (OFLC).

OFLC processes job certification applications for employers seeking to bring foreign workers into the United States and grants certifications in those cases where employers can demonstrate that there are not sufficient US workers available to perform the work at wages that meet or exceed the wage paid for the occupation in the area of intended employment.

Objective:¶

In FY 2016, the OFLC processed 775,979 employer applications for 1,699,957 positions for temporary and permanent labor certifications. This was a nine percent increase in the overall number of processed applications from the previous year. The process of reviewing every case is becoming a tedious task as the number of applicants is increasing every year.

The increasing number of applicants every year calls for a Machine Learning based solution that can help in shortlisting the candidates having higher chances of VISA approval. OFLC has hired your firm EasyVisa for data-driven solutions. You as a data scientist have to analyze the data provided and, with the help of a classification model:

  • Facilitate the process of visa approvals.
  • Recommend a suitable profile for the applicants for whom the visa should be certified or denied based on the drivers that significantly influence the case status.

Data Description¶

The data contains the different attributes of the employee and the employer. The detailed data dictionary is given below.

  • case_id: ID of each visa application
  • continent: Information of continent the employee
  • education_of_employee: Information of education of the employee
  • has_job_experience: Does the employee has any job experience? Y= Yes; N = No
  • requires_job_training: Does the employee require any job training? Y = Yes; N = No
  • no_of_employees: Number of employees in the employer's company
  • yr_of_estab: Year in which the employer's company was established
  • region_of_employment: Information of foreign worker's intended region of employment in the US.
  • prevailing_wage: Average wage paid to similarly employed workers in a specific occupation in the area of intended employment. The purpose of the prevailing wage is to ensure that the foreign worker is not underpaid compared to other workers offering the same or similar service in the same area of employment.
  • unit_of_wage: Unit of prevailing wage. Values include Hourly, Weekly, Monthly, and Yearly.
  • full_time_position: Is the position of work full-time? Y = Full Time Position; N = Part Time Position
  • case_status: Flag indicating if the Visa was certified or denied

Importing necessary libraries and data¶

In [1]:
# Installing the libraries with the specified version.
#!pip install numpy==1.25.2 pandas==1.5.3 scikit-learn==1.2.2 matplotlib==3.7.1 seaborn==0.13.1 xgboost==2.0.3

Note: After running the above cell, kindly restart the notebook kernel and run all cells sequentially from the start again.

In [2]:
import warnings

warnings.filterwarnings("ignore")

# Libraries to help with reading and manipulating data
import numpy as np
import pandas as pd

# Library to split data
from sklearn.model_selection import train_test_split

# libaries to help with data visualization
import matplotlib.pyplot as plt
import seaborn as sns

# Removes the limit for the number of displayed columns
pd.set_option("display.max_columns", None)
# Sets the limit for the number of displayed rows
pd.set_option("display.max_rows", 100)


# Libraries different ensemble classifiers
from sklearn.ensemble import (
    BaggingClassifier,
    RandomForestClassifier,
    AdaBoostClassifier,
    GradientBoostingClassifier,
    StackingClassifier,
)

from xgboost import XGBClassifier
from sklearn.tree import DecisionTreeClassifier

# Libraries to get different metric scores
from sklearn import metrics
from sklearn.metrics import (
    confusion_matrix,
    accuracy_score,
    precision_score,
    recall_score,
    f1_score,
)

# To tune different models
from sklearn.model_selection import GridSearchCV
from sklearn.experimental import enable_halving_search_cv

tab20_blue = '#1f77b4'
tab20_orange = '#ff7f0e'
tab20_green = '#2ca02c'
tab20_red = '#d62728'
tab20_puple = '#9467bd'
tab20_pink = '#e377c2'
tab20_grey = '#7f7f7f'
tab20_yellow = '#bcbd22'
tab20_teal = '#17becf'
sns.set_style("white")
In [3]:
visa = pd.read_csv('EasyVisa.csv')
In [4]:
data = visa.copy()

Data Overview¶

  • Observations
  • Sanity checks
In [5]:
data.shape
Out[5]:
(25480, 12)
In [6]:
data.head()
Out[6]:
case_id continent education_of_employee has_job_experience requires_job_training no_of_employees yr_of_estab region_of_employment prevailing_wage unit_of_wage full_time_position case_status
0 EZYV01 Asia High School N N 14513 2007 West 592.2029 Hour Y Denied
1 EZYV02 Asia Master's Y N 2412 2002 Northeast 83425.6500 Year Y Certified
2 EZYV03 Asia Bachelor's N Y 44444 2008 West 122996.8600 Year Y Denied
3 EZYV04 Asia Bachelor's N N 98 1897 West 83434.0300 Year Y Denied
4 EZYV05 Africa Master's Y N 1082 2005 South 149907.3900 Year Y Certified
In [7]:
data.tail()
Out[7]:
case_id continent education_of_employee has_job_experience requires_job_training no_of_employees yr_of_estab region_of_employment prevailing_wage unit_of_wage full_time_position case_status
25475 EZYV25476 Asia Bachelor's Y Y 2601 2008 South 77092.57 Year Y Certified
25476 EZYV25477 Asia High School Y N 3274 2006 Northeast 279174.79 Year Y Certified
25477 EZYV25478 Asia Master's Y N 1121 1910 South 146298.85 Year N Certified
25478 EZYV25479 Asia Master's Y Y 1918 1887 West 86154.77 Year Y Certified
25479 EZYV25480 Asia Bachelor's Y N 3195 1960 Midwest 70876.91 Year Y Certified
In [8]:
data.info()
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 25480 entries, 0 to 25479
Data columns (total 12 columns):
 #   Column                 Non-Null Count  Dtype  
---  ------                 --------------  -----  
 0   case_id                25480 non-null  object 
 1   continent              25480 non-null  object 
 2   education_of_employee  25480 non-null  object 
 3   has_job_experience     25480 non-null  object 
 4   requires_job_training  25480 non-null  object 
 5   no_of_employees        25480 non-null  int64  
 6   yr_of_estab            25480 non-null  int64  
 7   region_of_employment   25480 non-null  object 
 8   prevailing_wage        25480 non-null  float64
 9   unit_of_wage           25480 non-null  object 
 10  full_time_position     25480 non-null  object 
 11  case_status            25480 non-null  object 
dtypes: float64(1), int64(2), object(9)
memory usage: 2.3+ MB
In [9]:
data.describe().T
Out[9]:
count mean std min 25% 50% 75% max
no_of_employees 25480.0 5667.043210 22877.928848 -26.0000 1022.00 2109.00 3504.0000 602069.00
yr_of_estab 25480.0 1979.409929 42.366929 1800.0000 1976.00 1997.00 2005.0000 2016.00
prevailing_wage 25480.0 74455.814592 52815.942327 2.1367 34015.48 70308.21 107735.5125 319210.27
In [10]:
# Making a list of all catrgorical variables
cat_col = list(data.select_dtypes("object").columns)

# Printing number of count of each unique value in each column
for column in cat_col:
    print(column, data[column].nunique())
    print(data[column].value_counts())
    print("-" * 50)
case_id 25480
EZYV01       1
EZYV16995    1
EZYV16993    1
EZYV16992    1
EZYV16991    1
            ..
EZYV8492     1
EZYV8491     1
EZYV8490     1
EZYV8489     1
EZYV25480    1
Name: case_id, Length: 25480, dtype: int64
--------------------------------------------------
continent 6
Asia             16861
Europe            3732
North America     3292
South America      852
Africa             551
Oceania            192
Name: continent, dtype: int64
--------------------------------------------------
education_of_employee 4
Bachelor's     10234
Master's        9634
High School     3420
Doctorate       2192
Name: education_of_employee, dtype: int64
--------------------------------------------------
has_job_experience 2
Y    14802
N    10678
Name: has_job_experience, dtype: int64
--------------------------------------------------
requires_job_training 2
N    22525
Y     2955
Name: requires_job_training, dtype: int64
--------------------------------------------------
region_of_employment 5
Northeast    7195
South        7017
West         6586
Midwest      4307
Island        375
Name: region_of_employment, dtype: int64
--------------------------------------------------
unit_of_wage 4
Year     22962
Hour      2157
Week       272
Month       89
Name: unit_of_wage, dtype: int64
--------------------------------------------------
full_time_position 2
Y    22773
N     2707
Name: full_time_position, dtype: int64
--------------------------------------------------
case_status 2
Certified    17018
Denied        8462
Name: case_status, dtype: int64
--------------------------------------------------
In [11]:
data.drop('case_id', axis=1, inplace=True)
In [12]:
cat_col.remove('case_id')
for column in cat_col:
    data[column] = data[column].astype('category')
In [13]:
data.info()
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 25480 entries, 0 to 25479
Data columns (total 11 columns):
 #   Column                 Non-Null Count  Dtype   
---  ------                 --------------  -----   
 0   continent              25480 non-null  category
 1   education_of_employee  25480 non-null  category
 2   has_job_experience     25480 non-null  category
 3   requires_job_training  25480 non-null  category
 4   no_of_employees        25480 non-null  int64   
 5   yr_of_estab            25480 non-null  int64   
 6   region_of_employment   25480 non-null  category
 7   prevailing_wage        25480 non-null  float64 
 8   unit_of_wage           25480 non-null  category
 9   full_time_position     25480 non-null  category
 10  case_status            25480 non-null  category
dtypes: category(8), float64(1), int64(2)
memory usage: 797.7 KB
In [14]:
data.describe(include='category').T
Out[14]:
count unique top freq
continent 25480 6 Asia 16861
education_of_employee 25480 4 Bachelor's 10234
has_job_experience 25480 2 Y 14802
requires_job_training 25480 2 N 22525
region_of_employment 25480 5 Northeast 7195
unit_of_wage 25480 4 Year 22962
full_time_position 25480 2 Y 22773
case_status 25480 2 Certified 17018
In [15]:
data.isnull().sum()
Out[15]:
continent                0
education_of_employee    0
has_job_experience       0
requires_job_training    0
no_of_employees          0
yr_of_estab              0
region_of_employment     0
prevailing_wage          0
unit_of_wage             0
full_time_position       0
case_status              0
dtype: int64
In [16]:
data.duplicated().sum()
Out[16]:
0

Exploratory Data Analysis (EDA)¶

  • EDA is an important part of any project involving data.
  • It is important to investigate and understand the data better before building a model with it.
  • A few questions have been mentioned below which will help you approach the analysis in the right manner and generate insights from the data.
  • A thorough analysis of the data, in addition to the questions mentioned below, should be done.
In [17]:
def histogram_and_boxplot(data, feature, figsize=(15, 5), kde=False, bins=None, box_color=tab20_orange,
                          hist_color=tab20_blue, showmeans=True, title_fontsize=16, xlabel_fontsize=14,
                          ylabel_fontsize=14, rotation=0, fontsize=14, labelsize=12):
    """
    Boxplot and histogram combined

    data: dataframe
    feature: dataframe column
    figsize: size of figure (default (10,5))
    kde: whether to show the density curve (default False)
    bins: number of bins for histogram (default None)
    box_color: color of boxplot (default "tab20_orange")
    hist_color: color of histogram (default "tab20_blue")
    title_fontsize: font size for the title (default 16)
    xlabel_fontsize: font size for the x-axis label (default 16)
    ylabel_fontsize: font size for the y-axis label (default 16)
    rotation: rotation of x-axis labels (default 0 degrees)
    fontsize: font size of the axis tick labels (default 16)
    labelsize: font size of the annotation labels (default 16)
    """
    fig, (ax_box, ax_hist) = plt.subplots(
        nrows=2,
        sharex=True,
        gridspec_kw={"height_ratios": (0.25, 0.75)},
        figsize=figsize,
    )

    sns.boxplot(data=data, x=feature, ax=ax_box, showmeans=showmeans)

    sns.histplot(data=data, x=feature, kde=kde, ax=ax_hist, bins=bins
                 ) if bins else sns.histplot(data=data, x=feature, kde=kde, ax=ax_hist)

    mean = data[feature].mean()
    median = data[feature].median()
    ax_hist.axvline(mean, color=tab20_green, linestyle="--", label=f'Mean: {mean:.2f}')
    ax_hist.axvline(median, color=tab20_orange, linestyle="-", label=f'Median: {median:.2f}')
    ax_hist.legend(fontsize=labelsize)
    ax_box.set(title=f'Distribution of {" ".join(feature.split("_")).lower()}', xlabel='', ylabel='')

    ax_box.tick_params(axis='x', rotation=rotation, labelsize=fontsize)
    ax_box.tick_params(axis='y', labelsize=fontsize)
    ax_hist.set_xlabel(feature, fontsize=xlabel_fontsize)
    ax_hist.set_ylabel('Count', fontsize=ylabel_fontsize)
    ax_hist.tick_params(axis='x', rotation=rotation, labelsize=fontsize)
    ax_hist.tick_params(axis='y', labelsize=fontsize)
    plt.show()
In [18]:
def labeled_barplot(data, feature, perc=True, n=None, palette="tab10", figsize=(5, 5),
                    rotation=90, fontsize=14, labelsize=12, title_fontsize=16,
                    xlabel_fontsize=14, ylabel_fontsize=14):
    """
    Barplot with percentage at the top

    data: dataframe
    feature: dataframe column
    perc: whether to display percentages instead of count (default is False)
    n: displays the top n category levels (default is None, i.e., display all levels)
    """

    total = len(data[feature])
    unique_count = data[feature].nunique()

    temp_data = data.copy()

    if unique_count > 31:
        bins = np.linspace(temp_data[feature].min(), temp_data[feature].max(), 11)
        bins = np.round(bins).astype(int)
        temp_data[feature + '_binned'] = pd.cut(temp_data[feature], bins=bins, include_lowest=True)
        temp_data[feature + '_binned'] = temp_data[feature + '_binned'].apply(lambda x: f'{int(x.left)} - {int(x.right)}')
        feature = feature + '_binned'
        unique_count = temp_data[feature].nunique()

    plot_count = unique_count if n is None else min(n, unique_count)
    plt.figure(figsize=(max(plot_count + 2, figsize[0]), figsize[1]))
    plt.xticks(rotation=rotation, fontsize=fontsize)
    order = temp_data[feature].sort_values().unique()[:plot_count]
    ax = sns.countplot(
        data=temp_data,
        x=feature,
        order=order,
        palette=palette,
        hue='case_status'
    )

    for p in ax.patches:
        if p.get_height() == 0:
            continue
        if perc:
            label = "{:.1f}%".format(100 * p.get_height() / total)
        else:
            label = p.get_height()

        x = p.get_x() + p.get_width() / 2
        y = p.get_height()

        ax.annotate(
            label,
            (x, y),
            ha="center",
            va="center",
            size=labelsize,
            xytext=(0, 5),
            textcoords="offset points"
        )

    ax.set_title(f'{" ".join(feature.split("_")).lower()}', fontsize=title_fontsize)
    ax.set_xlabel(feature, fontsize=xlabel_fontsize)
    ax.set_ylabel('Count' if not perc else 'Percentage', fontsize=ylabel_fontsize)
    ax.tick_params(axis='x', rotation=rotation, labelsize=fontsize)
    ax.tick_params(axis='y', labelsize=fontsize)

    plt.show()
In [19]:
def get_variables_any_dataset(data):
    continuous_cols = data.select_dtypes(include=['float64']).columns
    int_cols = data.select_dtypes(include=['int64']).columns
    int_cols_with_many_uniques = int_cols[data[int_cols].nunique() > 31]
    continuous_cols = continuous_cols.union(int_cols_with_many_uniques).to_list()
    discreet_cols = data.select_dtypes(include=['int64']).columns.to_list()
    category_cols = data.select_dtypes(include=['category']).columns.to_list()
    bivariate_analysis_cols = [col for col in discreet_cols if col not in ['no_of_previous_bookings_not_canceled', 'lead_time']]
    return continuous_cols, discreet_cols, category_cols, bivariate_analysis_cols
In [20]:
continuous_cols, discreet_cols, category_cols, bivariate_analysis_cols = get_variables_any_dataset(data)
print('Continuous Columns:', continuous_cols)
print('Discreet Columns:', discreet_cols)
print('Category Columns:', category_cols)
print('Bivariate Analysis Columns:', bivariate_analysis_cols)
Continuous Columns: ['no_of_employees', 'prevailing_wage', 'yr_of_estab']
Discreet Columns: ['no_of_employees', 'yr_of_estab']
Category Columns: ['continent', 'education_of_employee', 'has_job_experience', 'requires_job_training', 'region_of_employment', 'unit_of_wage', 'full_time_position', 'case_status']
Bivariate Analysis Columns: ['no_of_employees', 'yr_of_estab']
In [21]:
for feature in continuous_cols:
    histogram_and_boxplot(data, feature)
No description has been provided for this image
No description has been provided for this image
No description has been provided for this image
In [22]:
for feature in discreet_cols:
    labeled_barplot(data, feature)
No description has been provided for this image
No description has been provided for this image
In [23]:
for feature in category_cols:
    labeled_barplot(data, feature)
No description has been provided for this image
No description has been provided for this image
No description has been provided for this image
No description has been provided for this image
No description has been provided for this image
No description has been provided for this image
No description has been provided for this image
No description has been provided for this image
In [24]:
plt.figure(figsize=(10, 5))
sns.boxplot(data=data, x="region_of_employment", y="prevailing_wage", hue='case_status')
plt.show()
No description has been provided for this image
In [25]:
plt.figure(figsize=(10, 5))
sns.boxplot(data=data, x="unit_of_wage", y="prevailing_wage", hue='case_status')
plt.show()
No description has been provided for this image
In [26]:
plt.figure(figsize=(10, 5))
sns.boxplot(data=data, x="continent", y="prevailing_wage", hue='case_status')
plt.show()
No description has been provided for this image
In [27]:
plt.figure(figsize=(10, 5))
sns.boxplot(data=data, x="education_of_employee", y="prevailing_wage", hue='case_status')
plt.show()
No description has been provided for this image
In [28]:
plt.figure(figsize=(10, 5))
sns.boxplot(data=data, x="full_time_position", y="prevailing_wage", hue='case_status')
plt.show()
No description has been provided for this image
In [29]:
plt.figure(figsize=(10, 5))
sns.boxplot(data=data, x="has_job_experience", y="prevailing_wage", hue='case_status')
plt.show()
No description has been provided for this image
In [30]:
plt.figure(figsize=(10, 5))
sns.boxplot(data=data, x="requires_job_training", y="prevailing_wage", hue='case_status')
plt.show()
No description has been provided for this image

Observations:

  • The higher level of education the more likely the visa is certified.
  • The most visa applications are coming from Asia with almost 3x the number of total applications and certifications
  • The percentage of cases certified is much higher among those that have job experience.
  • Cases paid yearly are far mor common and also have far more likely to be certified.
  • The median prevailing wage is almost always higher for certified cases across all categories including education, continent, region of employment.

Helper Methods¶

In [31]:
def confusion_matrix_sklearn(model, predictors, target):
    """
    To plot the confusion_matrix with percentages

    model: classifier
    predictors: independent variables
    target: dependent variable
    """
    y_pred = model.predict(predictors)
    cm = confusion_matrix(target, y_pred)
    labels = np.asarray(
        [
            ["{0:0.0f}".format(item) + "\n{0:.2%}".format(item / cm.flatten().sum())]
            for item in cm.flatten()
        ]
    ).reshape(2, 2)

    plt.figure(figsize=(6, 4))
    sns.heatmap(cm, annot=labels, fmt="")
    plt.ylabel("True label")
    plt.xlabel("Predicted label")
In [32]:
def model_performance_classification_sklearn(model, predictors, target):
    """
    Function to compute different metrics to check classification model performance

    model: classifier
    predictors: independent variables
    target: dependent variable
    """

    # predicting using the independent variables
    pred = model.predict(predictors)

    acc = accuracy_score(target, pred)  # to compute Accuracy
    recall = recall_score(target, pred)  # to compute Recall
    precision = precision_score(target, pred)  # to compute Precision
    f1 = f1_score(target, pred)  # to compute F1-score

    # creating a dataframe of metrics
    df_perf = pd.DataFrame(
        {
            "Accuracy": acc,
            "Recall": recall,
            "Precision": precision,
            "F1": f1,
        },
        index=[0],
    )

    return df_perf

Data Preparation for modeling¶

In [33]:
data['case_status'] = data['case_status'].apply(lambda x: 1 if x == "Certified" else 0)

X = data.drop('case_status',axis=1)
Y = data['case_status']
X = pd.get_dummies(X, drop_first=True)
X_train, X_test, y_train, y_test = train_test_split(X, Y, test_size=0.3, random_state=1, stratify=Y)
In [34]:
print("Shape of Training set : ", X_train.shape)
print("Shape of test set : ", X_test.shape)
print("Percentage of classes in training set:")
print(y_train.value_counts(normalize=True))
print("Percentage of classes in test set:")
print(y_test.value_counts(normalize=True))
Shape of Training set :  (17836, 21)
Shape of test set :  (7644, 21)
Percentage of classes in training set:
1    0.667919
0    0.332081
Name: case_status, dtype: float64
Percentage of classes in test set:
1    0.667844
0    0.332156
Name: case_status, dtype: float64

Building bagging and boosting models¶

Decision Tree¶

In [35]:
d_tree = DecisionTreeClassifier(random_state=1)
d_tree.fit(X_train,y_train)
confusion_matrix_sklearn(d_tree,X_test,y_test)
No description has been provided for this image
In [36]:
d_tree_model_train_perf=model_performance_classification_sklearn(d_tree,X_train,y_train)
print("Training performance:\n",d_tree_model_train_perf)
d_tree_model_test_perf=model_performance_classification_sklearn(d_tree,X_test,y_test)
print("Testing performance:\n",d_tree_model_test_perf)
Training performance:
    Accuracy  Recall  Precision   F1
0       1.0     1.0        1.0  1.0
Testing performance:
    Accuracy    Recall  Precision        F1
0  0.664443  0.742605   0.751884  0.747216

Hyperparameter Tuning

In [37]:
dtree_estimator = DecisionTreeClassifier(class_weight={0:0.18,1:0.72},random_state=1)

parameters = {
    'max_depth': np.arange(2,6),
    'min_samples_leaf': [1, 4, 7],
    'max_leaf_nodes' : [10, 15],
    'min_impurity_decrease': [0.0001,0.001]
}

scorer = metrics.make_scorer(metrics.f1_score)

grid_obj = GridSearchCV(dtree_estimator, parameters, scoring=scorer,n_jobs=-1)
grid_obj = grid_obj.fit(X_train, y_train)

dtree_estimator = grid_obj.best_estimator_
dtree_estimator.fit(X_train, y_train)
Out[37]:
DecisionTreeClassifier(class_weight={0: 0.18, 1: 0.72}, max_depth=4,
                       max_leaf_nodes=15, min_impurity_decrease=0.0001,
                       random_state=1)
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
DecisionTreeClassifier(class_weight={0: 0.18, 1: 0.72}, max_depth=4,
                       max_leaf_nodes=15, min_impurity_decrease=0.0001,
                       random_state=1)
In [38]:
confusion_matrix_sklearn(dtree_estimator,X_test,y_test)
dtree_estimator_model_train_perf=model_performance_classification_sklearn(dtree_estimator,X_train,y_train)
print("Training performance:\n",dtree_estimator_model_train_perf)
dtree_estimator_model_test_perf=model_performance_classification_sklearn(dtree_estimator,X_test,y_test)
print("Testing performance:\n",dtree_estimator_model_test_perf)
Training performance:
    Accuracy    Recall  Precision        F1
0  0.678908  0.995803    0.67634  0.805555
Testing performance:
    Accuracy    Recall  Precision        F1
0  0.674516  0.994711   0.673564  0.803227
No description has been provided for this image

Random Forest¶

In [39]:
rf_estimator = RandomForestClassifier(random_state=1)
rf_estimator.fit(X_train,y_train)
confusion_matrix_sklearn(rf_estimator,X_test,y_test)
rf_estimator_model_train_perf=model_performance_classification_sklearn(rf_estimator,X_train,y_train)
print("Training performance:\n",rf_estimator_model_train_perf)
rf_estimator_model_test_perf=model_performance_classification_sklearn(rf_estimator,X_test,y_test)
print("Testing performance:\n",rf_estimator_model_test_perf)
Training performance:
    Accuracy  Recall  Precision   F1
0       1.0     1.0        1.0  1.0
Testing performance:
    Accuracy    Recall  Precision        F1
0  0.720958  0.831342   0.769398  0.799171
No description has been provided for this image

Hyperparameter Tuning

In [40]:
# Hyperparameter Tuning
rf_tuned = RandomForestClassifier(class_weight={0:0.18,1:0.82},random_state=1,oob_score=True,bootstrap=True)
parameters = {
    "n_estimators": [50,110,25],
    "min_samples_leaf": np.arange(1, 4),
    "max_features": [np.arange(0.3, 0.6, 0.1),'sqrt'],
    "max_samples": np.arange(0.4, 0.7, 0.1)
}

scorer = metrics.make_scorer(metrics.f1_score)

grid_obj = GridSearchCV(rf_tuned, parameters, scoring=scorer, cv=5,n_jobs=-1)
grid_obj = grid_obj.fit(X_train, y_train)

rf_tuned = grid_obj.best_estimator_
rf_tuned.fit(X_train, y_train)
Out[40]:
RandomForestClassifier(class_weight={0: 0.18, 1: 0.82}, max_samples=0.5,
                       min_samples_leaf=2, n_estimators=50, oob_score=True,
                       random_state=1)
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
RandomForestClassifier(class_weight={0: 0.18, 1: 0.82}, max_samples=0.5,
                       min_samples_leaf=2, n_estimators=50, oob_score=True,
                       random_state=1)
In [41]:
confusion_matrix_sklearn(rf_tuned,X_test,y_test)
rf_tuned_model_train_perf=model_performance_classification_sklearn(rf_tuned,X_train,y_train)
print("Training performance:\n",rf_tuned_model_train_perf)
rf_tuned_model_test_perf=model_performance_classification_sklearn(rf_tuned,X_test,y_test)
print("Testing performance:\n",rf_tuned_model_test_perf)
Training performance:
    Accuracy    Recall  Precision        F1
0  0.799563  0.992697   0.772235  0.868697
Testing performance:
    Accuracy    Recall  Precision        F1
0  0.717949  0.942018   0.721098  0.816885
No description has been provided for this image

Bagging Classifier¶

In [42]:
bagging_classifier = BaggingClassifier(random_state=1)
bagging_classifier.fit(X_train,y_train)
confusion_matrix_sklearn(bagging_classifier,X_test,y_test)
bagging_classifier_model_train_perf=model_performance_classification_sklearn(bagging_classifier,X_train,y_train)
print(bagging_classifier_model_train_perf)
bagging_classifier_model_test_perf=model_performance_classification_sklearn(bagging_classifier,X_test,y_test)
print(bagging_classifier_model_test_perf)
   Accuracy    Recall  Precision        F1
0  0.985367  0.986066   0.991978  0.989013
   Accuracy    Recall  Precision        F1
0  0.693223  0.767091   0.772082  0.769578
No description has been provided for this image

Hyperparameter Tuning

In [43]:
bagging_estimator_tuned = BaggingClassifier(random_state=1)

parameters = {
    'max_samples': [0.8,0.9,1],
    'max_features': [0.7,0.8,0.9],
    'n_estimators' : [30,50,70],
             }

scorer = metrics.make_scorer(metrics.f1_score)

grid_obj = GridSearchCV(bagging_estimator_tuned, parameters, scoring=scorer,cv=5)
grid_obj = grid_obj.fit(X_train, y_train)

bagging_estimator_tuned = grid_obj.best_estimator_

bagging_estimator_tuned.fit(X_train, y_train)
Out[43]:
BaggingClassifier(max_features=0.7, max_samples=0.8, n_estimators=70,
                  random_state=1)
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
BaggingClassifier(max_features=0.7, max_samples=0.8, n_estimators=70,
                  random_state=1)
In [44]:
confusion_matrix_sklearn(bagging_estimator_tuned,X_test,y_test)
bagging_estimator_tuned_model_train_perf=model_performance_classification_sklearn(bagging_estimator_tuned,X_train,y_train)
print(bagging_estimator_tuned_model_train_perf)
bagging_estimator_tuned_model_test_perf=model_performance_classification_sklearn(bagging_estimator_tuned,X_test,y_test)
print(bagging_estimator_tuned_model_test_perf)
   Accuracy    Recall  Precision        F1
0  0.998654  0.999916   0.998073  0.998994
   Accuracy    Recall  Precision        F1
0  0.723836  0.886582   0.747111  0.810893
No description has been provided for this image

AdaBoost Classifier¶

In [45]:
ab_classifier = AdaBoostClassifier(random_state=1)
ab_classifier.fit(X_train,y_train)
confusion_matrix_sklearn(ab_classifier,X_test,y_test)
ab_classifier_model_train_perf=model_performance_classification_sklearn(ab_classifier,X_train,y_train)
print(ab_classifier_model_train_perf)
ab_classifier_model_test_perf=model_performance_classification_sklearn(ab_classifier,X_test,y_test)
print(ab_classifier_model_test_perf)
   Accuracy    Recall  Precision        F1
0  0.738058  0.887434   0.760411  0.819027
   Accuracy    Recall  Precision        F1
0  0.732993  0.885015    0.75653  0.815744
No description has been provided for this image

Hyperparameter Tuning

In [46]:
abc_tuned= AdaBoostClassifier(random_state=1)
parameters = {
    "n_estimators": np.arange(50,110,25),
    "learning_rate": [0.01,0.1,0.05],
    "base_estimator": [
        DecisionTreeClassifier(max_depth=2, random_state=1),
        DecisionTreeClassifier(max_depth=3, random_state=1),
    ],
}

scorer = metrics.make_scorer(metrics.f1_score)

grid_obj = GridSearchCV(abc_tuned, parameters, scoring='f1',cv=5)
grid_obj = grid_obj.fit(X_train, y_train)

abc_tuned = grid_obj.best_estimator_

abc_tuned.fit(X_train, y_train)
Out[46]:
AdaBoostClassifier(base_estimator=DecisionTreeClassifier(max_depth=3,
                                                         random_state=1),
                   learning_rate=0.1, random_state=1)
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
AdaBoostClassifier(base_estimator=DecisionTreeClassifier(max_depth=3,
                                                         random_state=1),
                   learning_rate=0.1, random_state=1)
DecisionTreeClassifier(max_depth=3, random_state=1)
DecisionTreeClassifier(max_depth=3, random_state=1)
In [47]:
confusion_matrix_sklearn(abc_tuned,X_test,y_test)
abc_tuned_model_train_perf=model_performance_classification_sklearn(abc_tuned,X_train,y_train)
print(abc_tuned_model_train_perf)
abc_tuned_model_test_perf=model_performance_classification_sklearn(abc_tuned,X_test,y_test)
print(abc_tuned_model_test_perf)
   Accuracy    Recall  Precision        F1
0   0.75314  0.888189   0.775051  0.827772
   Accuracy    Recall  Precision        F1
0  0.740842  0.881881   0.765646  0.819663
No description has been provided for this image

Gradient Boosting Classifier¶

In [48]:
gb_classifier = GradientBoostingClassifier(random_state=1)
gb_classifier.fit(X_train,y_train)
confusion_matrix_sklearn(gb_classifier,X_test,y_test)
gb_classifier_model_train_perf=model_performance_classification_sklearn(gb_classifier,X_train,y_train)
print("Training performance:\n",gb_classifier_model_train_perf)
gb_classifier_model_test_perf=model_performance_classification_sklearn(gb_classifier,X_test,y_test)
print("Testing performance:\n",gb_classifier_model_test_perf)
Training performance:
    Accuracy    Recall  Precision        F1
0  0.759419  0.882901   0.784106  0.830576
Testing performance:
    Accuracy    Recall  Precision       F1
0  0.744636  0.873262   0.773555  0.82039
No description has been provided for this image

Hyperparameter Tuning

In [49]:
gbc_tuned = GradientBoostingClassifier(init=AdaBoostClassifier(random_state=1),random_state=1)
parameters = {
    "init": [AdaBoostClassifier(random_state=1),DecisionTreeClassifier(random_state=1)],
    "n_estimators": np.arange(50,110,25),
    "learning_rate": [0.01,0.1,0.05],
    "subsample":[0.7,0.9],
    "max_features":[0.5,0.7,1],
}

scorer = metrics.make_scorer(metrics.f1_score)

grid_obj = GridSearchCV(gbc_tuned, parameters, scoring=scorer,cv=5)
grid_obj = grid_obj.fit(X_train, y_train)

gbc_tuned = grid_obj.best_estimator_
gbc_tuned.fit(X_train, y_train)
Out[49]:
GradientBoostingClassifier(init=AdaBoostClassifier(random_state=1),
                           learning_rate=0.05, max_features=0.5, random_state=1,
                           subsample=0.9)
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
GradientBoostingClassifier(init=AdaBoostClassifier(random_state=1),
                           learning_rate=0.05, max_features=0.5, random_state=1,
                           subsample=0.9)
AdaBoostClassifier(random_state=1)
AdaBoostClassifier(random_state=1)
In [50]:
confusion_matrix_sklearn(gbc_tuned,X_test,y_test)
gbc_tuned_model_train_perf=model_performance_classification_sklearn(gbc_tuned,X_train,y_train)
print("Training performance:\n",gbc_tuned_model_train_perf)
gbc_tuned_model_test_perf=model_performance_classification_sklearn(gbc_tuned,X_test,y_test)
print("Testing performance:\n",gbc_tuned_model_test_perf)
Training performance:
    Accuracy    Recall  Precision        F1
0  0.753756  0.884496   0.777466  0.827535
Testing performance:
    Accuracy    Recall  Precision        F1
0  0.744636  0.878355   0.771109  0.821245
No description has been provided for this image

XGBoost Classifier¶

In [51]:
xgb_classifier = XGBClassifier(random_state=1, eval_metric='logloss')
xgb_classifier.fit(X_train,y_train)
confusion_matrix_sklearn(xgb_classifier,X_test,y_test)
xgb_classifier_model_train_perf=model_performance_classification_sklearn(xgb_classifier,X_train,y_train)
print("Training performance:\n",xgb_classifier_model_train_perf)
xgb_classifier_model_test_perf=model_performance_classification_sklearn(xgb_classifier,X_test,y_test)
print("Testing performance:\n",xgb_classifier_model_test_perf)
Training performance:
    Accuracy    Recall  Precision        F1
0  0.843575  0.931084   0.849246  0.888284
Testing performance:
    Accuracy    Recall  Precision        F1
0  0.728545  0.855044   0.765789  0.807959
No description has been provided for this image

Hyperparameter Tuning

In [52]:
xgb_tuned = XGBClassifier(random_state=1, eval_metric='logloss')

parameters = {
  'n_estimators':np.arange(50,110,25),
  'scale_pos_weight':[1,2,5],
  'learning_rate':[0.01,0.1,0.05],
  'gamma':[1,3],
  'subsample':[0.7,0.9]
}

scorer = metrics.make_scorer(metrics.f1_score)

grid_obj = GridSearchCV(xgb_tuned, parameters,scoring=scorer,cv=5)
grid_obj = grid_obj.fit(X_train, y_train)

xgb_tuned = grid_obj.best_estimator_
xgb_tuned.fit(X_train, y_train)
Out[52]:
XGBClassifier(base_score=None, booster=None, callbacks=None,
              colsample_bylevel=None, colsample_bynode=None,
              colsample_bytree=None, device=None, early_stopping_rounds=None,
              enable_categorical=False, eval_metric='logloss',
              feature_types=None, gamma=3, grow_policy=None,
              importance_type=None, interaction_constraints=None,
              learning_rate=0.05, max_bin=None, max_cat_threshold=None,
              max_cat_to_onehot=None, max_delta_step=None, max_depth=None,
              max_leaves=None, min_child_weight=None, missing=nan,
              monotone_constraints=None, multi_strategy=None, n_estimators=50,
              n_jobs=None, num_parallel_tree=None, random_state=1, ...)
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
XGBClassifier(base_score=None, booster=None, callbacks=None,
              colsample_bylevel=None, colsample_bynode=None,
              colsample_bytree=None, device=None, early_stopping_rounds=None,
              enable_categorical=False, eval_metric='logloss',
              feature_types=None, gamma=3, grow_policy=None,
              importance_type=None, interaction_constraints=None,
              learning_rate=0.05, max_bin=None, max_cat_threshold=None,
              max_cat_to_onehot=None, max_delta_step=None, max_depth=None,
              max_leaves=None, min_child_weight=None, missing=nan,
              monotone_constraints=None, multi_strategy=None, n_estimators=50,
              n_jobs=None, num_parallel_tree=None, random_state=1, ...)
In [53]:
confusion_matrix_sklearn(xgb_tuned,X_test,y_test)
xgb_tuned_model_train_perf=model_performance_classification_sklearn(xgb_tuned,X_train,y_train)
print("Training performance:\n",xgb_tuned_model_train_perf)
xgb_tuned_model_test_perf=model_performance_classification_sklearn(xgb_tuned,X_test,y_test)
print("Testing performance:\n",xgb_tuned_model_test_perf)
Training performance:
    Accuracy    Recall  Precision       F1
0   0.76183  0.887602   0.784247  0.83273
Testing performance:
    Accuracy    Recall  Precision        F1
0  0.745683  0.877963   0.772359  0.821782
No description has been provided for this image

Stacking Model¶

In [54]:
estimators = [('Random Forest',rf_tuned), ('Gradient Boosting',gbc_tuned), ('Decision Tree',dtree_estimator)]
final_estimator = xgb_tuned
stacking_classifier= StackingClassifier(estimators=estimators,final_estimator=final_estimator)
stacking_classifier.fit(X_train,y_train)
Out[54]:
StackingClassifier(estimators=[('Random Forest',
                                RandomForestClassifier(class_weight={0: 0.18,
                                                                     1: 0.82},
                                                       max_samples=0.5,
                                                       min_samples_leaf=2,
                                                       n_estimators=50,
                                                       oob_score=True,
                                                       random_state=1)),
                               ('Gradient Boosting',
                                GradientBoostingClassifier(init=AdaBoostClassifier(random_state=1),
                                                           learning_rate=0.05,
                                                           max_features=0.5,
                                                           random_state=1,
                                                           subsample=0.9)),
                               ('Decision Tree...
                                                 feature_types=None, gamma=3,
                                                 grow_policy=None,
                                                 importance_type=None,
                                                 interaction_constraints=None,
                                                 learning_rate=0.05,
                                                 max_bin=None,
                                                 max_cat_threshold=None,
                                                 max_cat_to_onehot=None,
                                                 max_delta_step=None,
                                                 max_depth=None,
                                                 max_leaves=None,
                                                 min_child_weight=None,
                                                 missing=nan,
                                                 monotone_constraints=None,
                                                 multi_strategy=None,
                                                 n_estimators=50, n_jobs=None,
                                                 num_parallel_tree=None,
                                                 random_state=1, ...))
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
StackingClassifier(estimators=[('Random Forest',
                                RandomForestClassifier(class_weight={0: 0.18,
                                                                     1: 0.82},
                                                       max_samples=0.5,
                                                       min_samples_leaf=2,
                                                       n_estimators=50,
                                                       oob_score=True,
                                                       random_state=1)),
                               ('Gradient Boosting',
                                GradientBoostingClassifier(init=AdaBoostClassifier(random_state=1),
                                                           learning_rate=0.05,
                                                           max_features=0.5,
                                                           random_state=1,
                                                           subsample=0.9)),
                               ('Decision Tree...
                                                 feature_types=None, gamma=3,
                                                 grow_policy=None,
                                                 importance_type=None,
                                                 interaction_constraints=None,
                                                 learning_rate=0.05,
                                                 max_bin=None,
                                                 max_cat_threshold=None,
                                                 max_cat_to_onehot=None,
                                                 max_delta_step=None,
                                                 max_depth=None,
                                                 max_leaves=None,
                                                 min_child_weight=None,
                                                 missing=nan,
                                                 monotone_constraints=None,
                                                 multi_strategy=None,
                                                 n_estimators=50, n_jobs=None,
                                                 num_parallel_tree=None,
                                                 random_state=1, ...))
RandomForestClassifier(class_weight={0: 0.18, 1: 0.82}, max_samples=0.5,
                       min_samples_leaf=2, n_estimators=50, oob_score=True,
                       random_state=1)
AdaBoostClassifier(random_state=1)
AdaBoostClassifier(random_state=1)
DecisionTreeClassifier(class_weight={0: 0.18, 1: 0.72}, max_depth=4,
                       max_leaf_nodes=15, min_impurity_decrease=0.0001,
                       random_state=1)
XGBClassifier(base_score=None, booster=None, callbacks=None,
              colsample_bylevel=None, colsample_bynode=None,
              colsample_bytree=None, device=None, early_stopping_rounds=None,
              enable_categorical=False, eval_metric='logloss',
              feature_types=None, gamma=3, grow_policy=None,
              importance_type=None, interaction_constraints=None,
              learning_rate=0.05, max_bin=None, max_cat_threshold=None,
              max_cat_to_onehot=None, max_delta_step=None, max_depth=None,
              max_leaves=None, min_child_weight=None, missing=nan,
              monotone_constraints=None, multi_strategy=None, n_estimators=50,
              n_jobs=None, num_parallel_tree=None, random_state=1, ...)
In [55]:
confusion_matrix_sklearn(stacking_classifier,X_test,y_test)
stacking_classifier_model_train_perf=model_performance_classification_sklearn(stacking_classifier,X_train,y_train)
print("Training performance:\n",stacking_classifier_model_train_perf)
stacking_classifier_model_test_perf=model_performance_classification_sklearn(stacking_classifier,X_test,y_test)
print("Testing performance:\n",stacking_classifier_model_test_perf)
Training performance:
    Accuracy    Recall  Precision        F1
0  0.782855  0.900277   0.799776  0.847056
Testing performance:
    Accuracy    Recall  Precision        F1
0  0.745814  0.879138    0.77193  0.822053
No description has been provided for this image

Model Performance Comparison and Conclusions¶

In [56]:
models_train_comp_df = pd.concat(
    [d_tree_model_train_perf.T,dtree_estimator_model_train_perf.T,rf_estimator_model_train_perf.T,rf_tuned_model_train_perf.T,
     bagging_classifier_model_train_perf.T,bagging_estimator_tuned_model_train_perf.T,ab_classifier_model_train_perf.T,
     abc_tuned_model_train_perf.T,gb_classifier_model_train_perf.T,gbc_tuned_model_train_perf.T,xgb_classifier_model_train_perf.T,
    xgb_tuned_model_train_perf.T,stacking_classifier_model_train_perf.T],
    axis=1,
)
models_train_comp_df.columns = [
    "Decision Tree",
    "Decision Tree Estimator",
    "Random Forest Estimator",
    "Random Forest Tuned",
    "Bagging Classifier",
    "Bagging Estimator Tuned",
    "Adaboost Classifier",
    "Adabosst Classifier Tuned",
    "Gradient Boost Classifier",
    "Gradient Boost Classifier Tuned",
    "XGBoost Classifier",
    "XGBoost Classifier Tuned",
    "Stacking Classifier"]
print("Training performance comparison:")
models_train_comp_df
Training performance comparison:
Out[56]:
Decision Tree Decision Tree Estimator Random Forest Estimator Random Forest Tuned Bagging Classifier Bagging Estimator Tuned Adaboost Classifier Adabosst Classifier Tuned Gradient Boost Classifier Gradient Boost Classifier Tuned XGBoost Classifier XGBoost Classifier Tuned Stacking Classifier
Accuracy 1.0 0.678908 1.0 0.799563 0.985367 0.998654 0.738058 0.753140 0.759419 0.753756 0.843575 0.761830 0.782855
Recall 1.0 0.995803 1.0 0.992697 0.986066 0.999916 0.887434 0.888189 0.882901 0.884496 0.931084 0.887602 0.900277
Precision 1.0 0.676340 1.0 0.772235 0.991978 0.998073 0.760411 0.775051 0.784106 0.777466 0.849246 0.784247 0.799776
F1 1.0 0.805555 1.0 0.868697 0.989013 0.998994 0.819027 0.827772 0.830576 0.827535 0.888284 0.832730 0.847056
In [57]:
models_test_comp_df = pd.concat(
    [d_tree_model_test_perf.T,dtree_estimator_model_test_perf.T,rf_estimator_model_test_perf.T,rf_tuned_model_test_perf.T,
     bagging_classifier_model_test_perf.T,bagging_estimator_tuned_model_test_perf.T,ab_classifier_model_test_perf.T,
     abc_tuned_model_test_perf.T,gb_classifier_model_test_perf.T,gbc_tuned_model_test_perf.T,xgb_classifier_model_test_perf.T,
    xgb_tuned_model_test_perf.T,stacking_classifier_model_test_perf.T],
    axis=1,
)
models_test_comp_df.columns = [
    "Decision Tree",
    "Decision Tree Estimator",
    "Random Forest Estimator",
    "Random Forest Tuned",
    "Bagging Classifier",
    "Bagging Estimator Tuned",
    "Adaboost Classifier",
    "Adabosst Classifier Tuned",
    "Gradient Boost Classifier",
    "Gradient Boost Classifier Tuned",
    "XGBoost Classifier",
    "XGBoost Classifier Tuned",
    "Stacking Classifier"]
print("Testing performance comparison:")
models_test_comp_df
Testing performance comparison:
Out[57]:
Decision Tree Decision Tree Estimator Random Forest Estimator Random Forest Tuned Bagging Classifier Bagging Estimator Tuned Adaboost Classifier Adabosst Classifier Tuned Gradient Boost Classifier Gradient Boost Classifier Tuned XGBoost Classifier XGBoost Classifier Tuned Stacking Classifier
Accuracy 0.664443 0.674516 0.720958 0.717949 0.693223 0.723836 0.732993 0.740842 0.744636 0.744636 0.728545 0.745683 0.745814
Recall 0.742605 0.994711 0.831342 0.942018 0.767091 0.886582 0.885015 0.881881 0.873262 0.878355 0.855044 0.877963 0.879138
Precision 0.751884 0.673564 0.769398 0.721098 0.772082 0.747111 0.756530 0.765646 0.773555 0.771109 0.765789 0.772359 0.771930
F1 0.747216 0.803227 0.799171 0.816885 0.769578 0.810893 0.815744 0.819663 0.820390 0.821245 0.807959 0.821782 0.822053

Feature Importance of XGBoost¶

In [58]:
feature_names = X_train.columns
importances = xgb_classifier.feature_importances_
indices = np.argsort(importances)

plt.figure(figsize=(12,12))
plt.title('Feature Importances')
plt.barh(range(len(indices)), importances[indices], color='violet', align='center')
plt.yticks(range(len(indices)), [feature_names[i] for i in indices])
plt.xlabel('Relative Importance')
plt.show()
No description has been provided for this image

Actionable Insights and Recommendations¶

Feature Importance:¶

  1. Top Features: The features "education_of_employee_High School" and "education_of_employee_Doctorate" are the most influential for predicting visa approval, highlighting the significant role of educational background in the decision process.
  2. Significant Features: Other notable features include "unit_of_wage_Year," "continent_Europe," and "has_job_experience_Y," suggesting that factors like wage unit, geographical location, and job experience also play crucial roles.

Model Performance:¶

  • Training Performance:

    • The XGBoost classifier (tuned) shows a good balance between accuracy (0.763848), recall (0.896919), and precision (0.781696), with a high F1 score (0.835353).
    • Models like the Decision Tree and Random Forest show perfect accuracy (1.0) on training data but are likely overfitting.
  • Test Data Performance:

    • The Gradient Boost Classifier (tuned) and XGBoost Classifier (tuned) are among the best performing on the test data, showing good generalization with F1 scores around 0.82.
    • The high recall (0.994711) of the Decision Tree Estimator suggests it’s good at identifying positive cases but might also generate false positives due to lower precision (0.673564).

Actionable Insights:¶

  • Focus on Education and Job Experience: The OFLC can prioritize applicants with higher education and relevant job experience, especially those from regions like Europe where the likelihood of visa approval seems higher.
  • Wage and Employment Terms: Focus on yearly wage units could could streamline application processes.
  • Model Selection: For automated decision-making, models like Gradient Boost and XGBoost (tuned versions) are recommended due to their balanced performance, minimizing overfitting while maintaining accuracy.

These insights can guide OFLC in making data-driven decisions for visa applications, improving efficiency in handling the increasing number of applications.