Credit Card Users Churn PredictionΒΆ

Problem StatementΒΆ

Business ContextΒΆ

The Thera bank recently saw a steep decline in the number of users of their credit card, credit cards are a good source of income for banks because of different kinds of fees charged by the banks like annual fees, balance transfer fees, and cash advance fees, late payment fees, foreign transaction fees, and others. Some fees are charged to every user irrespective of usage, while others are charged under specified circumstances.

Customers’ leaving credit cards services would lead bank to loss, so the bank wants to analyze the data of customers and identify the customers who will leave their credit card services and reason for same – so that bank could improve upon those areas

You as a Data scientist at Thera bank need to come up with a classification model that will help the bank improve its services so that customers do not renounce their credit cards

Data DescriptionΒΆ

  • CLIENTNUM: Client number. Unique identifier for the customer holding the account
  • Attrition_Flag: Internal event (customer activity) variable - if the account is closed then "Attrited Customer" else "Existing Customer"
  • Customer_Age: Age in Years
  • Gender: Gender of the account holder
  • Dependent_count: Number of dependents
  • Education_Level: Educational Qualification of the account holder - Graduate, High School, Unknown, Uneducated, College(refers to college student), Post-Graduate, Doctorate
  • Marital_Status: Marital Status of the account holder
  • Income_Category: Annual Income Category of the account holder
  • Card_Category: Type of Card
  • Months_on_book: Period of relationship with the bank (in months)
  • Total_Relationship_Count: Total no. of products held by the customer
  • Months_Inactive_12_mon: No. of months inactive in the last 12 months
  • Contacts_Count_12_mon: No. of Contacts in the last 12 months
  • Credit_Limit: Credit Limit on the Credit Card
  • Total_Revolving_Bal: Total Revolving Balance on the Credit Card
  • Avg_Open_To_Buy: Open to Buy Credit Line (Average of last 12 months)
  • Total_Amt_Chng_Q4_Q1: Change in Transaction Amount (Q4 over Q1)
  • Total_Trans_Amt: Total Transaction Amount (Last 12 months)
  • Total_Trans_Ct: Total Transaction Count (Last 12 months)
  • Total_Ct_Chng_Q4_Q1: Change in Transaction Count (Q4 over Q1)
  • Avg_Utilization_Ratio: Average Card Utilization Ratio

What Is a Revolving Balance?ΒΆ

  • If we don't pay the balance of the revolving credit account in full every month, the unpaid portion carries over to the next month. That's called a revolving balance
What is the Average Open to buy?ΒΆ
  • 'Open to Buy' means the amount left on your credit card to use. Now, this column represents the average of this value for the last 12 months.
What is the Average utilization Ratio?ΒΆ
  • The Avg_Utilization_Ratio represents how much of the available credit the customer spent. This is useful for calculating credit scores.
Relation b/w Avg_Open_To_Buy, Credit_Limit and Avg_Utilization_Ratio:ΒΆ
  • ( Avg_Open_To_Buy / Credit_Limit ) + Avg_Utilization_Ratio = 1

Please read the instructions carefully before starting the project.ΒΆ

This is a commented Jupyter IPython Notebook file in which all the instructions and tasks to be performed are mentioned.

  • Blanks '_______' are provided in the notebook that

needs to be filled with an appropriate code to get the correct result. With every '_______' blank, there is a comment that briefly describes what needs to be filled in the blank space.

  • Identify the task to be performed correctly, and only then proceed to write the required code.
  • Fill the code wherever asked by the commented lines like "# write your code here" or "# complete the code". Running incomplete code may throw error.
  • Please run the codes in a sequential manner from the beginning to avoid any unnecessary errors.
  • Add the results/observations (wherever mentioned) derived from the analysis in the presentation and submit the same.

Importing necessary librariesΒΆ

InΒ [Β ]:
# Installing the libraries with the specified version.
# uncomment and run the following line if Google Colab is being used
# !pip install scikit-learn==1.2.2 seaborn==0.13.1 matplotlib==3.7.1 numpy==1.25.2 pandas==1.5.3 imbalanced-learn==0.10.1 xgboost==2.0.3 -q --user
InΒ [Β ]:
# Installing the libraries with the specified version.
# uncomment and run the following lines if Jupyter Notebook is being used
# !pip install scikit-learn==1.2.2 seaborn==0.13.1 matplotlib==3.7.1 numpy==1.25.2 pandas==1.5.3 imblearn==0.12.0 xgboost==2.0.3 -q --user
# !pip install --upgrade -q threadpoolctl

Note: After running the above cell, kindly restart the notebook kernel and run all cells sequentially from the start again.

InΒ [1]:
import pandas as pd
import numpy as np
from sklearn import metrics
import matplotlib.pyplot as plt
%matplotlib inline
import warnings
warnings.filterwarnings('ignore')
import seaborn as sns
from sklearn.model_selection import train_test_split
from sklearn.model_selection import GridSearchCV
from sklearn import metrics
from sklearn.ensemble import BaggingClassifier, RandomForestClassifier, AdaBoostClassifier, GradientBoostingClassifier
from sklearn.tree import DecisionTreeClassifier
from sklearn.linear_model import LogisticRegression
from xgboost import XGBClassifier
from sklearn.metrics import accuracy_score, recall_score, precision_score, f1_score, confusion_matrix
from imblearn.over_sampling import SMOTE
from imblearn.under_sampling import RandomUnderSampler

# To tune a model
from sklearn.model_selection import GridSearchCV
from sklearn.model_selection import RandomizedSearchCV

# To suppress warnings
import warnings
warnings.filterwarnings("ignore")
InΒ [2]:
# Set display options to avoid scientific notation
pd.set_option('display.float_format', '{:.2f}'.format)

# Set the option to display all columns
pd.set_option('display.max_columns', None)

Loading the datasetΒΆ

InΒ [3]:
# uncomment and run the following line if using Google Colab
from google.colab import drive
drive.mount('/content/drive')
Mounted at /content/drive
InΒ [4]:
bankData = pd.read_csv("/content/drive/MyDrive/Personal/UT Austin/Project 3 - AML/BankChurners.csv")

data = bankData.copy()

CLIENTNUM won't be used in the analysis, so can remove it.

InΒ [5]:
#Remove CLIENTNUM column from data
data.drop(['CLIENTNUM'], axis=1, inplace=True)

Data OverviewΒΆ

  • Observations
  • Sanity checks

Viewing the first and last 5 rows of the datasetΒΆ

InΒ [Β ]:
data.head(5)
Out[Β ]:
CLIENTNUM Attrition_Flag Customer_Age Gender Dependent_count Education_Level Marital_Status Income_Category Card_Category Months_on_book ... Months_Inactive_12_mon Contacts_Count_12_mon Credit_Limit Total_Revolving_Bal Avg_Open_To_Buy Total_Amt_Chng_Q4_Q1 Total_Trans_Amt Total_Trans_Ct Total_Ct_Chng_Q4_Q1 Avg_Utilization_Ratio
0 768805383 Existing Customer 45 M 3 High School Married $60K - $80K Blue 39 ... 1 3 12691.00 777 11914.00 1.33 1144 42 1.62 0.06
1 818770008 Existing Customer 49 F 5 Graduate Single Less than $40K Blue 44 ... 1 2 8256.00 864 7392.00 1.54 1291 33 3.71 0.10
2 713982108 Existing Customer 51 M 3 Graduate Married $80K - $120K Blue 36 ... 1 0 3418.00 0 3418.00 2.59 1887 20 2.33 0.00
3 769911858 Existing Customer 40 F 4 High School NaN Less than $40K Blue 34 ... 4 1 3313.00 2517 796.00 1.41 1171 20 2.33 0.76
4 709106358 Existing Customer 40 M 3 Uneducated Married $60K - $80K Blue 21 ... 1 0 4716.00 0 4716.00 2.17 816 28 2.50 0.00

5 rows Γ— 21 columns

InΒ [Β ]:
data.tail(5)
Out[Β ]:
CLIENTNUM Attrition_Flag Customer_Age Gender Dependent_count Education_Level Marital_Status Income_Category Card_Category Months_on_book ... Months_Inactive_12_mon Contacts_Count_12_mon Credit_Limit Total_Revolving_Bal Avg_Open_To_Buy Total_Amt_Chng_Q4_Q1 Total_Trans_Amt Total_Trans_Ct Total_Ct_Chng_Q4_Q1 Avg_Utilization_Ratio
10122 772366833 Existing Customer 50 M 2 Graduate Single $40K - $60K Blue 40 ... 2 3 4003.00 1851 2152.00 0.70 15476 117 0.86 0.46
10123 710638233 Attrited Customer 41 M 2 NaN Divorced $40K - $60K Blue 25 ... 2 3 4277.00 2186 2091.00 0.80 8764 69 0.68 0.51
10124 716506083 Attrited Customer 44 F 1 High School Married Less than $40K Blue 36 ... 3 4 5409.00 0 5409.00 0.82 10291 60 0.82 0.00
10125 717406983 Attrited Customer 30 M 2 Graduate NaN $40K - $60K Blue 36 ... 3 3 5281.00 0 5281.00 0.54 8395 62 0.72 0.00
10126 714337233 Attrited Customer 43 F 2 Graduate Married Less than $40K Silver 25 ... 2 4 10388.00 1961 8427.00 0.70 10294 61 0.65 0.19

5 rows Γ— 21 columns

Checking the shape of the dataset, attribute types, and statistical summaryΒΆ

InΒ [Β ]:
data.shape
Out[Β ]:
(10127, 21)
InΒ [Β ]:
data.info()
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 10127 entries, 0 to 10126
Data columns (total 21 columns):
 #   Column                    Non-Null Count  Dtype  
---  ------                    --------------  -----  
 0   CLIENTNUM                 10127 non-null  int64  
 1   Attrition_Flag            10127 non-null  object 
 2   Customer_Age              10127 non-null  int64  
 3   Gender                    10127 non-null  object 
 4   Dependent_count           10127 non-null  int64  
 5   Education_Level           8608 non-null   object 
 6   Marital_Status            9378 non-null   object 
 7   Income_Category           10127 non-null  object 
 8   Card_Category             10127 non-null  object 
 9   Months_on_book            10127 non-null  int64  
 10  Total_Relationship_Count  10127 non-null  int64  
 11  Months_Inactive_12_mon    10127 non-null  int64  
 12  Contacts_Count_12_mon     10127 non-null  int64  
 13  Credit_Limit              10127 non-null  float64
 14  Total_Revolving_Bal       10127 non-null  int64  
 15  Avg_Open_To_Buy           10127 non-null  float64
 16  Total_Amt_Chng_Q4_Q1      10127 non-null  float64
 17  Total_Trans_Amt           10127 non-null  int64  
 18  Total_Trans_Ct            10127 non-null  int64  
 19  Total_Ct_Chng_Q4_Q1       10127 non-null  float64
 20  Avg_Utilization_Ratio     10127 non-null  float64
dtypes: float64(5), int64(10), object(6)
memory usage: 1.6+ MB
InΒ [Β ]:
#Get Summary Statistics for the numerical columns
data.describe().T
Out[Β ]:
count mean std min 25% 50% 75% max
CLIENTNUM 10127.00 739177606.33 36903783.45 708082083.00 713036770.50 717926358.00 773143533.00 828343083.00
Customer_Age 10127.00 46.33 8.02 26.00 41.00 46.00 52.00 73.00
Dependent_count 10127.00 2.35 1.30 0.00 1.00 2.00 3.00 5.00
Months_on_book 10127.00 35.93 7.99 13.00 31.00 36.00 40.00 56.00
Total_Relationship_Count 10127.00 3.81 1.55 1.00 3.00 4.00 5.00 6.00
Months_Inactive_12_mon 10127.00 2.34 1.01 0.00 2.00 2.00 3.00 6.00
Contacts_Count_12_mon 10127.00 2.46 1.11 0.00 2.00 2.00 3.00 6.00
Credit_Limit 10127.00 8631.95 9088.78 1438.30 2555.00 4549.00 11067.50 34516.00
Total_Revolving_Bal 10127.00 1162.81 814.99 0.00 359.00 1276.00 1784.00 2517.00
Avg_Open_To_Buy 10127.00 7469.14 9090.69 3.00 1324.50 3474.00 9859.00 34516.00
Total_Amt_Chng_Q4_Q1 10127.00 0.76 0.22 0.00 0.63 0.74 0.86 3.40
Total_Trans_Amt 10127.00 4404.09 3397.13 510.00 2155.50 3899.00 4741.00 18484.00
Total_Trans_Ct 10127.00 64.86 23.47 10.00 45.00 67.00 81.00 139.00
Total_Ct_Chng_Q4_Q1 10127.00 0.71 0.24 0.00 0.58 0.70 0.82 3.71
Avg_Utilization_Ratio 10127.00 0.27 0.28 0.00 0.02 0.18 0.50 1.00

Checking for missing values or duplicate valuesΒΆ

InΒ [Β ]:
# checking for null values
data.isnull().sum()
Out[Β ]:
0
CLIENTNUM 0
Attrition_Flag 0
Customer_Age 0
Gender 0
Dependent_count 0
Education_Level 1519
Marital_Status 749
Income_Category 0
Card_Category 0
Months_on_book 0
Total_Relationship_Count 0
Months_Inactive_12_mon 0
Contacts_Count_12_mon 0
Credit_Limit 0
Total_Revolving_Bal 0
Avg_Open_To_Buy 0
Total_Amt_Chng_Q4_Q1 0
Total_Trans_Amt 0
Total_Trans_Ct 0
Total_Ct_Chng_Q4_Q1 0
Avg_Utilization_Ratio 0

  • Education_Level (N/A) has ~15% missing data.
  • Martial_Status (NaN) has ~7.5% missing data
  • Although not evident in the above Income_Category has ~11% missing values that show up as 'abc'.
InΒ [Β ]:
#Checking for duplicate values
data.duplicated().sum()
Out[Β ]:
0

There are no duplicate rows in the dataset

Checking the mix of customers who attrited vs remained with the credit card.ΒΆ

InΒ [Β ]:
# Get counts of each unique value
data['Attrition_Flag'].value_counts()
Out[Β ]:
count
Attrition_Flag
Existing Customer 8500
Attrited Customer 1627

InΒ [Β ]:
# Get percentages of each unique value
data['Attrition_Flag'].value_counts(normalize=True) * 100
Out[Β ]:
proportion
Attrition_Flag
Existing Customer 83.93
Attrited Customer 16.07

Attrited Customers make up 16% of the total data, so there is definitely an imbalance between Existing and Attrited the model will need to consider.

Exploratory Data Analysis (EDA)ΒΆ

  • EDA is an important part of any project involving data.
  • It is important to investigate and understand the data better before building a model with it.
  • A few questions have been mentioned below which will help you approach the analysis in the right manner and generate insights from the data.
  • A thorough analysis of the data, in addition to the questions mentioned below, should be done.

Questions:

  1. How is the total transaction amount distributed?
  2. What is the distribution of the level of education of customers?
  3. What is the distribution of the level of income of customers?
  4. How does the change in transaction amount between Q4 and Q1 (total_ct_change_Q4_Q1) vary by the customer's account status (Attrition_Flag)?
  5. How does the number of months a customer was inactive in the last 12 months (Months_Inactive_12_mon) vary by the customer's account status (Attrition_Flag)?
  6. What are the attributes that have a strong correlation with each other?

The below functions need to be defined to carry out the Exploratory Data Analysis.ΒΆ

InΒ [Β ]:
# function to plot a boxplot and a histogram along the same scale.


def histogram_boxplot(data, feature, figsize=(12, 7), kde=False, bins=None):
    """
    Boxplot and histogram combined

    data: dataframe
    feature: dataframe column
    figsize: size of figure (default (12,7))
    kde: whether to the show density curve (default False)
    bins: number of bins for histogram (default None)
    """
    f2, (ax_box2, ax_hist2) = plt.subplots(
        nrows=2,  # Number of rows of the subplot grid= 2
        sharex=True,  # x-axis will be shared among all subplots
        gridspec_kw={"height_ratios": (0.25, 0.75)},
        figsize=figsize,
    )  # creating the 2 subplots
    sns.boxplot(
        data=data, x=feature, ax=ax_box2, showmeans=True, color="violet"
    )  # boxplot will be created and a triangle will indicate the mean value of the column
    sns.histplot(
        data=data, x=feature, kde=kde, ax=ax_hist2, bins=bins, palette="winter"
    ) if bins else sns.histplot(
        data=data, x=feature, kde=kde, ax=ax_hist2
    )  # For histogram
    ax_hist2.axvline(
        data[feature].mean(), color="green", linestyle="--"
    )  # Add mean to the histogram
    ax_hist2.axvline(
        data[feature].median(), color="black", linestyle="-"
    )  # Add median to the histogram
InΒ [Β ]:
# function to create labeled barplots


def labeled_barplot(data, feature, perc=False, n=None):
    """
    Barplot with percentage at the top

    data: dataframe
    feature: dataframe column
    perc: whether to display percentages instead of count (default is False)
    n: displays the top n category levels (default is None, i.e., display all levels)
    """

    total = len(data[feature])  # length of the column
    count = data[feature].nunique()
    if n is None:
        plt.figure(figsize=(count + 1, 5))
    else:
        plt.figure(figsize=(n + 1, 5))

    plt.xticks(rotation=90, fontsize=15)
    ax = sns.countplot(
        data=data,
        x=feature,
        palette="Paired",
        order=data[feature].value_counts().index[:n].sort_values(),
    )

    for p in ax.patches:
        if perc == True:
            label = "{:.1f}%".format(
                100 * p.get_height() / total
            )  # percentage of each class of the category
        else:
            label = p.get_height()  # count of each level of the category

        x = p.get_x() + p.get_width() / 2  # width of the plot
        y = p.get_height()  # height of the plot

        ax.annotate(
            label,
            (x, y),
            ha="center",
            va="center",
            size=12,
            xytext=(0, 5),
            textcoords="offset points",
        )  # annotate the percentage

    plt.show()  # show the plot
InΒ [Β ]:
# function to plot stacked bar chart

def stacked_barplot(data, predictor, target):
    """
    Print the category counts and plot a stacked bar chart

    data: dataframe
    predictor: independent variable
    target: target variable
    """
    count = data[predictor].nunique()
    sorter = data[target].value_counts().index[-1]
    tab1 = pd.crosstab(data[predictor], data[target], margins=True).sort_values(
        by=sorter, ascending=False
    )
    print(tab1)
    print("-" * 120)
    tab = pd.crosstab(data[predictor], data[target], normalize="index").sort_values(
        by=sorter, ascending=False
    )
    tab.plot(kind="bar", stacked=True, figsize=(count + 1, 5))
    plt.legend(
        loc="lower left", frameon=False,
    )
    plt.legend(loc="upper left", bbox_to_anchor=(1, 1))
    plt.show()
InΒ [Β ]:
### Function to plot distributions

def distribution_plot_wrt_target(data, predictor, target):

    fig, axs = plt.subplots(2, 2, figsize=(12, 10))

    target_uniq = data[target].unique()

    axs[0, 0].set_title("Distribution of target for target=" + str(target_uniq[0]))
    sns.histplot(
        data=data[data[target] == target_uniq[0]],
        x=predictor,
        kde=True,
        ax=axs[0, 0],
        color="teal",
    )

    axs[0, 1].set_title("Distribution of target for target=" + str(target_uniq[1]))
    sns.histplot(
        data=data[data[target] == target_uniq[1]],
        x=predictor,
        kde=True,
        ax=axs[0, 1],
        color="orange",
    )

    axs[1, 0].set_title("Boxplot w.r.t target")
    sns.boxplot(data=data, x=target, y=predictor, ax=axs[1, 0], palette="gist_rainbow")

    axs[1, 1].set_title("Boxplot (without outliers) w.r.t target")
    sns.boxplot(
        data=data,
        x=target,
        y=predictor,
        ax=axs[1, 1],
        showfliers=False,
        palette="gist_rainbow",
    )

    plt.tight_layout()
    plt.show()

Univariate AnalysisΒΆ

Numerical VariablesΒΆ

InΒ [Β ]:
histogram_boxplot(data, 'Customer_Age')
No description has been provided for this image
InΒ [Β ]:
histogram_boxplot(data, 'Dependent_count')
No description has been provided for this image
InΒ [Β ]:
histogram_boxplot(data, 'Months_on_book')
No description has been provided for this image
InΒ [Β ]:
histogram_boxplot(data, 'Total_Relationship_Count')
No description has been provided for this image
InΒ [Β ]:
histogram_boxplot(data, 'Months_Inactive_12_mon')
No description has been provided for this image
InΒ [Β ]:
histogram_boxplot(data, 'Contacts_Count_12_mon')
No description has been provided for this image
InΒ [Β ]:
histogram_boxplot(data, 'Credit_Limit')
No description has been provided for this image
InΒ [Β ]:
histogram_boxplot(data, 'Total_Revolving_Bal')
No description has been provided for this image
InΒ [Β ]:
histogram_boxplot(data, 'Avg_Open_To_Buy')
No description has been provided for this image
InΒ [Β ]:
histogram_boxplot(data, 'Total_Amt_Chng_Q4_Q1')
No description has been provided for this image
InΒ [Β ]:
histogram_boxplot(data, 'Total_Trans_Amt')
No description has been provided for this image
InΒ [Β ]:
histogram_boxplot(data, 'Total_Trans_Ct')
No description has been provided for this image
InΒ [Β ]:
histogram_boxplot(data, 'Total_Ct_Chng_Q4_Q1')
No description has been provided for this image
InΒ [Β ]:
histogram_boxplot(data, 'Avg_Utilization_Ratio')
No description has been provided for this image

Categorical VariablesΒΆ

InΒ [Β ]:
labeled_barplot(data, 'Gender')
No description has been provided for this image
InΒ [Β ]:
labeled_barplot(data, 'Education_Level')
No description has been provided for this image
InΒ [Β ]:
labeled_barplot(data, 'Marital_Status')
No description has been provided for this image
InΒ [Β ]:
labeled_barplot(data, 'Income_Category')
No description has been provided for this image
InΒ [Β ]:
labeled_barplot(data, 'Card_Category')
No description has been provided for this image

Bivariate AnalysisΒΆ

Numerical VariablesΒΆ

Let's look at the correlation matrix for all the numeric variables.

InΒ [Β ]:
# plotting the correlation heatmap

plt.figure(figsize=(12, 7))
sns.heatmap(data.corr(numeric_only = True), annot=True, fmt='0.2f', cmap='Spectral')
plt.show()
No description has been provided for this image

Some notable correlations:

  1. Months_on_book is correlated with the Customer_Age which makes sense given the older someone is the better chance they have to have a longer relationship with the bank.
  2. Total_Trans_Ct is correlated with Total_Trans_Amt which makes sense in that the more transactions you have the more the $ amount will be.
  3. Avg_Utilization_Ratio is correlated with Total_Revolving_Balance which makes sense in that when you aren't paying your balance in full each month, you will be utilizing more of your available credit.
InΒ [Β ]:
sns.pairplot(data,hue='Attrition_Flag')
plt.show()
Output hidden; open in https://colab.research.google.com to view.
InΒ [Β ]:
distribution_plot_wrt_target(data, 'Customer_Age', 'Attrition_Flag')
No description has been provided for this image
InΒ [Β ]:
distribution_plot_wrt_target(data, 'Dependent_count', 'Attrition_Flag')
No description has been provided for this image
InΒ [Β ]:
distribution_plot_wrt_target(data, 'Months_on_book', 'Attrition_Flag')
No description has been provided for this image
InΒ [Β ]:
distribution_plot_wrt_target(data, 'Total_Relationship_Count', 'Attrition_Flag')
No description has been provided for this image
InΒ [Β ]:
distribution_plot_wrt_target(data, 'Months_Inactive_12_mon', 'Attrition_Flag')
No description has been provided for this image
InΒ [Β ]:
distribution_plot_wrt_target(data, 'Contacts_Count_12_mon', 'Attrition_Flag')
No description has been provided for this image
InΒ [Β ]:
distribution_plot_wrt_target(data, 'Credit_Limit', 'Attrition_Flag')
No description has been provided for this image
InΒ [Β ]:
distribution_plot_wrt_target(data, 'Total_Revolving_Bal', 'Attrition_Flag')
No description has been provided for this image
InΒ [Β ]:
distribution_plot_wrt_target(data, 'Avg_Open_To_Buy', 'Attrition_Flag')
No description has been provided for this image
InΒ [Β ]:
distribution_plot_wrt_target(data, 'Total_Amt_Chng_Q4_Q1', 'Attrition_Flag')
No description has been provided for this image
InΒ [Β ]:
distribution_plot_wrt_target(data, 'Total_Trans_Amt', 'Attrition_Flag')
No description has been provided for this image
InΒ [Β ]:
distribution_plot_wrt_target(data, 'Total_Trans_Ct', 'Attrition_Flag')
No description has been provided for this image
InΒ [Β ]:
distribution_plot_wrt_target(data, 'Total_Ct_Chng_Q4_Q1', 'Attrition_Flag')
No description has been provided for this image
InΒ [Β ]:
distribution_plot_wrt_target(data, 'Avg_Utilization_Ratio', 'Attrition_Flag')
No description has been provided for this image

Categorical VariablesΒΆ

InΒ [Β ]:
stacked_barplot(data, 'Gender', 'Attrition_Flag')
Attrition_Flag  Attrited Customer  Existing Customer    All
Gender                                                     
All                          1627               8500  10127
F                             930               4428   5358
M                             697               4072   4769
------------------------------------------------------------------------------------------------------------------------
No description has been provided for this image
InΒ [Β ]:
# Even though Dependent_count is technically a numerical variables, there are so few values, it can be instructure to view as a categorical variable.
stacked_barplot(data, 'Dependent_count', 'Attrition_Flag')
Attrition_Flag   Attrited Customer  Existing Customer    All
Dependent_count                                             
All                           1627               8500  10127
3                              482               2250   2732
2                              417               2238   2655
1                              269               1569   1838
4                              260               1314   1574
0                              135                769    904
5                               64                360    424
------------------------------------------------------------------------------------------------------------------------
No description has been provided for this image
InΒ [Β ]:
stacked_barplot(data, 'Education_Level', 'Attrition_Flag')
Attrition_Flag   Attrited Customer  Existing Customer   All
Education_Level                                            
All                           1371               7237  8608
Graduate                       487               2641  3128
High School                    306               1707  2013
Uneducated                     237               1250  1487
College                        154                859  1013
Doctorate                       95                356   451
Post-Graduate                   92                424   516
------------------------------------------------------------------------------------------------------------------------
No description has been provided for this image
InΒ [Β ]:
stacked_barplot(data, 'Marital_Status', 'Attrition_Flag')
Attrition_Flag  Attrited Customer  Existing Customer   All
Marital_Status                                            
All                          1498               7880  9378
Married                       709               3978  4687
Single                        668               3275  3943
Divorced                      121                627   748
------------------------------------------------------------------------------------------------------------------------
No description has been provided for this image
InΒ [Β ]:
stacked_barplot(data, 'Income_Category', 'Attrition_Flag')
Attrition_Flag   Attrited Customer  Existing Customer    All
Income_Category                                             
All                           1627               8500  10127
Less than $40K                 612               2949   3561
$40K - $60K                    271               1519   1790
$80K - $120K                   242               1293   1535
$60K - $80K                    189               1213   1402
abc                            187                925   1112
$120K +                        126                601    727
------------------------------------------------------------------------------------------------------------------------
No description has been provided for this image
InΒ [Β ]:
stacked_barplot(data, 'Card_Category', 'Attrition_Flag')
Attrition_Flag  Attrited Customer  Existing Customer    All
Card_Category                                              
All                          1627               8500  10127
Blue                         1519               7917   9436
Silver                         82                473    555
Gold                           21                 95    116
Platinum                        5                 15     20
------------------------------------------------------------------------------------------------------------------------
No description has been provided for this image

Data Pre-processingΒΆ

Missing value imputationΒΆ

As noted earlier, Education_Level and Martial_Status, have a small percentage of missing values. Income_Category also has missing values which appear as 'abc'. There are missing Education_Level values which show up as 'N/A'.

Based on our stacked bar plot of Education_Level and Attrition_Flag, we saw that Uneducated, High School, College, and Graduate all had about the same prevalence of Attrited Customer, so let's set the missing values here to the mode which is Graduate.

InΒ [Β ]:
# Set NaN values of Education_Level to 'Graduate'
data['Education_Level'] = data['Education_Level'].fillna('Graduate')
InΒ [Β ]:
print(data.Education_Level.value_counts(dropna=False))
Education_Level
Graduate         4647
High School      2013
Uneducated       1487
College          1013
Post-Graduate     516
Doctorate         451
Name: count, dtype: int64

Based on our stacked bar plot of Marital_status and Attrition_Flag, we saw that the three values all had about the same prevalence of Attrited Customer, so let's set the missing values here to the mode which is Married.

InΒ [Β ]:
# Set NaN values of Education_Level to 'Graduate'
data['Marital_Status'] = data['Marital_Status'].fillna('Married')
InΒ [Β ]:
print(data.Marital_Status.value_counts(dropna=False))
Marital_Status
Married     5436
Single      3943
Divorced     748
Name: count, dtype: int64

To decide what to do with the abc values in Income_Category, we could reasonably assume some correlation between Income_Category and Total_Trans_Amt and Customer_Age. In the below violin plots we can see that "Less than $40K" is the closest match for Total_Trans_Amt, while for Customer_Age, "Less than $40K" is also a good match. This value also happens to be the mode, so let's go with these value for the missing values.

InΒ [Β ]:
plt.figure(figsize=(14, 8))
sns.violinplot(data=data, x='Income_Category', y='Total_Trans_Amt', orient='v');
No description has been provided for this image
InΒ [Β ]:
plt.figure(figsize=(14, 8))
sns.violinplot(data=data, x='Income_Category', y='Customer_Age', orient='v');
No description has been provided for this image
InΒ [Β ]:
# Replace 'abc' with Less than $40K
data['Income_Category'].replace('abc', 'Less than $40K', inplace=True)

Conversion from Object to CategoryΒΆ

Lets convert the columns with an 'object' datatype into categorical variables, to reduce the space required to store the dataframe and improve processing efficieny.

InΒ [Β ]:
for feature in data.columns: # Loop through all columns in the dataframe
    if data[feature].dtype == 'object': # Only apply for columns with categorical strings
        data[feature] = pd.Categorical(data[feature]) # Convert the column to a categorical type
InΒ [Β ]:
# Confirm that column of type object have been converted to category
data.info()
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 10127 entries, 0 to 10126
Data columns (total 20 columns):
 #   Column                    Non-Null Count  Dtype   
---  ------                    --------------  -----   
 0   Attrition_Flag            10127 non-null  category
 1   Customer_Age              10127 non-null  int64   
 2   Gender                    10127 non-null  category
 3   Dependent_count           10127 non-null  int64   
 4   Education_Level           10127 non-null  category
 5   Marital_Status            10127 non-null  category
 6   Income_Category           10127 non-null  category
 7   Card_Category             10127 non-null  category
 8   Months_on_book            10127 non-null  int64   
 9   Total_Relationship_Count  10127 non-null  int64   
 10  Months_Inactive_12_mon    10127 non-null  int64   
 11  Contacts_Count_12_mon     10127 non-null  int64   
 12  Credit_Limit              10127 non-null  float64 
 13  Total_Revolving_Bal       10127 non-null  int64   
 14  Avg_Open_To_Buy           10127 non-null  float64 
 15  Total_Amt_Chng_Q4_Q1      10127 non-null  float64 
 16  Total_Trans_Amt           10127 non-null  int64   
 17  Total_Trans_Ct            10127 non-null  int64   
 18  Total_Ct_Chng_Q4_Q1       10127 non-null  float64 
 19  Avg_Utilization_Ratio     10127 non-null  float64 
dtypes: category(6), float64(5), int64(9)
memory usage: 1.1 MB
InΒ [Β ]:
print(data.Attrition_Flag.value_counts(dropna=False))
print()
print(data.Gender.value_counts(dropna=False))
print()
print(data.Education_Level.value_counts(dropna=False))
print()
print(data.Marital_Status.value_counts(dropna=False))
print()
print(data.Income_Category.value_counts(dropna=False))
print()
print(data.Card_Category.value_counts(dropna=False))
Attrition_Flag
Existing Customer    8500
Attrited Customer    1627
Name: count, dtype: int64

Gender
F    5358
M    4769
Name: count, dtype: int64

Education_Level
Graduate         4647
High School      2013
Uneducated       1487
College          1013
Post-Graduate     516
Doctorate         451
Name: count, dtype: int64

Marital_Status
Married     5436
Single      3943
Divorced     748
Name: count, dtype: int64

Income_Category
Less than $40K    4673
$40K - $60K       1790
$80K - $120K      1535
$60K - $80K       1402
$120K +            727
Name: count, dtype: int64

Card_Category
Blue        9436
Silver       555
Gold         116
Platinum      20
Name: count, dtype: int64
InΒ [Β ]:
print(data.Income_Category.value_counts(dropna=False))
Income_Category
Less than $40K    4673
$40K - $60K       1790
$80K - $120K      1535
$60K - $80K       1402
$120K +            727
Name: count, dtype: int64

One-Hot EncodingΒΆ

InΒ [Β ]:
replaceStruct = {
                "Attrition_Flag": {"Existing Customer": 1, "Attrited Customer": 2},
                "Gender": {"M": 1, "F":2},
                 "Education_Level": {"Uneducated": 1, "High School":2 , "College": 3, "Graduate": 4, "Post-Graduate": 5,"Doctorate": 6},
                 "Marital_Status": {"Married": 1, "Single":2 , "Divorced": 3},
                 "Income_Category":     {"Less than $40K": 1, "$40K - $60K": 2 ,"$60K - $80K": 3 ,"$80K - $120K": 4 ,"$120K +": 5},
                "Card_Category":     {"Blue": 1, "Silver": 2, "Gold": 3, "Platinum": 4 }
                    }
oneHotCols=["Attrition_Flag","Gender","Education_Level","Marital_Status","Income_Category", "Card_Category"]
InΒ [Β ]:
data=data.replace(replaceStruct)
data=pd.get_dummies(data, columns=oneHotCols, drop_first=True)
data.head(10)
Out[Β ]:
Customer_Age Dependent_count Months_on_book Total_Relationship_Count Months_Inactive_12_mon Contacts_Count_12_mon Credit_Limit Total_Revolving_Bal Avg_Open_To_Buy Total_Amt_Chng_Q4_Q1 Total_Trans_Amt Total_Trans_Ct Total_Ct_Chng_Q4_Q1 Avg_Utilization_Ratio Attrition_Flag_1 Gender_1 Education_Level_6 Education_Level_4 Education_Level_2 Education_Level_5 Education_Level_1 Marital_Status_1 Marital_Status_2 Income_Category_2 Income_Category_3 Income_Category_4 Income_Category_1 Income_Category_abc Card_Category_3 Card_Category_4 Card_Category_2
0 45 3 39 5 1 3 12691.00 777 11914.00 1.33 1144 42 1.62 0.06 True True False False True False False True False False True False False False False False False
1 49 5 44 6 1 2 8256.00 864 7392.00 1.54 1291 33 3.71 0.10 True False False True False False False False True False False False True False False False False
2 51 3 36 4 1 0 3418.00 0 3418.00 2.59 1887 20 2.33 0.00 True True False True False False False True False False False True False False False False False
3 40 4 34 3 4 1 3313.00 2517 796.00 1.41 1171 20 2.33 0.76 True False False False True False False True False False False False True False False False False
4 40 3 21 5 1 0 4716.00 0 4716.00 2.17 816 28 2.50 0.00 True True False False False False True True False False True False False False False False False
5 44 2 36 3 1 2 4010.00 1247 2763.00 1.38 1088 24 0.85 0.31 True True False True False False False True False True False False False False False False False
6 51 4 46 6 1 3 34516.00 2264 32252.00 1.98 1330 31 0.72 0.07 True True False True False False False True False False False False False False True False False
7 32 0 27 2 2 2 29081.00 1396 27685.00 2.20 1538 36 0.71 0.05 True True False False True False False True False False True False False False False False True
8 37 3 36 5 2 0 22352.00 2517 19835.00 3.35 1350 24 1.18 0.11 True True False False False False True False True False True False False False False False False
9 48 2 36 6 3 3 11656.00 1677 9979.00 1.52 1441 32 0.88 0.14 True True False True False False False False True False False True False False False False False

Outlier TreatmentΒΆ

Although some of the variables have outliers as event in the EDA, we will keep these values since they are valid data.

Model BuildingΒΆ

Model evaluation criterionΒΆ

The nature of predictions made by the classification model will translate as follows:

  • True positives (TP) are failures correctly predicted by the model.
  • False negatives (FN) are real failures in a generator where there is no detection by model.
  • False positives (FP) are failure detections in a generator where there is no failure.

Which metric to optimize?

  • We need to choose the metric which will ensure that the maximum number of generator failures are predicted correctly by the model.
  • We would want Recall to be maximized as greater the Recall, the higher the chances of minimizing false negatives.
  • We want to minimize false negatives because if a model predicts that a machine will have no failure when there will be a failure, it will increase the maintenance cost.

Let's define a function to output different metrics (including recall) on the train and test set and a function to show confusion matrix so that we do not have to use the same code repetitively while evaluating models.

InΒ [Β ]:
# defining a function to compute different metrics to check performance of a classification model built using sklearn
def model_performance_classification_sklearn(model, predictors, target):
    """
    Function to compute different metrics to check classification model performance

    model: classifier
    predictors: independent variables
    target: dependent variable
    """

    # predicting using the independent variables
    pred = model.predict(predictors)

    acc = accuracy_score(target, pred)  # to compute Accuracy
    recall = recall_score(target, pred)  # to compute Recall
    precision = precision_score(target, pred)  # to compute Precision
    f1 = f1_score(target, pred)  # to compute F1-score

    # creating a dataframe of metrics
    df_perf = pd.DataFrame(
        {
            "Accuracy": acc,
            "Recall": recall,
            "Precision": precision,
            "F1": f1

        },
        index=[0],
    )

    return df_perf
InΒ [Β ]:
def plot_confusion_matrix(model, predictors, target):
    """
    To plot the confusion_matrix with percentages
    model: classifier
    predictors: independent variables
    target: dependent variable
    """
    # Predict the target values using the provided model and predictors
    y_pred = model.predict(predictors)
    # Compute the confusion matrix comparing the true target values with the predicted values
    cm = confusion_matrix(target, y_pred)
    # Create labels for each cell in the confusion matrix with both count and percentage
    labels = np.asarray(
        [
            ["{0:0.0f}".format(item) + "\n{0:.2%}".format(item / cm.flatten().sum())]
            for item in cm.flatten()
        ]
    ).reshape(2, 2)    # reshaping to a matrix
    # Set the figure size for the plot
    plt.figure(figsize=(6, 4))
    # Plot the confusion matrix as a heatmap with the labels
    sns.heatmap(cm, annot=labels, fmt="")
    # Add a label to the y-axis
    plt.ylabel("True label")
    # Add a label to the x-axis
    plt.xlabel("Predicted label")

Split the data into train and test setsΒΆ

  • Since we have an imbalance in the distribution of the target class (Attrition_Flag), it is good to use stratified sampling to ensure that relative class frequencies are approximately preserved in train and test sets.
  • We do this by setting the stratify parameter to target variable in the train_test_split function.
InΒ [Β ]:
X = data.drop("Attrition_Flag_1" , axis=1)
y = data.pop("Attrition_Flag_1")
InΒ [Β ]:
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=.30, random_state=1,stratify=y)
InΒ [Β ]:
#Show the proportion of value in y
y_train.value_counts(normalize=False)
Out[Β ]:
count
Attrition_Flag_1
True 5949
False 1139

Note: For this project I elected to go with the simpler approach of using just a Train and Test split. I acknowledge though another approach is to split the data per Train, Validation, and Test, where Test can be reserved for a check on the final best model.

Model Building with Original dataΒΆ

Build 5 models (from decision trees, bagging and boosting methods) - Comment on the model performance * You can choose NOT to build XGBoost if you are facing issues with the installation

BaggingΒΆ

InΒ [Β ]:
bagging_estimator=BaggingClassifier(random_state=1)
bagging_estimator.fit(X_train,y_train)
Out[Β ]:
BaggingClassifier(random_state=1)
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
BaggingClassifier(random_state=1)
InΒ [Β ]:
bagging_estimator_train_perf = model_performance_classification_sklearn(bagging_estimator, X_train, y_train)
print("Training performance \n",bagging_estimator_train_perf)
Training performance 
    Accuracy  Recall  Precision   F1
0      1.00    1.00       1.00 1.00
InΒ [Β ]:
bagging_estimator_test_perf = model_performance_classification_sklearn(bagging_estimator, X_test, y_test)
print("Testing performance \n",bagging_estimator_test_perf)
Testing performance 
    Accuracy  Recall  Precision   F1
0      0.95    0.97       0.97 0.97

Random ForestΒΆ

InΒ [Β ]:
rf = RandomForestClassifier(random_state=1)
rf.fit(X_train,y_train)
Out[Β ]:
RandomForestClassifier(random_state=1)
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
RandomForestClassifier(random_state=1)
InΒ [Β ]:
rf_train_perf = model_performance_classification_sklearn(rf, X_train, y_train)
print("Training performance \n",rf_train_perf)
Training performance 
    Accuracy  Recall  Precision   F1
0      1.00    1.00       1.00 1.00
InΒ [Β ]:
rf_test_perf = model_performance_classification_sklearn(rf, X_test, y_test)
print("Testing performance \n",rf_test_perf)
Testing performance 
    Accuracy  Recall  Precision   F1
0      0.95    0.99       0.96 0.97

AdaBoostΒΆ

InΒ [Β ]:
ab_regressor = AdaBoostClassifier(random_state=1)
ab_regressor.fit(X_train,y_train)
Out[Β ]:
AdaBoostClassifier(random_state=1)
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
AdaBoostClassifier(random_state=1)
InΒ [Β ]:
#Train performance
ab_regressor_train_perf = model_performance_classification_sklearn(ab_regressor, X_train, y_train)
print("Training performance \n",ab_regressor_train_perf)
Training performance 
    Accuracy  Recall  Precision   F1
0      0.96    0.98       0.97 0.98
InΒ [Β ]:
#Test performance
ab_regressor_test_perf = model_performance_classification_sklearn(ab_regressor, X_test, y_test)
print("Testing performance \n",ab_regressor_test_perf)
Testing performance 
    Accuracy  Recall  Precision   F1
0      0.96    0.98       0.97 0.98

GradientBoostingΒΆ

InΒ [Β ]:
#GradientBoost Classifier
gb_estimator = GradientBoostingClassifier(random_state=1)
gb_estimator.fit(X_train,y_train)
Out[Β ]:
GradientBoostingClassifier(random_state=1)
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
GradientBoostingClassifier(random_state=1)
InΒ [Β ]:
#Train performance
gb_estimator_train_perf = model_performance_classification_sklearn(gb_estimator, X_train, y_train)
print("Training performance \n",gb_estimator_train_perf)
Training performance 
    Accuracy  Recall  Precision   F1
0      0.98    0.99       0.98 0.99
InΒ [Β ]:
#Test performance
gb_estimator_test_perf = model_performance_classification_sklearn(gb_estimator, X_test, y_test)
print("Testing performance \n",gb_estimator_test_perf)
Testing performance 
    Accuracy  Recall  Precision   F1
0      0.96    0.99       0.97 0.98

Decision TreeΒΆ

InΒ [Β ]:
dtree=DecisionTreeClassifier(random_state=1)
dtree.fit(X_train,y_train)
Out[Β ]:
DecisionTreeClassifier(random_state=1)
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
DecisionTreeClassifier(random_state=1)
InΒ [Β ]:
dtree_model_train_perf = model_performance_classification_sklearn(dtree, X_train, y_train)
print("Training performance \n",dtree_model_train_perf)
Training performance 
    Accuracy  Recall  Precision   F1
0      1.00    1.00       1.00 1.00
InΒ [Β ]:
dtree_model_test_perf = model_performance_classification_sklearn(dtree, X_test, y_test)
print("Testing performance \n",dtree_model_test_perf)
Testing performance 
    Accuracy  Recall  Precision   F1
0      0.94    0.96       0.96 0.96

Model Building with Oversampled dataΒΆ

Oversample the train data - Build 5 models (from decision trees, bagging and boosting methods) - Comment on the model performance * You can choose NOT to build XGBoost if you are facing issues with the installation

InΒ [Β ]:
# Synthetic Minority Over Sampling Technique
sm = SMOTE(sampling_strategy=1, k_neighbors=5, random_state=1)
X_train_over, y_train_over = sm.fit_resample(X_train, y_train)
InΒ [Β ]:
print("Before OverSampling, count of label '1': {}".format(sum(y_train == 1)))
print("Before OverSampling, count of label '0': {} \n".format(sum(y_train == 0)))

print("After OverSampling, count of label '1': {}".format(sum(y_train_over == 1)))
print("After OverSampling, count of label '0': {} \n".format(sum(y_train_over == 0)))

print("After OverSampling, the shape of train_X: {}".format(X_train_over.shape))
print("After OverSampling, the shape of train_y: {} \n".format(y_train_over.shape))
Before OverSampling, count of label '1': 5949
Before OverSampling, count of label '0': 1139 

After OverSampling, count of label '1': 5949
After OverSampling, count of label '0': 5949 

After OverSampling, the shape of train_X: (11898, 30)
After OverSampling, the shape of train_y: (11898,) 

Bagging (Over Sampled)ΒΆ

InΒ [Β ]:
bagging_estimator_o=BaggingClassifier(random_state=1)
bagging_estimator_o.fit(X_train_over,y_train_over)
Out[Β ]:
BaggingClassifier(random_state=1)
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
BaggingClassifier(random_state=1)
InΒ [Β ]:
bagging_estimator_o_train_perf = model_performance_classification_sklearn(bagging_estimator_o, X_train_over, y_train_over)
print("Training performance \n",bagging_estimator_o_train_perf)
Training performance 
    Accuracy  Recall  Precision   F1
0      1.00    1.00       1.00 1.00
InΒ [Β ]:
bagging_estimator_o_test_perf = model_performance_classification_sklearn(bagging_estimator_o, X_test, y_test)
print("Testing performance \n",bagging_estimator_o_test_perf)
Testing performance 
    Accuracy  Recall  Precision   F1
0      0.95    0.96       0.98 0.97

Random Forest (Over Sampled)ΒΆ

InΒ [Β ]:
rf_o = RandomForestClassifier(random_state=1)
rf_o.fit(X_train_over,y_train_over)
Out[Β ]:
RandomForestClassifier(random_state=1)
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
RandomForestClassifier(random_state=1)
InΒ [Β ]:
rf_train_perf_o = model_performance_classification_sklearn(rf_o, X_train_over, y_train_over)
print("Training performance \n",rf_train_perf_o)
Training performance 
    Accuracy  Recall  Precision   F1
0      1.00    1.00       1.00 1.00
InΒ [Β ]:
rf_test_perf_o = model_performance_classification_sklearn(rf_o, X_test, y_test)
print("Testing performance \n",rf_test_perf_o)
Testing performance 
    Accuracy  Recall  Precision   F1
0      0.96    0.97       0.98 0.97

AdaBoost (Over Sampled)ΒΆ

InΒ [Β ]:
ab_regressor_o = AdaBoostClassifier(random_state=1)
ab_regressor_o.fit(X_train_over,y_train_over)
Out[Β ]:
AdaBoostClassifier(random_state=1)
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
AdaBoostClassifier(random_state=1)
InΒ [Β ]:
#Train performance
ab_regressor_o_train_perf = model_performance_classification_sklearn(ab_regressor_o, X_train_over, y_train_over)
print("Training performance \n",ab_regressor_o_train_perf)
Training performance 
    Accuracy  Recall  Precision   F1
0      0.96    0.96       0.96 0.96
InΒ [Β ]:
#Test performance
ab_regressor_o_test_perf = model_performance_classification_sklearn(ab_regressor_o, X_test, y_test)
print("Testing performance \n",ab_regressor_o_test_perf)
Testing performance 
    Accuracy  Recall  Precision   F1
0      0.95    0.96       0.98 0.97

GradientBoosting (Over Sampled)ΒΆ

InΒ [Β ]:
#GradientBoost Classifier
gb_estimator_o = GradientBoostingClassifier(random_state=1)
gb_estimator_o.fit(X_train_over,y_train_over)
Out[Β ]:
GradientBoostingClassifier(random_state=1)
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
GradientBoostingClassifier(random_state=1)
InΒ [Β ]:
#Train performance
gb_estimator_o_train_perf = model_performance_classification_sklearn(gb_estimator_o, X_train_over, y_train_over)
print("Training performance \n",gb_estimator_o_train_perf)
Training performance 
    Accuracy  Recall  Precision   F1
0      0.98    0.97       0.98 0.98
InΒ [Β ]:
#Test performance
gb_estimator_o_test_perf = model_performance_classification_sklearn(gb_estimator_o, X_test, y_test)
print("Testing performance \n",gb_estimator_o_test_perf)
Testing performance 
    Accuracy  Recall  Precision   F1
0      0.96    0.97       0.98 0.98

Decision Tree (Over Sampled)ΒΆ

InΒ [Β ]:
dtree_o=DecisionTreeClassifier(random_state=1)
dtree_o.fit(X_train_over,y_train_over)
Out[Β ]:
DecisionTreeClassifier(random_state=1)
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
DecisionTreeClassifier(random_state=1)
InΒ [Β ]:
dtree_model_o_train_perf = model_performance_classification_sklearn(dtree_o, X_train_over, y_train_over)
print("Training performance \n",dtree_model_o_train_perf)
Training performance 
    Accuracy  Recall  Precision   F1
0      1.00    1.00       1.00 1.00
InΒ [Β ]:
dtree_model_o_test_perf = model_performance_classification_sklearn(dtree_o, X_test, y_test)
print("Testing performance \n",dtree_model_o_test_perf)
Testing performance 
    Accuracy  Recall  Precision   F1
0      0.93    0.95       0.97 0.96

Model Building with Undersampled dataΒΆ

Undersample the train data - Build 5 models (from decision trees, bagging and boosting methods) - Comment on the model performance * You can choose NOT to build XGBoost if you are facing issues with the installation

InΒ [Β ]:
# Random undersampler for under sampling the data
rus = RandomUnderSampler(random_state=1, sampling_strategy=1)
X_train_un, y_train_un = rus.fit_resample(X_train, y_train)
InΒ [Β ]:
print("Before UnderSampling, count of label '1': {}".format(sum(y_train == 1)))
print("Before UnderSampling, count of label '0': {} \n".format(sum(y_train == 0)))

print("After UnderSampling, count of label '1': {}".format(sum(y_train_un == 1)))
print("After UnderSampling, count of label '0': {} \n".format(sum(y_train_un == 0)))

print("After UnderSampling, the shape of train_X: {}".format(X_train_un.shape))
print("After UnderSampling, the shape of train_y: {} \n".format(y_train_un.shape))
Before UnderSampling, count of label '1': 5949
Before UnderSampling, count of label '0': 1139 

After UnderSampling, count of label '1': 1139
After UnderSampling, count of label '0': 1139 

After UnderSampling, the shape of train_X: (2278, 30)
After UnderSampling, the shape of train_y: (2278,) 

Bagging (Under Sampled)ΒΆ

InΒ [Β ]:
bagging_estimator_u=BaggingClassifier(random_state=1)
bagging_estimator_u.fit(X_train_un,y_train_un)
Out[Β ]:
BaggingClassifier(random_state=1)
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
BaggingClassifier(random_state=1)
InΒ [Β ]:
bagging_estimator_u_train_perf = model_performance_classification_sklearn(bagging_estimator_u, X_train_un, y_train_un)
print("Training performance \n",bagging_estimator_u_train_perf)
Training performance 
    Accuracy  Recall  Precision   F1
0      1.00    0.99       1.00 1.00
InΒ [Β ]:
bagging_estimator_u_test_perf = model_performance_classification_sklearn(bagging_estimator_u, X_test, y_test)
print("Testing performance \n",bagging_estimator_u_test_perf)
Testing performance 
    Accuracy  Recall  Precision   F1
0      0.91    0.91       0.98 0.95

Random Forest (Under Sampled)ΒΆ

InΒ [Β ]:
rf_u = RandomForestClassifier(random_state=1)
rf_u.fit(X_train_un,y_train_un)
Out[Β ]:
RandomForestClassifier(random_state=1)
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
RandomForestClassifier(random_state=1)
InΒ [Β ]:
rf_train_perf_u = model_performance_classification_sklearn(rf_u, X_train_un, y_train_un)
print("Training performance \n",rf_train_perf_u)
Training performance 
    Accuracy  Recall  Precision   F1
0      1.00    1.00       1.00 1.00
InΒ [Β ]:
rf_test_perf_u = model_performance_classification_sklearn(rf_u, X_test, y_test)
print("Testing performance \n",rf_test_perf_u)
Testing performance 
    Accuracy  Recall  Precision   F1
0      0.92    0.92       0.99 0.95

AdaBoost (Under Sampled)ΒΆ

InΒ [Β ]:
ab_regressor_u = AdaBoostClassifier(random_state=1)
ab_regressor_u.fit(X_train_un,y_train_un)
Out[Β ]:
AdaBoostClassifier(random_state=1)
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
AdaBoostClassifier(random_state=1)
InΒ [Β ]:
#Train performance
ab_regressor_u_train_perf = model_performance_classification_sklearn(ab_regressor_u, X_train_un, y_train_un)
print("Training performance \n",ab_regressor_u_train_perf)
Training performance 
    Accuracy  Recall  Precision   F1
0      0.95    0.95       0.96 0.95
InΒ [Β ]:
#Test performance
ab_regressor_u_test_perf = model_performance_classification_sklearn(ab_regressor_u, X_test, y_test)
print("Testing performance \n",ab_regressor_u_test_perf)
Testing performance 
    Accuracy  Recall  Precision   F1
0      0.93    0.93       0.99 0.96

GradientBoosting (Under Sampled)ΒΆ

InΒ [Β ]:
#GradientBoost Classifier
gb_estimator_u = GradientBoostingClassifier(random_state=1)
gb_estimator_u.fit(X_train_un,y_train_un)
Out[Β ]:
GradientBoostingClassifier(random_state=1)
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
GradientBoostingClassifier(random_state=1)
InΒ [Β ]:
#Train performance
gb_estimator_u_train_perf = model_performance_classification_sklearn(gb_estimator_u, X_train_un, y_train_un)
print("Training performance \n",gb_estimator_u_train_perf)
Training performance 
    Accuracy  Recall  Precision   F1
0      0.98    0.97       0.98 0.98
InΒ [Β ]:
#Test performance
gb_estimator_u_test_perf = model_performance_classification_sklearn(gb_estimator_u, X_test, y_test)
print("Testing performance \n",gb_estimator_u_test_perf)
Testing performance 
    Accuracy  Recall  Precision   F1
0      0.95    0.95       0.99 0.97

Decision Tree (Under Sampled)ΒΆ

InΒ [Β ]:
dtree_u=DecisionTreeClassifier(random_state=1)
dtree_u.fit(X_train_un,y_train_un)
Out[Β ]:
DecisionTreeClassifier(random_state=1)
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
DecisionTreeClassifier(random_state=1)
InΒ [Β ]:
dtree_model_u_train_perf = model_performance_classification_sklearn(dtree_u, X_train_un, y_train_un)
print("Training performance \n",dtree_model_o_train_perf)
Training performance 
    Accuracy  Recall  Precision   F1
0      1.00    1.00       1.00 1.00
InΒ [Β ]:
dtree_model_u_test_perf = model_performance_classification_sklearn(dtree_u, X_test, y_test)
print("Testing performance \n",dtree_model_u_test_perf)
Testing performance 
    Accuracy  Recall  Precision   F1
0      0.90    0.91       0.98 0.94

Model Comparison to Choose Best 3 Models for Next StepΒΆ

InΒ [Β ]:
# training performance comparison

models_train_comp_df = pd.concat(
    [bagging_estimator_train_perf.T, bagging_estimator_o_train_perf.T, bagging_estimator_u_train_perf.T,
    rf_train_perf.T, rf_train_perf_o.T, rf_train_perf_u.T,
     ab_regressor_train_perf.T, ab_regressor_o_train_perf.T, ab_regressor_u_train_perf.T,
     gb_estimator_train_perf.T, gb_estimator_o_train_perf.T, gb_estimator_u_train_perf.T,
     dtree_model_train_perf.T, dtree_model_o_train_perf.T, dtree_model_u_train_perf.T],
    axis=1,
)

models_train_comp_df.columns = [
    "Bagging Estimator",
    "Bagging Estimator O",
    "Bagging Estimator U",
    "Random Forest",
    "Random Forest O",
    "Random Forest U",
    "AdaBoost",
    "AdaBoost O",
    "AdaBoost U",
    "GradientBoost",
    "GradientBoost O",
    "GradientBoost U",
    "Decision Tree",
    "Decision Tree O",
    "Decision Tree U"
]

# Sort columns based on the 'Recall' row
models_train_comp_df = models_train_comp_df.T.sort_values(by='Recall', ascending=False).T


print("Training performance comparison:")
models_train_comp_df
Training performance comparison:
Out[Β ]:
Random Forest Random Forest O Random Forest U Decision Tree Decision Tree O Decision Tree U Bagging Estimator Bagging Estimator O Bagging Estimator U GradientBoost AdaBoost GradientBoost O GradientBoost U AdaBoost O AdaBoost U
Accuracy 1.00 1.00 1.00 1.00 1.00 1.00 1.00 1.00 1.00 0.98 0.96 0.98 0.98 0.96 0.95
Recall 1.00 1.00 1.00 1.00 1.00 1.00 1.00 1.00 0.99 0.99 0.98 0.97 0.97 0.96 0.95
Precision 1.00 1.00 1.00 1.00 1.00 1.00 1.00 1.00 1.00 0.98 0.97 0.98 0.98 0.96 0.96
F1 1.00 1.00 1.00 1.00 1.00 1.00 1.00 1.00 1.00 0.99 0.98 0.98 0.98 0.96 0.95
InΒ [Β ]:
# test performance comparison

models_test_comp_df = pd.concat(
    [bagging_estimator_test_perf.T, bagging_estimator_o_test_perf.T, bagging_estimator_u_test_perf.T,
    rf_test_perf.T, rf_test_perf_o.T, rf_test_perf_u.T,
     ab_regressor_test_perf.T, ab_regressor_o_test_perf.T, ab_regressor_u_test_perf.T,
     gb_estimator_test_perf.T, gb_estimator_o_test_perf.T, gb_estimator_u_test_perf.T,
     dtree_model_test_perf.T, dtree_model_o_test_perf.T, dtree_model_u_test_perf.T],
    axis=1,
)

models_test_comp_df.columns = [
    "Bagging Estimator",
    "Bagging Estimator O",
    "Bagging Estimator U",
    "Random Forest",
    "Random Forest O",
    "Random Forest U",
    "AdaBoost",
    "AdaBoost O",
    "AdaBoost U",
    "GradientBoost",
    "GradientBoost O",
    "GradientBoost U",
    "Decision Tree",
    "Decision Tree O",
    "Decision Tree U"
]

# Sort columns based on the 'Recall' row
models_test_comp_df = models_test_comp_df.T.sort_values(by='Recall', ascending=False).T


print("Testing performance comparison:")
models_test_comp_df
Testing performance comparison:
Out[Β ]:
Random Forest GradientBoost AdaBoost Random Forest O GradientBoost O Bagging Estimator Decision Tree Bagging Estimator O AdaBoost O Decision Tree O GradientBoost U AdaBoost U Random Forest U Bagging Estimator U Decision Tree U
Accuracy 0.95 0.96 0.96 0.96 0.96 0.95 0.94 0.95 0.95 0.93 0.95 0.93 0.92 0.91 0.90
Recall 0.99 0.99 0.98 0.97 0.97 0.97 0.96 0.96 0.96 0.95 0.95 0.93 0.92 0.91 0.91
Precision 0.96 0.97 0.97 0.98 0.98 0.97 0.96 0.98 0.98 0.97 0.99 0.99 0.99 0.98 0.98
F1 0.97 0.98 0.98 0.97 0.98 0.97 0.96 0.97 0.97 0.96 0.97 0.96 0.95 0.95 0.94

In this particular problem, Recall is the best peformance measure to optimize, since attempts to maxium the True Positive Rate, so we don't overlook or miss any customer who is a risk of cancelling their credit cards. Per the above our top 3 performers are

  • Random Forest
  • GradientBoost
  • AdaBoost

HyperparameterTuningΒΆ

Choose models that might perform better after tuning (tune at least 3 models out of 15 built in the previous steps) - Provide proper reasoning for tuning that model - Tune the best 3 models obtained above using randomized search and metric of interest - Check the performance of 3 tuned models.

Sample Parameter GridsΒΆ

Note

  1. Sample parameter grids have been provided to do necessary hyperparameter tuning. These sample grids are expected to provide a balance between model performance improvement and execution time. One can extend/reduce the parameter grid based on execution time and system configuration.
  • Please note that if the parameter grid is extended to improve the model performance further, the execution time will increase
  • For Gradient Boosting:
param_grid = {
    "init": [AdaBoostClassifier(random_state=1),DecisionTreeClassifier(random_state=1)],
    "n_estimators": np.arange(50,110,25),
    "learning_rate": [0.01,0.1,0.05],
    "subsample":[0.7,0.9],
    "max_features":[0.5,0.7,1],
}
  • For Adaboost:
param_grid = {
    "n_estimators": np.arange(50,110,25),
    "learning_rate": [0.01,0.1,0.05],
    "base_estimator": [
        DecisionTreeClassifier(max_depth=2, random_state=1),
        DecisionTreeClassifier(max_depth=3, random_state=1),
    ],
}
  • For Bagging Classifier:
param_grid = {
    'max_samples': [0.8,0.9,1],
    'max_features': [0.7,0.8,0.9],
    'n_estimators' : [30,50,70],
}
  • For Random Forest:
param_grid = {
    "n_estimators": [50,110,25],
    "min_samples_leaf": np.arange(1, 4),
    "max_features": [np.arange(0.3, 0.6, 0.1),'sqrt'],
    "max_samples": np.arange(0.4, 0.7, 0.1)
}
  • For Decision Trees:
param_grid = {
    'max_depth': np.arange(2,6),
    'min_samples_leaf': [1, 4, 7],
    'max_leaf_nodes' : [10, 15],
    'min_impurity_decrease': [0.0001,0.001]
}
  • For XGBoost (optional):
param_grid={'n_estimators':np.arange(50,110,25),
            'scale_pos_weight':[1,2,5],
            'learning_rate':[0.01,0.1,0.05],
            'gamma':[1,3],
            'subsample':[0.7,0.9]
}

Tuning of the Random Forest ModelΒΆ

InΒ [Β ]:
# defining model
rf = RandomForestClassifier(random_state=1)

# Parameter grid to pass in RandomSearchCV
param_grid = {
    "n_estimators": [50,110,25],
    "min_samples_leaf": np.arange(1, 4),
    "max_features": [np.arange(0.3, 0.6, 0.1),'sqrt'],
    "max_samples": np.arange(0.4, 0.7, 0.1)
}

# Type of scoring used to compare parameter combinations
scorer = metrics.make_scorer(metrics.recall_score)

#Calling RandomizedSearchCV
randomized_cv = RandomizedSearchCV(estimator=rf, param_distributions=param_grid, n_iter=10, n_jobs = -1, scoring=scorer, cv=5, random_state=1)

#Fitting parameters in RandomizedSearchCV
randomized_cv.fit(X_train,y_train)

print("Best parameters are {} with CV score={}:" .format(randomized_cv.best_params_,randomized_cv.best_score_))
Best parameters are {'n_estimators': 50, 'min_samples_leaf': 2, 'max_samples': 0.4, 'max_features': 'sqrt'} with CV score=0.9873935444657258:
InΒ [Β ]:
# Set the model to the best combination of parameters
rf_tuned = RandomForestClassifier(
    random_state=1,
    n_estimators=50,
    min_samples_leaf=2,
    max_samples=0.4,
    max_features='sqrt'
)

# Fit the best algorithm to the data
rf_tuned.fit(X_train, y_train)
Out[Β ]:
RandomForestClassifier(max_samples=0.4, min_samples_leaf=2, n_estimators=50,
                       random_state=1)
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
RandomForestClassifier(max_samples=0.4, min_samples_leaf=2, n_estimators=50,
                       random_state=1)
InΒ [Β ]:
rf_tuned_train_perf = model_performance_classification_sklearn(rf_tuned, X_train, y_train)
print("Training performance \n",rf_tuned_train_perf)
Training performance 
    Accuracy  Recall  Precision   F1
0      0.98    0.99       0.98 0.99
InΒ [Β ]:
rf_tuned_test_perf = model_performance_classification_sklearn(rf_tuned, X_test, y_test)
print("Testing performance \n",rf_tuned_test_perf)
Training performance 
    Accuracy  Recall  Precision   F1
0      0.94    0.99       0.95 0.97

Tuning of the GradientBoost ModelΒΆ

InΒ [Β ]:
# defining model
gb_estimator = GradientBoostingClassifier(random_state=1)

# Parameter grid to pass in RandomSearchCV
param_grid = {
    "init": [AdaBoostClassifier(random_state=1),DecisionTreeClassifier(random_state=1)],
    "n_estimators": np.arange(50,110,25),
    "learning_rate": [0.01,0.1,0.05],
    "subsample":[0.7,0.9],
    "max_features":[0.5,0.7,1],
}

# Type of scoring used to compare parameter combinations
scorer = metrics.make_scorer(metrics.recall_score)

#Calling RandomizedSearchCV
randomized_cv = RandomizedSearchCV(estimator=gb_estimator, param_distributions=param_grid, n_iter=10, n_jobs = -1, scoring=scorer, cv=5, random_state=1)

#Fitting parameters in RandomizedSearchCV
randomized_cv.fit(X_train,y_train)

print("Best parameters are {} with CV score={}:" .format(randomized_cv.best_params_,randomized_cv.best_score_))
Best parameters are {'subsample': 0.7, 'n_estimators': 50, 'max_features': 1, 'learning_rate': 0.05, 'init': AdaBoostClassifier(random_state=1)} with CV score=0.9996637241944718:
InΒ [Β ]:
# Set the model to the best combination of parameters
gb_estimator_tuned = GradientBoostingClassifier(
    random_state=1,
    subsample=0.7,
    n_estimators=50,
    max_features=1,
    learning_rate=0.05,
    init=AdaBoostClassifier(random_state=1)
)

# Fit the best algorithm to the data
gb_estimator_tuned.fit(X_train, y_train)
Out[Β ]:
GradientBoostingClassifier(init=AdaBoostClassifier(random_state=1),
                           learning_rate=0.05, max_features=1, n_estimators=50,
                           random_state=1, subsample=0.7)
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
GradientBoostingClassifier(init=AdaBoostClassifier(random_state=1),
                           learning_rate=0.05, max_features=1, n_estimators=50,
                           random_state=1, subsample=0.7)
AdaBoostClassifier(random_state=1)
AdaBoostClassifier(random_state=1)
InΒ [Β ]:
gb_tuned_train_perf = model_performance_classification_sklearn(gb_estimator_tuned, X_train, y_train)
print("Training performance \n",gb_tuned_train_perf)
Training performance 
    Accuracy  Recall  Precision   F1
0      0.86    1.00       0.86 0.92
InΒ [Β ]:
gb_tuned_test_perf = model_performance_classification_sklearn(gb_estimator_tuned, X_test, y_test)
print("Testing performance \n",gb_tuned_test_perf)
Testing performance 
    Accuracy  Recall  Precision   F1
0      0.86    1.00       0.86 0.92

Tuning of the AdaBoost ModelΒΆ

InΒ [Β ]:
# defining model
ab = AdaBoostClassifier(random_state=1)

# Parameter grid to pass in RandomSearchCV
param_grid = {
    "n_estimators": np.arange(50,110,25),
    "learning_rate": [0.01,0.1,0.05],
    "base_estimator": [
        DecisionTreeClassifier(max_depth=2, random_state=1),
        DecisionTreeClassifier(max_depth=3, random_state=1),
    ],
}

# Type of scoring used to compare parameter combinations
scorer = metrics.make_scorer(metrics.recall_score)

#Calling RandomizedSearchCV
randomized_cv = RandomizedSearchCV(estimator=ab, param_distributions=param_grid, n_iter=10, n_jobs = -1, scoring=scorer, cv=5, random_state=1)

#Fitting parameters in RandomizedSearchCV
randomized_cv.fit(X_train,y_train)

print("Best parameters are {} with CV score={}:" .format(randomized_cv.best_params_,randomized_cv.best_score_))
Best parameters are {'n_estimators': 75, 'learning_rate': 0.1, 'base_estimator': DecisionTreeClassifier(max_depth=3, random_state=1)} with CV score=0.9889060081559957:
InΒ [Β ]:
# Set the model to the best combination of parameters
ab_tuned = AdaBoostClassifier(
    random_state=1,
    n_estimators=75,
    learning_rate=0.1,
    base_estimator=DecisionTreeClassifier(max_depth=3, random_state=1)
)

# Fit the best algorithm to the data
ab_tuned.fit(X_train, y_train)
Out[Β ]:
AdaBoostClassifier(base_estimator=DecisionTreeClassifier(max_depth=3,
                                                         random_state=1),
                   learning_rate=0.1, n_estimators=75, random_state=1)
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
AdaBoostClassifier(base_estimator=DecisionTreeClassifier(max_depth=3,
                                                         random_state=1),
                   learning_rate=0.1, n_estimators=75, random_state=1)
DecisionTreeClassifier(max_depth=3, random_state=1)
DecisionTreeClassifier(max_depth=3, random_state=1)
InΒ [Β ]:
ab_tuned_train_perf = model_performance_classification_sklearn(ab_tuned, X_train, y_train)
print("Training performance \n",ab_tuned_train_perf)
Training performance 
    Accuracy  Recall  Precision   F1
0      0.98    0.99       0.98 0.99
InΒ [Β ]:
ab_tuned_test_perf = model_performance_classification_sklearn(ab_tuned, X_test, y_test)
print("Testing performance \n",ab_tuned_test_perf)
Testing performance 
    Accuracy  Recall  Precision   F1
0      0.97    0.99       0.97 0.98

Final Model Comparison of Tuned ModelsΒΆ

InΒ [Β ]:
# training performance comparison

models_train_comp_df = pd.concat(
    [rf_train_perf.T, rf_tuned_train_perf.T,
     ab_regressor_train_perf.T, ab_tuned_train_perf.T,
     gb_estimator_train_perf.T, gb_tuned_train_perf.T],
    axis=1,
)

models_train_comp_df.columns = [
    "Random Forest",
    "Random Forest Tuned",
    "AdaBoost",
    "AdaBoost Tuned",
    "GradientBoost",
    "GradientBoost Tuned"
]

# Sort columns based on the 'Recall' row
#models_train_comp_df = models_train_comp_df.T.sort_values(by='Recall', ascending=False).T


print("Training performance comparison:")
models_train_comp_df
Training performance comparison:
Out[Β ]:
Random Forest Random Forest Tuned AdaBoost AdaBoost Tuned GradientBoost GradientBoost Tuned
Accuracy 1.00 0.98 0.96 0.98 0.98 0.86
Recall 1.00 0.99 0.98 0.99 0.99 1.00
Precision 1.00 0.98 0.97 0.98 0.98 0.86
F1 1.00 0.99 0.98 0.99 0.99 0.92
InΒ [Β ]:
# testing performance comparison

models_test_comp_df = pd.concat(
    [rf_test_perf.T, rf_tuned_test_perf.T,
     ab_regressor_test_perf.T, ab_tuned_test_perf.T,
     gb_estimator_test_perf.T, gb_tuned_test_perf.T],
    axis=1,
)

models_test_comp_df.columns = [
    "Random Forest",
    "Random Forest Tuned",
    "AdaBoost",
    "AdaBoost Tuned",
    "GradientBoost",
    "GradientBoost Tuned"
]

# Sort columns based on the 'Recall' row
#models_test_comp_df = models_train_comp_df.T.sort_values(by='Recall', ascending=False).T


print("Testing performance comparison:")
models_test_comp_df
Testing performance comparison:
Out[Β ]:
Random Forest Random Forest Tuned AdaBoost AdaBoost Tuned GradientBoost GradientBoost Tuned
Accuracy 0.95 0.94 0.96 0.97 0.96 0.86
Recall 0.99 0.99 0.98 0.99 0.99 1.00
Precision 0.96 0.95 0.97 0.97 0.97 0.86
F1 0.97 0.97 0.98 0.98 0.98 0.92

GradientBoost Tuned appears to give the best performance with a perfect Recall score of 1 on both Test and Train data stamples. Let's look at the confusion matrices for this model along with the feature importances.

InΒ [Β ]:
#Train Confusion Matrix
plot_confusion_matrix(gb_estimator_tuned, X_train, y_train)
No description has been provided for this image
InΒ [Β ]:
confusion_matrix(y_train, gb_estimator_tuned.predict(X_train))
Out[Β ]:
array([[ 164,  975],
       [   1, 5948]])
InΒ [Β ]:
confusion_matrix(y_test, gb_estimator_tuned.predict(X_test))
Out[Β ]:
array([[  61,  427],
       [   0, 2551]])
InΒ [Β ]:
#Test Confusion Matrix
plot_confusion_matrix(gb_estimator_tuned, X_test, y_test)
No description has been provided for this image
InΒ [Β ]:
feature_names = X_train.columns
importances = gb_estimator_tuned.feature_importances_
indices = np.argsort(importances)
plt.figure(figsize=(12,12))
plt.title('Feature Importances')
plt.barh(range(len(indices)), importances[indices], color='violet', align='center')
plt.yticks(range(len(indices)), [feature_names[i] for i in indices])
plt.xlabel('Relative Importance')
plt.show()
No description has been provided for this image

Business Insights and ConclusionsΒΆ

The dominant factors for predicting if a customer will drop their credit cards are:

  • Total_Trans_Amt: Total Transaction Amount (Last 12 months)
  • Total_Revolving_Bal: The balance that carries over from one month to the next
  • Total_Ct_Chng_Q4_Q1: Ratio of the total transaction count in 4th quarter and the total transaction count in the 1st quarter
  • Avg_Utilization_Ratio: Represents how much of the available credit the customer spent

Customers with lower Total_Trans_Amt are more likely to drop. Customers with a lower Total_Revolving_Bal are more likely to drop. Customers with a lower Total_Ct_Chng_Q4_Q1 are more likely to drop. Customers with a low Avg_Utilization_Ratio are more likelty to drop.