Credit Card Users Churn PredictionΒΆ
Problem StatementΒΆ
Business ContextΒΆ
The Thera bank recently saw a steep decline in the number of users of their credit card, credit cards are a good source of income for banks because of different kinds of fees charged by the banks like annual fees, balance transfer fees, and cash advance fees, late payment fees, foreign transaction fees, and others. Some fees are charged to every user irrespective of usage, while others are charged under specified circumstances.
Customersβ leaving credit cards services would lead bank to loss, so the bank wants to analyze the data of customers and identify the customers who will leave their credit card services and reason for same β so that bank could improve upon those areas
You as a Data scientist at Thera bank need to come up with a classification model that will help the bank improve its services so that customers do not renounce their credit cards
Data DescriptionΒΆ
- CLIENTNUM: Client number. Unique identifier for the customer holding the account
- Attrition_Flag: Internal event (customer activity) variable - if the account is closed then "Attrited Customer" else "Existing Customer"
- Customer_Age: Age in Years
- Gender: Gender of the account holder
- Dependent_count: Number of dependents
- Education_Level: Educational Qualification of the account holder - Graduate, High School, Unknown, Uneducated, College(refers to college student), Post-Graduate, Doctorate
- Marital_Status: Marital Status of the account holder
- Income_Category: Annual Income Category of the account holder
- Card_Category: Type of Card
- Months_on_book: Period of relationship with the bank (in months)
- Total_Relationship_Count: Total no. of products held by the customer
- Months_Inactive_12_mon: No. of months inactive in the last 12 months
- Contacts_Count_12_mon: No. of Contacts in the last 12 months
- Credit_Limit: Credit Limit on the Credit Card
- Total_Revolving_Bal: Total Revolving Balance on the Credit Card
- Avg_Open_To_Buy: Open to Buy Credit Line (Average of last 12 months)
- Total_Amt_Chng_Q4_Q1: Change in Transaction Amount (Q4 over Q1)
- Total_Trans_Amt: Total Transaction Amount (Last 12 months)
- Total_Trans_Ct: Total Transaction Count (Last 12 months)
- Total_Ct_Chng_Q4_Q1: Change in Transaction Count (Q4 over Q1)
- Avg_Utilization_Ratio: Average Card Utilization Ratio
What Is a Revolving Balance?ΒΆ
- If we don't pay the balance of the revolving credit account in full every month, the unpaid portion carries over to the next month. That's called a revolving balance
What is the Average Open to buy?ΒΆ
- 'Open to Buy' means the amount left on your credit card to use. Now, this column represents the average of this value for the last 12 months.
What is the Average utilization Ratio?ΒΆ
- The Avg_Utilization_Ratio represents how much of the available credit the customer spent. This is useful for calculating credit scores.
Relation b/w Avg_Open_To_Buy, Credit_Limit and Avg_Utilization_Ratio:ΒΆ
- ( Avg_Open_To_Buy / Credit_Limit ) + Avg_Utilization_Ratio = 1
Please read the instructions carefully before starting the project.ΒΆ
This is a commented Jupyter IPython Notebook file in which all the instructions and tasks to be performed are mentioned.
- Blanks '_______' are provided in the notebook that
needs to be filled with an appropriate code to get the correct result. With every '_______' blank, there is a comment that briefly describes what needs to be filled in the blank space.
- Identify the task to be performed correctly, and only then proceed to write the required code.
- Fill the code wherever asked by the commented lines like "# write your code here" or "# complete the code". Running incomplete code may throw error.
- Please run the codes in a sequential manner from the beginning to avoid any unnecessary errors.
- Add the results/observations (wherever mentioned) derived from the analysis in the presentation and submit the same.
Importing necessary librariesΒΆ
# Installing the libraries with the specified version.
# uncomment and run the following line if Google Colab is being used
# !pip install scikit-learn==1.2.2 seaborn==0.13.1 matplotlib==3.7.1 numpy==1.25.2 pandas==1.5.3 imbalanced-learn==0.10.1 xgboost==2.0.3 -q --user
# Installing the libraries with the specified version.
# uncomment and run the following lines if Jupyter Notebook is being used
# !pip install scikit-learn==1.2.2 seaborn==0.13.1 matplotlib==3.7.1 numpy==1.25.2 pandas==1.5.3 imblearn==0.12.0 xgboost==2.0.3 -q --user
# !pip install --upgrade -q threadpoolctl
Note: After running the above cell, kindly restart the notebook kernel and run all cells sequentially from the start again.
import pandas as pd
import numpy as np
from sklearn import metrics
import matplotlib.pyplot as plt
%matplotlib inline
import warnings
warnings.filterwarnings('ignore')
import seaborn as sns
from sklearn.model_selection import train_test_split
from sklearn.model_selection import GridSearchCV
from sklearn import metrics
from sklearn.ensemble import BaggingClassifier, RandomForestClassifier, AdaBoostClassifier, GradientBoostingClassifier
from sklearn.tree import DecisionTreeClassifier
from sklearn.linear_model import LogisticRegression
from xgboost import XGBClassifier
from sklearn.metrics import accuracy_score, recall_score, precision_score, f1_score, confusion_matrix
from imblearn.over_sampling import SMOTE
from imblearn.under_sampling import RandomUnderSampler
# To tune a model
from sklearn.model_selection import GridSearchCV
from sklearn.model_selection import RandomizedSearchCV
# To suppress warnings
import warnings
warnings.filterwarnings("ignore")
# Set display options to avoid scientific notation
pd.set_option('display.float_format', '{:.2f}'.format)
# Set the option to display all columns
pd.set_option('display.max_columns', None)
Loading the datasetΒΆ
# uncomment and run the following line if using Google Colab
from google.colab import drive
drive.mount('/content/drive')
Mounted at /content/drive
bankData = pd.read_csv("/content/drive/MyDrive/Personal/UT Austin/Project 3 - AML/BankChurners.csv")
data = bankData.copy()
CLIENTNUM won't be used in the analysis, so can remove it.
#Remove CLIENTNUM column from data
data.drop(['CLIENTNUM'], axis=1, inplace=True)
Data OverviewΒΆ
- Observations
- Sanity checks
Viewing the first and last 5 rows of the datasetΒΆ
data.head(5)
| CLIENTNUM | Attrition_Flag | Customer_Age | Gender | Dependent_count | Education_Level | Marital_Status | Income_Category | Card_Category | Months_on_book | ... | Months_Inactive_12_mon | Contacts_Count_12_mon | Credit_Limit | Total_Revolving_Bal | Avg_Open_To_Buy | Total_Amt_Chng_Q4_Q1 | Total_Trans_Amt | Total_Trans_Ct | Total_Ct_Chng_Q4_Q1 | Avg_Utilization_Ratio | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 0 | 768805383 | Existing Customer | 45 | M | 3 | High School | Married | $60K - $80K | Blue | 39 | ... | 1 | 3 | 12691.00 | 777 | 11914.00 | 1.33 | 1144 | 42 | 1.62 | 0.06 |
| 1 | 818770008 | Existing Customer | 49 | F | 5 | Graduate | Single | Less than $40K | Blue | 44 | ... | 1 | 2 | 8256.00 | 864 | 7392.00 | 1.54 | 1291 | 33 | 3.71 | 0.10 |
| 2 | 713982108 | Existing Customer | 51 | M | 3 | Graduate | Married | $80K - $120K | Blue | 36 | ... | 1 | 0 | 3418.00 | 0 | 3418.00 | 2.59 | 1887 | 20 | 2.33 | 0.00 |
| 3 | 769911858 | Existing Customer | 40 | F | 4 | High School | NaN | Less than $40K | Blue | 34 | ... | 4 | 1 | 3313.00 | 2517 | 796.00 | 1.41 | 1171 | 20 | 2.33 | 0.76 |
| 4 | 709106358 | Existing Customer | 40 | M | 3 | Uneducated | Married | $60K - $80K | Blue | 21 | ... | 1 | 0 | 4716.00 | 0 | 4716.00 | 2.17 | 816 | 28 | 2.50 | 0.00 |
5 rows Γ 21 columns
data.tail(5)
| CLIENTNUM | Attrition_Flag | Customer_Age | Gender | Dependent_count | Education_Level | Marital_Status | Income_Category | Card_Category | Months_on_book | ... | Months_Inactive_12_mon | Contacts_Count_12_mon | Credit_Limit | Total_Revolving_Bal | Avg_Open_To_Buy | Total_Amt_Chng_Q4_Q1 | Total_Trans_Amt | Total_Trans_Ct | Total_Ct_Chng_Q4_Q1 | Avg_Utilization_Ratio | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 10122 | 772366833 | Existing Customer | 50 | M | 2 | Graduate | Single | $40K - $60K | Blue | 40 | ... | 2 | 3 | 4003.00 | 1851 | 2152.00 | 0.70 | 15476 | 117 | 0.86 | 0.46 |
| 10123 | 710638233 | Attrited Customer | 41 | M | 2 | NaN | Divorced | $40K - $60K | Blue | 25 | ... | 2 | 3 | 4277.00 | 2186 | 2091.00 | 0.80 | 8764 | 69 | 0.68 | 0.51 |
| 10124 | 716506083 | Attrited Customer | 44 | F | 1 | High School | Married | Less than $40K | Blue | 36 | ... | 3 | 4 | 5409.00 | 0 | 5409.00 | 0.82 | 10291 | 60 | 0.82 | 0.00 |
| 10125 | 717406983 | Attrited Customer | 30 | M | 2 | Graduate | NaN | $40K - $60K | Blue | 36 | ... | 3 | 3 | 5281.00 | 0 | 5281.00 | 0.54 | 8395 | 62 | 0.72 | 0.00 |
| 10126 | 714337233 | Attrited Customer | 43 | F | 2 | Graduate | Married | Less than $40K | Silver | 25 | ... | 2 | 4 | 10388.00 | 1961 | 8427.00 | 0.70 | 10294 | 61 | 0.65 | 0.19 |
5 rows Γ 21 columns
Checking the shape of the dataset, attribute types, and statistical summaryΒΆ
data.shape
(10127, 21)
data.info()
<class 'pandas.core.frame.DataFrame'> RangeIndex: 10127 entries, 0 to 10126 Data columns (total 21 columns): # Column Non-Null Count Dtype --- ------ -------------- ----- 0 CLIENTNUM 10127 non-null int64 1 Attrition_Flag 10127 non-null object 2 Customer_Age 10127 non-null int64 3 Gender 10127 non-null object 4 Dependent_count 10127 non-null int64 5 Education_Level 8608 non-null object 6 Marital_Status 9378 non-null object 7 Income_Category 10127 non-null object 8 Card_Category 10127 non-null object 9 Months_on_book 10127 non-null int64 10 Total_Relationship_Count 10127 non-null int64 11 Months_Inactive_12_mon 10127 non-null int64 12 Contacts_Count_12_mon 10127 non-null int64 13 Credit_Limit 10127 non-null float64 14 Total_Revolving_Bal 10127 non-null int64 15 Avg_Open_To_Buy 10127 non-null float64 16 Total_Amt_Chng_Q4_Q1 10127 non-null float64 17 Total_Trans_Amt 10127 non-null int64 18 Total_Trans_Ct 10127 non-null int64 19 Total_Ct_Chng_Q4_Q1 10127 non-null float64 20 Avg_Utilization_Ratio 10127 non-null float64 dtypes: float64(5), int64(10), object(6) memory usage: 1.6+ MB
#Get Summary Statistics for the numerical columns
data.describe().T
| count | mean | std | min | 25% | 50% | 75% | max | |
|---|---|---|---|---|---|---|---|---|
| CLIENTNUM | 10127.00 | 739177606.33 | 36903783.45 | 708082083.00 | 713036770.50 | 717926358.00 | 773143533.00 | 828343083.00 |
| Customer_Age | 10127.00 | 46.33 | 8.02 | 26.00 | 41.00 | 46.00 | 52.00 | 73.00 |
| Dependent_count | 10127.00 | 2.35 | 1.30 | 0.00 | 1.00 | 2.00 | 3.00 | 5.00 |
| Months_on_book | 10127.00 | 35.93 | 7.99 | 13.00 | 31.00 | 36.00 | 40.00 | 56.00 |
| Total_Relationship_Count | 10127.00 | 3.81 | 1.55 | 1.00 | 3.00 | 4.00 | 5.00 | 6.00 |
| Months_Inactive_12_mon | 10127.00 | 2.34 | 1.01 | 0.00 | 2.00 | 2.00 | 3.00 | 6.00 |
| Contacts_Count_12_mon | 10127.00 | 2.46 | 1.11 | 0.00 | 2.00 | 2.00 | 3.00 | 6.00 |
| Credit_Limit | 10127.00 | 8631.95 | 9088.78 | 1438.30 | 2555.00 | 4549.00 | 11067.50 | 34516.00 |
| Total_Revolving_Bal | 10127.00 | 1162.81 | 814.99 | 0.00 | 359.00 | 1276.00 | 1784.00 | 2517.00 |
| Avg_Open_To_Buy | 10127.00 | 7469.14 | 9090.69 | 3.00 | 1324.50 | 3474.00 | 9859.00 | 34516.00 |
| Total_Amt_Chng_Q4_Q1 | 10127.00 | 0.76 | 0.22 | 0.00 | 0.63 | 0.74 | 0.86 | 3.40 |
| Total_Trans_Amt | 10127.00 | 4404.09 | 3397.13 | 510.00 | 2155.50 | 3899.00 | 4741.00 | 18484.00 |
| Total_Trans_Ct | 10127.00 | 64.86 | 23.47 | 10.00 | 45.00 | 67.00 | 81.00 | 139.00 |
| Total_Ct_Chng_Q4_Q1 | 10127.00 | 0.71 | 0.24 | 0.00 | 0.58 | 0.70 | 0.82 | 3.71 |
| Avg_Utilization_Ratio | 10127.00 | 0.27 | 0.28 | 0.00 | 0.02 | 0.18 | 0.50 | 1.00 |
Checking for missing values or duplicate valuesΒΆ
# checking for null values
data.isnull().sum()
| 0 | |
|---|---|
| CLIENTNUM | 0 |
| Attrition_Flag | 0 |
| Customer_Age | 0 |
| Gender | 0 |
| Dependent_count | 0 |
| Education_Level | 1519 |
| Marital_Status | 749 |
| Income_Category | 0 |
| Card_Category | 0 |
| Months_on_book | 0 |
| Total_Relationship_Count | 0 |
| Months_Inactive_12_mon | 0 |
| Contacts_Count_12_mon | 0 |
| Credit_Limit | 0 |
| Total_Revolving_Bal | 0 |
| Avg_Open_To_Buy | 0 |
| Total_Amt_Chng_Q4_Q1 | 0 |
| Total_Trans_Amt | 0 |
| Total_Trans_Ct | 0 |
| Total_Ct_Chng_Q4_Q1 | 0 |
| Avg_Utilization_Ratio | 0 |
- Education_Level (N/A) has ~15% missing data.
- Martial_Status (NaN) has ~7.5% missing data
- Although not evident in the above Income_Category has ~11% missing values that show up as 'abc'.
#Checking for duplicate values
data.duplicated().sum()
0
There are no duplicate rows in the dataset
Checking the mix of customers who attrited vs remained with the credit card.ΒΆ
# Get counts of each unique value
data['Attrition_Flag'].value_counts()
| count | |
|---|---|
| Attrition_Flag | |
| Existing Customer | 8500 |
| Attrited Customer | 1627 |
# Get percentages of each unique value
data['Attrition_Flag'].value_counts(normalize=True) * 100
| proportion | |
|---|---|
| Attrition_Flag | |
| Existing Customer | 83.93 |
| Attrited Customer | 16.07 |
Attrited Customers make up 16% of the total data, so there is definitely an imbalance between Existing and Attrited the model will need to consider.
Exploratory Data Analysis (EDA)ΒΆ
- EDA is an important part of any project involving data.
- It is important to investigate and understand the data better before building a model with it.
- A few questions have been mentioned below which will help you approach the analysis in the right manner and generate insights from the data.
- A thorough analysis of the data, in addition to the questions mentioned below, should be done.
Questions:
- How is the total transaction amount distributed?
- What is the distribution of the level of education of customers?
- What is the distribution of the level of income of customers?
- How does the change in transaction amount between Q4 and Q1 (
total_ct_change_Q4_Q1) vary by the customer's account status (Attrition_Flag)? - How does the number of months a customer was inactive in the last 12 months (
Months_Inactive_12_mon) vary by the customer's account status (Attrition_Flag)? - What are the attributes that have a strong correlation with each other?
The below functions need to be defined to carry out the Exploratory Data Analysis.ΒΆ
# function to plot a boxplot and a histogram along the same scale.
def histogram_boxplot(data, feature, figsize=(12, 7), kde=False, bins=None):
"""
Boxplot and histogram combined
data: dataframe
feature: dataframe column
figsize: size of figure (default (12,7))
kde: whether to the show density curve (default False)
bins: number of bins for histogram (default None)
"""
f2, (ax_box2, ax_hist2) = plt.subplots(
nrows=2, # Number of rows of the subplot grid= 2
sharex=True, # x-axis will be shared among all subplots
gridspec_kw={"height_ratios": (0.25, 0.75)},
figsize=figsize,
) # creating the 2 subplots
sns.boxplot(
data=data, x=feature, ax=ax_box2, showmeans=True, color="violet"
) # boxplot will be created and a triangle will indicate the mean value of the column
sns.histplot(
data=data, x=feature, kde=kde, ax=ax_hist2, bins=bins, palette="winter"
) if bins else sns.histplot(
data=data, x=feature, kde=kde, ax=ax_hist2
) # For histogram
ax_hist2.axvline(
data[feature].mean(), color="green", linestyle="--"
) # Add mean to the histogram
ax_hist2.axvline(
data[feature].median(), color="black", linestyle="-"
) # Add median to the histogram
# function to create labeled barplots
def labeled_barplot(data, feature, perc=False, n=None):
"""
Barplot with percentage at the top
data: dataframe
feature: dataframe column
perc: whether to display percentages instead of count (default is False)
n: displays the top n category levels (default is None, i.e., display all levels)
"""
total = len(data[feature]) # length of the column
count = data[feature].nunique()
if n is None:
plt.figure(figsize=(count + 1, 5))
else:
plt.figure(figsize=(n + 1, 5))
plt.xticks(rotation=90, fontsize=15)
ax = sns.countplot(
data=data,
x=feature,
palette="Paired",
order=data[feature].value_counts().index[:n].sort_values(),
)
for p in ax.patches:
if perc == True:
label = "{:.1f}%".format(
100 * p.get_height() / total
) # percentage of each class of the category
else:
label = p.get_height() # count of each level of the category
x = p.get_x() + p.get_width() / 2 # width of the plot
y = p.get_height() # height of the plot
ax.annotate(
label,
(x, y),
ha="center",
va="center",
size=12,
xytext=(0, 5),
textcoords="offset points",
) # annotate the percentage
plt.show() # show the plot
# function to plot stacked bar chart
def stacked_barplot(data, predictor, target):
"""
Print the category counts and plot a stacked bar chart
data: dataframe
predictor: independent variable
target: target variable
"""
count = data[predictor].nunique()
sorter = data[target].value_counts().index[-1]
tab1 = pd.crosstab(data[predictor], data[target], margins=True).sort_values(
by=sorter, ascending=False
)
print(tab1)
print("-" * 120)
tab = pd.crosstab(data[predictor], data[target], normalize="index").sort_values(
by=sorter, ascending=False
)
tab.plot(kind="bar", stacked=True, figsize=(count + 1, 5))
plt.legend(
loc="lower left", frameon=False,
)
plt.legend(loc="upper left", bbox_to_anchor=(1, 1))
plt.show()
### Function to plot distributions
def distribution_plot_wrt_target(data, predictor, target):
fig, axs = plt.subplots(2, 2, figsize=(12, 10))
target_uniq = data[target].unique()
axs[0, 0].set_title("Distribution of target for target=" + str(target_uniq[0]))
sns.histplot(
data=data[data[target] == target_uniq[0]],
x=predictor,
kde=True,
ax=axs[0, 0],
color="teal",
)
axs[0, 1].set_title("Distribution of target for target=" + str(target_uniq[1]))
sns.histplot(
data=data[data[target] == target_uniq[1]],
x=predictor,
kde=True,
ax=axs[0, 1],
color="orange",
)
axs[1, 0].set_title("Boxplot w.r.t target")
sns.boxplot(data=data, x=target, y=predictor, ax=axs[1, 0], palette="gist_rainbow")
axs[1, 1].set_title("Boxplot (without outliers) w.r.t target")
sns.boxplot(
data=data,
x=target,
y=predictor,
ax=axs[1, 1],
showfliers=False,
palette="gist_rainbow",
)
plt.tight_layout()
plt.show()
Univariate AnalysisΒΆ
Numerical VariablesΒΆ
histogram_boxplot(data, 'Customer_Age')
histogram_boxplot(data, 'Dependent_count')
histogram_boxplot(data, 'Months_on_book')
histogram_boxplot(data, 'Total_Relationship_Count')
histogram_boxplot(data, 'Months_Inactive_12_mon')
histogram_boxplot(data, 'Contacts_Count_12_mon')
histogram_boxplot(data, 'Credit_Limit')
histogram_boxplot(data, 'Total_Revolving_Bal')
histogram_boxplot(data, 'Avg_Open_To_Buy')
histogram_boxplot(data, 'Total_Amt_Chng_Q4_Q1')
histogram_boxplot(data, 'Total_Trans_Amt')
histogram_boxplot(data, 'Total_Trans_Ct')
histogram_boxplot(data, 'Total_Ct_Chng_Q4_Q1')
histogram_boxplot(data, 'Avg_Utilization_Ratio')
Categorical VariablesΒΆ
labeled_barplot(data, 'Gender')
labeled_barplot(data, 'Education_Level')
labeled_barplot(data, 'Marital_Status')
labeled_barplot(data, 'Income_Category')
labeled_barplot(data, 'Card_Category')
Bivariate AnalysisΒΆ
Numerical VariablesΒΆ
Let's look at the correlation matrix for all the numeric variables.
# plotting the correlation heatmap
plt.figure(figsize=(12, 7))
sns.heatmap(data.corr(numeric_only = True), annot=True, fmt='0.2f', cmap='Spectral')
plt.show()
Some notable correlations:
- Months_on_book is correlated with the Customer_Age which makes sense given the older someone is the better chance they have to have a longer relationship with the bank.
- Total_Trans_Ct is correlated with Total_Trans_Amt which makes sense in that the more transactions you have the more the $ amount will be.
- Avg_Utilization_Ratio is correlated with Total_Revolving_Balance which makes sense in that when you aren't paying your balance in full each month, you will be utilizing more of your available credit.
sns.pairplot(data,hue='Attrition_Flag')
plt.show()
Output hidden; open in https://colab.research.google.com to view.
distribution_plot_wrt_target(data, 'Customer_Age', 'Attrition_Flag')
distribution_plot_wrt_target(data, 'Dependent_count', 'Attrition_Flag')
distribution_plot_wrt_target(data, 'Months_on_book', 'Attrition_Flag')
distribution_plot_wrt_target(data, 'Total_Relationship_Count', 'Attrition_Flag')
distribution_plot_wrt_target(data, 'Months_Inactive_12_mon', 'Attrition_Flag')
distribution_plot_wrt_target(data, 'Contacts_Count_12_mon', 'Attrition_Flag')
distribution_plot_wrt_target(data, 'Credit_Limit', 'Attrition_Flag')
distribution_plot_wrt_target(data, 'Total_Revolving_Bal', 'Attrition_Flag')
distribution_plot_wrt_target(data, 'Avg_Open_To_Buy', 'Attrition_Flag')
distribution_plot_wrt_target(data, 'Total_Amt_Chng_Q4_Q1', 'Attrition_Flag')
distribution_plot_wrt_target(data, 'Total_Trans_Amt', 'Attrition_Flag')
distribution_plot_wrt_target(data, 'Total_Trans_Ct', 'Attrition_Flag')
distribution_plot_wrt_target(data, 'Total_Ct_Chng_Q4_Q1', 'Attrition_Flag')
distribution_plot_wrt_target(data, 'Avg_Utilization_Ratio', 'Attrition_Flag')
Categorical VariablesΒΆ
stacked_barplot(data, 'Gender', 'Attrition_Flag')
Attrition_Flag Attrited Customer Existing Customer All Gender All 1627 8500 10127 F 930 4428 5358 M 697 4072 4769 ------------------------------------------------------------------------------------------------------------------------
# Even though Dependent_count is technically a numerical variables, there are so few values, it can be instructure to view as a categorical variable.
stacked_barplot(data, 'Dependent_count', 'Attrition_Flag')
Attrition_Flag Attrited Customer Existing Customer All Dependent_count All 1627 8500 10127 3 482 2250 2732 2 417 2238 2655 1 269 1569 1838 4 260 1314 1574 0 135 769 904 5 64 360 424 ------------------------------------------------------------------------------------------------------------------------
stacked_barplot(data, 'Education_Level', 'Attrition_Flag')
Attrition_Flag Attrited Customer Existing Customer All Education_Level All 1371 7237 8608 Graduate 487 2641 3128 High School 306 1707 2013 Uneducated 237 1250 1487 College 154 859 1013 Doctorate 95 356 451 Post-Graduate 92 424 516 ------------------------------------------------------------------------------------------------------------------------
stacked_barplot(data, 'Marital_Status', 'Attrition_Flag')
Attrition_Flag Attrited Customer Existing Customer All Marital_Status All 1498 7880 9378 Married 709 3978 4687 Single 668 3275 3943 Divorced 121 627 748 ------------------------------------------------------------------------------------------------------------------------
stacked_barplot(data, 'Income_Category', 'Attrition_Flag')
Attrition_Flag Attrited Customer Existing Customer All Income_Category All 1627 8500 10127 Less than $40K 612 2949 3561 $40K - $60K 271 1519 1790 $80K - $120K 242 1293 1535 $60K - $80K 189 1213 1402 abc 187 925 1112 $120K + 126 601 727 ------------------------------------------------------------------------------------------------------------------------
stacked_barplot(data, 'Card_Category', 'Attrition_Flag')
Attrition_Flag Attrited Customer Existing Customer All Card_Category All 1627 8500 10127 Blue 1519 7917 9436 Silver 82 473 555 Gold 21 95 116 Platinum 5 15 20 ------------------------------------------------------------------------------------------------------------------------
Data Pre-processingΒΆ
Missing value imputationΒΆ
As noted earlier, Education_Level and Martial_Status, have a small percentage of missing values. Income_Category also has missing values which appear as 'abc'. There are missing Education_Level values which show up as 'N/A'.
Based on our stacked bar plot of Education_Level and Attrition_Flag, we saw that Uneducated, High School, College, and Graduate all had about the same prevalence of Attrited Customer, so let's set the missing values here to the mode which is Graduate.
# Set NaN values of Education_Level to 'Graduate'
data['Education_Level'] = data['Education_Level'].fillna('Graduate')
print(data.Education_Level.value_counts(dropna=False))
Education_Level Graduate 4647 High School 2013 Uneducated 1487 College 1013 Post-Graduate 516 Doctorate 451 Name: count, dtype: int64
Based on our stacked bar plot of Marital_status and Attrition_Flag, we saw that the three values all had about the same prevalence of Attrited Customer, so let's set the missing values here to the mode which is Married.
# Set NaN values of Education_Level to 'Graduate'
data['Marital_Status'] = data['Marital_Status'].fillna('Married')
print(data.Marital_Status.value_counts(dropna=False))
Marital_Status Married 5436 Single 3943 Divorced 748 Name: count, dtype: int64
To decide what to do with the abc values in Income_Category, we could reasonably assume some correlation between Income_Category and Total_Trans_Amt and Customer_Age. In the below violin plots we can see that "Less than $40K" is the closest match for Total_Trans_Amt, while for Customer_Age, "Less than $40K" is also a good match. This value also happens to be the mode, so let's go with these value for the missing values.
plt.figure(figsize=(14, 8))
sns.violinplot(data=data, x='Income_Category', y='Total_Trans_Amt', orient='v');
plt.figure(figsize=(14, 8))
sns.violinplot(data=data, x='Income_Category', y='Customer_Age', orient='v');
# Replace 'abc' with Less than $40K
data['Income_Category'].replace('abc', 'Less than $40K', inplace=True)
Conversion from Object to CategoryΒΆ
Lets convert the columns with an 'object' datatype into categorical variables, to reduce the space required to store the dataframe and improve processing efficieny.
for feature in data.columns: # Loop through all columns in the dataframe
if data[feature].dtype == 'object': # Only apply for columns with categorical strings
data[feature] = pd.Categorical(data[feature]) # Convert the column to a categorical type
# Confirm that column of type object have been converted to category
data.info()
<class 'pandas.core.frame.DataFrame'> RangeIndex: 10127 entries, 0 to 10126 Data columns (total 20 columns): # Column Non-Null Count Dtype --- ------ -------------- ----- 0 Attrition_Flag 10127 non-null category 1 Customer_Age 10127 non-null int64 2 Gender 10127 non-null category 3 Dependent_count 10127 non-null int64 4 Education_Level 10127 non-null category 5 Marital_Status 10127 non-null category 6 Income_Category 10127 non-null category 7 Card_Category 10127 non-null category 8 Months_on_book 10127 non-null int64 9 Total_Relationship_Count 10127 non-null int64 10 Months_Inactive_12_mon 10127 non-null int64 11 Contacts_Count_12_mon 10127 non-null int64 12 Credit_Limit 10127 non-null float64 13 Total_Revolving_Bal 10127 non-null int64 14 Avg_Open_To_Buy 10127 non-null float64 15 Total_Amt_Chng_Q4_Q1 10127 non-null float64 16 Total_Trans_Amt 10127 non-null int64 17 Total_Trans_Ct 10127 non-null int64 18 Total_Ct_Chng_Q4_Q1 10127 non-null float64 19 Avg_Utilization_Ratio 10127 non-null float64 dtypes: category(6), float64(5), int64(9) memory usage: 1.1 MB
print(data.Attrition_Flag.value_counts(dropna=False))
print()
print(data.Gender.value_counts(dropna=False))
print()
print(data.Education_Level.value_counts(dropna=False))
print()
print(data.Marital_Status.value_counts(dropna=False))
print()
print(data.Income_Category.value_counts(dropna=False))
print()
print(data.Card_Category.value_counts(dropna=False))
Attrition_Flag Existing Customer 8500 Attrited Customer 1627 Name: count, dtype: int64 Gender F 5358 M 4769 Name: count, dtype: int64 Education_Level Graduate 4647 High School 2013 Uneducated 1487 College 1013 Post-Graduate 516 Doctorate 451 Name: count, dtype: int64 Marital_Status Married 5436 Single 3943 Divorced 748 Name: count, dtype: int64 Income_Category Less than $40K 4673 $40K - $60K 1790 $80K - $120K 1535 $60K - $80K 1402 $120K + 727 Name: count, dtype: int64 Card_Category Blue 9436 Silver 555 Gold 116 Platinum 20 Name: count, dtype: int64
print(data.Income_Category.value_counts(dropna=False))
Income_Category Less than $40K 4673 $40K - $60K 1790 $80K - $120K 1535 $60K - $80K 1402 $120K + 727 Name: count, dtype: int64
One-Hot EncodingΒΆ
replaceStruct = {
"Attrition_Flag": {"Existing Customer": 1, "Attrited Customer": 2},
"Gender": {"M": 1, "F":2},
"Education_Level": {"Uneducated": 1, "High School":2 , "College": 3, "Graduate": 4, "Post-Graduate": 5,"Doctorate": 6},
"Marital_Status": {"Married": 1, "Single":2 , "Divorced": 3},
"Income_Category": {"Less than $40K": 1, "$40K - $60K": 2 ,"$60K - $80K": 3 ,"$80K - $120K": 4 ,"$120K +": 5},
"Card_Category": {"Blue": 1, "Silver": 2, "Gold": 3, "Platinum": 4 }
}
oneHotCols=["Attrition_Flag","Gender","Education_Level","Marital_Status","Income_Category", "Card_Category"]
data=data.replace(replaceStruct)
data=pd.get_dummies(data, columns=oneHotCols, drop_first=True)
data.head(10)
| Customer_Age | Dependent_count | Months_on_book | Total_Relationship_Count | Months_Inactive_12_mon | Contacts_Count_12_mon | Credit_Limit | Total_Revolving_Bal | Avg_Open_To_Buy | Total_Amt_Chng_Q4_Q1 | Total_Trans_Amt | Total_Trans_Ct | Total_Ct_Chng_Q4_Q1 | Avg_Utilization_Ratio | Attrition_Flag_1 | Gender_1 | Education_Level_6 | Education_Level_4 | Education_Level_2 | Education_Level_5 | Education_Level_1 | Marital_Status_1 | Marital_Status_2 | Income_Category_2 | Income_Category_3 | Income_Category_4 | Income_Category_1 | Income_Category_abc | Card_Category_3 | Card_Category_4 | Card_Category_2 | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 0 | 45 | 3 | 39 | 5 | 1 | 3 | 12691.00 | 777 | 11914.00 | 1.33 | 1144 | 42 | 1.62 | 0.06 | True | True | False | False | True | False | False | True | False | False | True | False | False | False | False | False | False |
| 1 | 49 | 5 | 44 | 6 | 1 | 2 | 8256.00 | 864 | 7392.00 | 1.54 | 1291 | 33 | 3.71 | 0.10 | True | False | False | True | False | False | False | False | True | False | False | False | True | False | False | False | False |
| 2 | 51 | 3 | 36 | 4 | 1 | 0 | 3418.00 | 0 | 3418.00 | 2.59 | 1887 | 20 | 2.33 | 0.00 | True | True | False | True | False | False | False | True | False | False | False | True | False | False | False | False | False |
| 3 | 40 | 4 | 34 | 3 | 4 | 1 | 3313.00 | 2517 | 796.00 | 1.41 | 1171 | 20 | 2.33 | 0.76 | True | False | False | False | True | False | False | True | False | False | False | False | True | False | False | False | False |
| 4 | 40 | 3 | 21 | 5 | 1 | 0 | 4716.00 | 0 | 4716.00 | 2.17 | 816 | 28 | 2.50 | 0.00 | True | True | False | False | False | False | True | True | False | False | True | False | False | False | False | False | False |
| 5 | 44 | 2 | 36 | 3 | 1 | 2 | 4010.00 | 1247 | 2763.00 | 1.38 | 1088 | 24 | 0.85 | 0.31 | True | True | False | True | False | False | False | True | False | True | False | False | False | False | False | False | False |
| 6 | 51 | 4 | 46 | 6 | 1 | 3 | 34516.00 | 2264 | 32252.00 | 1.98 | 1330 | 31 | 0.72 | 0.07 | True | True | False | True | False | False | False | True | False | False | False | False | False | False | True | False | False |
| 7 | 32 | 0 | 27 | 2 | 2 | 2 | 29081.00 | 1396 | 27685.00 | 2.20 | 1538 | 36 | 0.71 | 0.05 | True | True | False | False | True | False | False | True | False | False | True | False | False | False | False | False | True |
| 8 | 37 | 3 | 36 | 5 | 2 | 0 | 22352.00 | 2517 | 19835.00 | 3.35 | 1350 | 24 | 1.18 | 0.11 | True | True | False | False | False | False | True | False | True | False | True | False | False | False | False | False | False |
| 9 | 48 | 2 | 36 | 6 | 3 | 3 | 11656.00 | 1677 | 9979.00 | 1.52 | 1441 | 32 | 0.88 | 0.14 | True | True | False | True | False | False | False | False | True | False | False | True | False | False | False | False | False |
Outlier TreatmentΒΆ
Although some of the variables have outliers as event in the EDA, we will keep these values since they are valid data.
Model BuildingΒΆ
Model evaluation criterionΒΆ
The nature of predictions made by the classification model will translate as follows:
- True positives (TP) are failures correctly predicted by the model.
- False negatives (FN) are real failures in a generator where there is no detection by model.
- False positives (FP) are failure detections in a generator where there is no failure.
Which metric to optimize?
- We need to choose the metric which will ensure that the maximum number of generator failures are predicted correctly by the model.
- We would want Recall to be maximized as greater the Recall, the higher the chances of minimizing false negatives.
- We want to minimize false negatives because if a model predicts that a machine will have no failure when there will be a failure, it will increase the maintenance cost.
Let's define a function to output different metrics (including recall) on the train and test set and a function to show confusion matrix so that we do not have to use the same code repetitively while evaluating models.
# defining a function to compute different metrics to check performance of a classification model built using sklearn
def model_performance_classification_sklearn(model, predictors, target):
"""
Function to compute different metrics to check classification model performance
model: classifier
predictors: independent variables
target: dependent variable
"""
# predicting using the independent variables
pred = model.predict(predictors)
acc = accuracy_score(target, pred) # to compute Accuracy
recall = recall_score(target, pred) # to compute Recall
precision = precision_score(target, pred) # to compute Precision
f1 = f1_score(target, pred) # to compute F1-score
# creating a dataframe of metrics
df_perf = pd.DataFrame(
{
"Accuracy": acc,
"Recall": recall,
"Precision": precision,
"F1": f1
},
index=[0],
)
return df_perf
def plot_confusion_matrix(model, predictors, target):
"""
To plot the confusion_matrix with percentages
model: classifier
predictors: independent variables
target: dependent variable
"""
# Predict the target values using the provided model and predictors
y_pred = model.predict(predictors)
# Compute the confusion matrix comparing the true target values with the predicted values
cm = confusion_matrix(target, y_pred)
# Create labels for each cell in the confusion matrix with both count and percentage
labels = np.asarray(
[
["{0:0.0f}".format(item) + "\n{0:.2%}".format(item / cm.flatten().sum())]
for item in cm.flatten()
]
).reshape(2, 2) # reshaping to a matrix
# Set the figure size for the plot
plt.figure(figsize=(6, 4))
# Plot the confusion matrix as a heatmap with the labels
sns.heatmap(cm, annot=labels, fmt="")
# Add a label to the y-axis
plt.ylabel("True label")
# Add a label to the x-axis
plt.xlabel("Predicted label")
Split the data into train and test setsΒΆ
- Since we have an imbalance in the distribution of the target class (Attrition_Flag), it is good to use stratified sampling to ensure that relative class frequencies are approximately preserved in train and test sets.
- We do this by setting the
stratifyparameter to target variable in the train_test_split function.
X = data.drop("Attrition_Flag_1" , axis=1)
y = data.pop("Attrition_Flag_1")
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=.30, random_state=1,stratify=y)
#Show the proportion of value in y
y_train.value_counts(normalize=False)
| count | |
|---|---|
| Attrition_Flag_1 | |
| True | 5949 |
| False | 1139 |
Note: For this project I elected to go with the simpler approach of using just a Train and Test split. I acknowledge though another approach is to split the data per Train, Validation, and Test, where Test can be reserved for a check on the final best model.
Model Building with Original dataΒΆ
Build 5 models (from decision trees, bagging and boosting methods) - Comment on the model performance * You can choose NOT to build XGBoost if you are facing issues with the installation
BaggingΒΆ
bagging_estimator=BaggingClassifier(random_state=1)
bagging_estimator.fit(X_train,y_train)
BaggingClassifier(random_state=1)In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
BaggingClassifier(random_state=1)
bagging_estimator_train_perf = model_performance_classification_sklearn(bagging_estimator, X_train, y_train)
print("Training performance \n",bagging_estimator_train_perf)
Training performance
Accuracy Recall Precision F1
0 1.00 1.00 1.00 1.00
bagging_estimator_test_perf = model_performance_classification_sklearn(bagging_estimator, X_test, y_test)
print("Testing performance \n",bagging_estimator_test_perf)
Testing performance
Accuracy Recall Precision F1
0 0.95 0.97 0.97 0.97
Random ForestΒΆ
rf = RandomForestClassifier(random_state=1)
rf.fit(X_train,y_train)
RandomForestClassifier(random_state=1)In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
RandomForestClassifier(random_state=1)
rf_train_perf = model_performance_classification_sklearn(rf, X_train, y_train)
print("Training performance \n",rf_train_perf)
Training performance
Accuracy Recall Precision F1
0 1.00 1.00 1.00 1.00
rf_test_perf = model_performance_classification_sklearn(rf, X_test, y_test)
print("Testing performance \n",rf_test_perf)
Testing performance
Accuracy Recall Precision F1
0 0.95 0.99 0.96 0.97
AdaBoostΒΆ
ab_regressor = AdaBoostClassifier(random_state=1)
ab_regressor.fit(X_train,y_train)
AdaBoostClassifier(random_state=1)In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
AdaBoostClassifier(random_state=1)
#Train performance
ab_regressor_train_perf = model_performance_classification_sklearn(ab_regressor, X_train, y_train)
print("Training performance \n",ab_regressor_train_perf)
Training performance
Accuracy Recall Precision F1
0 0.96 0.98 0.97 0.98
#Test performance
ab_regressor_test_perf = model_performance_classification_sklearn(ab_regressor, X_test, y_test)
print("Testing performance \n",ab_regressor_test_perf)
Testing performance
Accuracy Recall Precision F1
0 0.96 0.98 0.97 0.98
GradientBoostingΒΆ
#GradientBoost Classifier
gb_estimator = GradientBoostingClassifier(random_state=1)
gb_estimator.fit(X_train,y_train)
GradientBoostingClassifier(random_state=1)In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
GradientBoostingClassifier(random_state=1)
#Train performance
gb_estimator_train_perf = model_performance_classification_sklearn(gb_estimator, X_train, y_train)
print("Training performance \n",gb_estimator_train_perf)
Training performance
Accuracy Recall Precision F1
0 0.98 0.99 0.98 0.99
#Test performance
gb_estimator_test_perf = model_performance_classification_sklearn(gb_estimator, X_test, y_test)
print("Testing performance \n",gb_estimator_test_perf)
Testing performance
Accuracy Recall Precision F1
0 0.96 0.99 0.97 0.98
Decision TreeΒΆ
dtree=DecisionTreeClassifier(random_state=1)
dtree.fit(X_train,y_train)
DecisionTreeClassifier(random_state=1)In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
DecisionTreeClassifier(random_state=1)
dtree_model_train_perf = model_performance_classification_sklearn(dtree, X_train, y_train)
print("Training performance \n",dtree_model_train_perf)
Training performance
Accuracy Recall Precision F1
0 1.00 1.00 1.00 1.00
dtree_model_test_perf = model_performance_classification_sklearn(dtree, X_test, y_test)
print("Testing performance \n",dtree_model_test_perf)
Testing performance
Accuracy Recall Precision F1
0 0.94 0.96 0.96 0.96
Model Building with Oversampled dataΒΆ
Oversample the train data - Build 5 models (from decision trees, bagging and boosting methods) - Comment on the model performance * You can choose NOT to build XGBoost if you are facing issues with the installation
# Synthetic Minority Over Sampling Technique
sm = SMOTE(sampling_strategy=1, k_neighbors=5, random_state=1)
X_train_over, y_train_over = sm.fit_resample(X_train, y_train)
print("Before OverSampling, count of label '1': {}".format(sum(y_train == 1)))
print("Before OverSampling, count of label '0': {} \n".format(sum(y_train == 0)))
print("After OverSampling, count of label '1': {}".format(sum(y_train_over == 1)))
print("After OverSampling, count of label '0': {} \n".format(sum(y_train_over == 0)))
print("After OverSampling, the shape of train_X: {}".format(X_train_over.shape))
print("After OverSampling, the shape of train_y: {} \n".format(y_train_over.shape))
Before OverSampling, count of label '1': 5949 Before OverSampling, count of label '0': 1139 After OverSampling, count of label '1': 5949 After OverSampling, count of label '0': 5949 After OverSampling, the shape of train_X: (11898, 30) After OverSampling, the shape of train_y: (11898,)
Bagging (Over Sampled)ΒΆ
bagging_estimator_o=BaggingClassifier(random_state=1)
bagging_estimator_o.fit(X_train_over,y_train_over)
BaggingClassifier(random_state=1)In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
BaggingClassifier(random_state=1)
bagging_estimator_o_train_perf = model_performance_classification_sklearn(bagging_estimator_o, X_train_over, y_train_over)
print("Training performance \n",bagging_estimator_o_train_perf)
Training performance
Accuracy Recall Precision F1
0 1.00 1.00 1.00 1.00
bagging_estimator_o_test_perf = model_performance_classification_sklearn(bagging_estimator_o, X_test, y_test)
print("Testing performance \n",bagging_estimator_o_test_perf)
Testing performance
Accuracy Recall Precision F1
0 0.95 0.96 0.98 0.97
Random Forest (Over Sampled)ΒΆ
rf_o = RandomForestClassifier(random_state=1)
rf_o.fit(X_train_over,y_train_over)
RandomForestClassifier(random_state=1)In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
RandomForestClassifier(random_state=1)
rf_train_perf_o = model_performance_classification_sklearn(rf_o, X_train_over, y_train_over)
print("Training performance \n",rf_train_perf_o)
Training performance
Accuracy Recall Precision F1
0 1.00 1.00 1.00 1.00
rf_test_perf_o = model_performance_classification_sklearn(rf_o, X_test, y_test)
print("Testing performance \n",rf_test_perf_o)
Testing performance
Accuracy Recall Precision F1
0 0.96 0.97 0.98 0.97
AdaBoost (Over Sampled)ΒΆ
ab_regressor_o = AdaBoostClassifier(random_state=1)
ab_regressor_o.fit(X_train_over,y_train_over)
AdaBoostClassifier(random_state=1)In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
AdaBoostClassifier(random_state=1)
#Train performance
ab_regressor_o_train_perf = model_performance_classification_sklearn(ab_regressor_o, X_train_over, y_train_over)
print("Training performance \n",ab_regressor_o_train_perf)
Training performance
Accuracy Recall Precision F1
0 0.96 0.96 0.96 0.96
#Test performance
ab_regressor_o_test_perf = model_performance_classification_sklearn(ab_regressor_o, X_test, y_test)
print("Testing performance \n",ab_regressor_o_test_perf)
Testing performance
Accuracy Recall Precision F1
0 0.95 0.96 0.98 0.97
GradientBoosting (Over Sampled)ΒΆ
#GradientBoost Classifier
gb_estimator_o = GradientBoostingClassifier(random_state=1)
gb_estimator_o.fit(X_train_over,y_train_over)
GradientBoostingClassifier(random_state=1)In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
GradientBoostingClassifier(random_state=1)
#Train performance
gb_estimator_o_train_perf = model_performance_classification_sklearn(gb_estimator_o, X_train_over, y_train_over)
print("Training performance \n",gb_estimator_o_train_perf)
Training performance
Accuracy Recall Precision F1
0 0.98 0.97 0.98 0.98
#Test performance
gb_estimator_o_test_perf = model_performance_classification_sklearn(gb_estimator_o, X_test, y_test)
print("Testing performance \n",gb_estimator_o_test_perf)
Testing performance
Accuracy Recall Precision F1
0 0.96 0.97 0.98 0.98
Decision Tree (Over Sampled)ΒΆ
dtree_o=DecisionTreeClassifier(random_state=1)
dtree_o.fit(X_train_over,y_train_over)
DecisionTreeClassifier(random_state=1)In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
DecisionTreeClassifier(random_state=1)
dtree_model_o_train_perf = model_performance_classification_sklearn(dtree_o, X_train_over, y_train_over)
print("Training performance \n",dtree_model_o_train_perf)
Training performance
Accuracy Recall Precision F1
0 1.00 1.00 1.00 1.00
dtree_model_o_test_perf = model_performance_classification_sklearn(dtree_o, X_test, y_test)
print("Testing performance \n",dtree_model_o_test_perf)
Testing performance
Accuracy Recall Precision F1
0 0.93 0.95 0.97 0.96
Model Building with Undersampled dataΒΆ
Undersample the train data - Build 5 models (from decision trees, bagging and boosting methods) - Comment on the model performance * You can choose NOT to build XGBoost if you are facing issues with the installation
# Random undersampler for under sampling the data
rus = RandomUnderSampler(random_state=1, sampling_strategy=1)
X_train_un, y_train_un = rus.fit_resample(X_train, y_train)
print("Before UnderSampling, count of label '1': {}".format(sum(y_train == 1)))
print("Before UnderSampling, count of label '0': {} \n".format(sum(y_train == 0)))
print("After UnderSampling, count of label '1': {}".format(sum(y_train_un == 1)))
print("After UnderSampling, count of label '0': {} \n".format(sum(y_train_un == 0)))
print("After UnderSampling, the shape of train_X: {}".format(X_train_un.shape))
print("After UnderSampling, the shape of train_y: {} \n".format(y_train_un.shape))
Before UnderSampling, count of label '1': 5949 Before UnderSampling, count of label '0': 1139 After UnderSampling, count of label '1': 1139 After UnderSampling, count of label '0': 1139 After UnderSampling, the shape of train_X: (2278, 30) After UnderSampling, the shape of train_y: (2278,)
Bagging (Under Sampled)ΒΆ
bagging_estimator_u=BaggingClassifier(random_state=1)
bagging_estimator_u.fit(X_train_un,y_train_un)
BaggingClassifier(random_state=1)In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
BaggingClassifier(random_state=1)
bagging_estimator_u_train_perf = model_performance_classification_sklearn(bagging_estimator_u, X_train_un, y_train_un)
print("Training performance \n",bagging_estimator_u_train_perf)
Training performance
Accuracy Recall Precision F1
0 1.00 0.99 1.00 1.00
bagging_estimator_u_test_perf = model_performance_classification_sklearn(bagging_estimator_u, X_test, y_test)
print("Testing performance \n",bagging_estimator_u_test_perf)
Testing performance
Accuracy Recall Precision F1
0 0.91 0.91 0.98 0.95
Random Forest (Under Sampled)ΒΆ
rf_u = RandomForestClassifier(random_state=1)
rf_u.fit(X_train_un,y_train_un)
RandomForestClassifier(random_state=1)In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
RandomForestClassifier(random_state=1)
rf_train_perf_u = model_performance_classification_sklearn(rf_u, X_train_un, y_train_un)
print("Training performance \n",rf_train_perf_u)
Training performance
Accuracy Recall Precision F1
0 1.00 1.00 1.00 1.00
rf_test_perf_u = model_performance_classification_sklearn(rf_u, X_test, y_test)
print("Testing performance \n",rf_test_perf_u)
Testing performance
Accuracy Recall Precision F1
0 0.92 0.92 0.99 0.95
AdaBoost (Under Sampled)ΒΆ
ab_regressor_u = AdaBoostClassifier(random_state=1)
ab_regressor_u.fit(X_train_un,y_train_un)
AdaBoostClassifier(random_state=1)In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
AdaBoostClassifier(random_state=1)
#Train performance
ab_regressor_u_train_perf = model_performance_classification_sklearn(ab_regressor_u, X_train_un, y_train_un)
print("Training performance \n",ab_regressor_u_train_perf)
Training performance
Accuracy Recall Precision F1
0 0.95 0.95 0.96 0.95
#Test performance
ab_regressor_u_test_perf = model_performance_classification_sklearn(ab_regressor_u, X_test, y_test)
print("Testing performance \n",ab_regressor_u_test_perf)
Testing performance
Accuracy Recall Precision F1
0 0.93 0.93 0.99 0.96
GradientBoosting (Under Sampled)ΒΆ
#GradientBoost Classifier
gb_estimator_u = GradientBoostingClassifier(random_state=1)
gb_estimator_u.fit(X_train_un,y_train_un)
GradientBoostingClassifier(random_state=1)In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
GradientBoostingClassifier(random_state=1)
#Train performance
gb_estimator_u_train_perf = model_performance_classification_sklearn(gb_estimator_u, X_train_un, y_train_un)
print("Training performance \n",gb_estimator_u_train_perf)
Training performance
Accuracy Recall Precision F1
0 0.98 0.97 0.98 0.98
#Test performance
gb_estimator_u_test_perf = model_performance_classification_sklearn(gb_estimator_u, X_test, y_test)
print("Testing performance \n",gb_estimator_u_test_perf)
Testing performance
Accuracy Recall Precision F1
0 0.95 0.95 0.99 0.97
Decision Tree (Under Sampled)ΒΆ
dtree_u=DecisionTreeClassifier(random_state=1)
dtree_u.fit(X_train_un,y_train_un)
DecisionTreeClassifier(random_state=1)In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
DecisionTreeClassifier(random_state=1)
dtree_model_u_train_perf = model_performance_classification_sklearn(dtree_u, X_train_un, y_train_un)
print("Training performance \n",dtree_model_o_train_perf)
Training performance
Accuracy Recall Precision F1
0 1.00 1.00 1.00 1.00
dtree_model_u_test_perf = model_performance_classification_sklearn(dtree_u, X_test, y_test)
print("Testing performance \n",dtree_model_u_test_perf)
Testing performance
Accuracy Recall Precision F1
0 0.90 0.91 0.98 0.94
Model Comparison to Choose Best 3 Models for Next StepΒΆ
# training performance comparison
models_train_comp_df = pd.concat(
[bagging_estimator_train_perf.T, bagging_estimator_o_train_perf.T, bagging_estimator_u_train_perf.T,
rf_train_perf.T, rf_train_perf_o.T, rf_train_perf_u.T,
ab_regressor_train_perf.T, ab_regressor_o_train_perf.T, ab_regressor_u_train_perf.T,
gb_estimator_train_perf.T, gb_estimator_o_train_perf.T, gb_estimator_u_train_perf.T,
dtree_model_train_perf.T, dtree_model_o_train_perf.T, dtree_model_u_train_perf.T],
axis=1,
)
models_train_comp_df.columns = [
"Bagging Estimator",
"Bagging Estimator O",
"Bagging Estimator U",
"Random Forest",
"Random Forest O",
"Random Forest U",
"AdaBoost",
"AdaBoost O",
"AdaBoost U",
"GradientBoost",
"GradientBoost O",
"GradientBoost U",
"Decision Tree",
"Decision Tree O",
"Decision Tree U"
]
# Sort columns based on the 'Recall' row
models_train_comp_df = models_train_comp_df.T.sort_values(by='Recall', ascending=False).T
print("Training performance comparison:")
models_train_comp_df
Training performance comparison:
| Random Forest | Random Forest O | Random Forest U | Decision Tree | Decision Tree O | Decision Tree U | Bagging Estimator | Bagging Estimator O | Bagging Estimator U | GradientBoost | AdaBoost | GradientBoost O | GradientBoost U | AdaBoost O | AdaBoost U | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Accuracy | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 0.98 | 0.96 | 0.98 | 0.98 | 0.96 | 0.95 |
| Recall | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 0.99 | 0.99 | 0.98 | 0.97 | 0.97 | 0.96 | 0.95 |
| Precision | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 0.98 | 0.97 | 0.98 | 0.98 | 0.96 | 0.96 |
| F1 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 0.99 | 0.98 | 0.98 | 0.98 | 0.96 | 0.95 |
# test performance comparison
models_test_comp_df = pd.concat(
[bagging_estimator_test_perf.T, bagging_estimator_o_test_perf.T, bagging_estimator_u_test_perf.T,
rf_test_perf.T, rf_test_perf_o.T, rf_test_perf_u.T,
ab_regressor_test_perf.T, ab_regressor_o_test_perf.T, ab_regressor_u_test_perf.T,
gb_estimator_test_perf.T, gb_estimator_o_test_perf.T, gb_estimator_u_test_perf.T,
dtree_model_test_perf.T, dtree_model_o_test_perf.T, dtree_model_u_test_perf.T],
axis=1,
)
models_test_comp_df.columns = [
"Bagging Estimator",
"Bagging Estimator O",
"Bagging Estimator U",
"Random Forest",
"Random Forest O",
"Random Forest U",
"AdaBoost",
"AdaBoost O",
"AdaBoost U",
"GradientBoost",
"GradientBoost O",
"GradientBoost U",
"Decision Tree",
"Decision Tree O",
"Decision Tree U"
]
# Sort columns based on the 'Recall' row
models_test_comp_df = models_test_comp_df.T.sort_values(by='Recall', ascending=False).T
print("Testing performance comparison:")
models_test_comp_df
Testing performance comparison:
| Random Forest | GradientBoost | AdaBoost | Random Forest O | GradientBoost O | Bagging Estimator | Decision Tree | Bagging Estimator O | AdaBoost O | Decision Tree O | GradientBoost U | AdaBoost U | Random Forest U | Bagging Estimator U | Decision Tree U | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Accuracy | 0.95 | 0.96 | 0.96 | 0.96 | 0.96 | 0.95 | 0.94 | 0.95 | 0.95 | 0.93 | 0.95 | 0.93 | 0.92 | 0.91 | 0.90 |
| Recall | 0.99 | 0.99 | 0.98 | 0.97 | 0.97 | 0.97 | 0.96 | 0.96 | 0.96 | 0.95 | 0.95 | 0.93 | 0.92 | 0.91 | 0.91 |
| Precision | 0.96 | 0.97 | 0.97 | 0.98 | 0.98 | 0.97 | 0.96 | 0.98 | 0.98 | 0.97 | 0.99 | 0.99 | 0.99 | 0.98 | 0.98 |
| F1 | 0.97 | 0.98 | 0.98 | 0.97 | 0.98 | 0.97 | 0.96 | 0.97 | 0.97 | 0.96 | 0.97 | 0.96 | 0.95 | 0.95 | 0.94 |
In this particular problem, Recall is the best peformance measure to optimize, since attempts to maxium the True Positive Rate, so we don't overlook or miss any customer who is a risk of cancelling their credit cards. Per the above our top 3 performers are
- Random Forest
- GradientBoost
- AdaBoost
HyperparameterTuningΒΆ
Choose models that might perform better after tuning (tune at least 3 models out of 15 built in the previous steps) - Provide proper reasoning for tuning that model - Tune the best 3 models obtained above using randomized search and metric of interest - Check the performance of 3 tuned models.
Sample Parameter GridsΒΆ
Note
- Sample parameter grids have been provided to do necessary hyperparameter tuning. These sample grids are expected to provide a balance between model performance improvement and execution time. One can extend/reduce the parameter grid based on execution time and system configuration.
- Please note that if the parameter grid is extended to improve the model performance further, the execution time will increase
- For Gradient Boosting:
param_grid = {
"init": [AdaBoostClassifier(random_state=1),DecisionTreeClassifier(random_state=1)],
"n_estimators": np.arange(50,110,25),
"learning_rate": [0.01,0.1,0.05],
"subsample":[0.7,0.9],
"max_features":[0.5,0.7,1],
}
- For Adaboost:
param_grid = {
"n_estimators": np.arange(50,110,25),
"learning_rate": [0.01,0.1,0.05],
"base_estimator": [
DecisionTreeClassifier(max_depth=2, random_state=1),
DecisionTreeClassifier(max_depth=3, random_state=1),
],
}
- For Bagging Classifier:
param_grid = {
'max_samples': [0.8,0.9,1],
'max_features': [0.7,0.8,0.9],
'n_estimators' : [30,50,70],
}
- For Random Forest:
param_grid = {
"n_estimators": [50,110,25],
"min_samples_leaf": np.arange(1, 4),
"max_features": [np.arange(0.3, 0.6, 0.1),'sqrt'],
"max_samples": np.arange(0.4, 0.7, 0.1)
}
- For Decision Trees:
param_grid = {
'max_depth': np.arange(2,6),
'min_samples_leaf': [1, 4, 7],
'max_leaf_nodes' : [10, 15],
'min_impurity_decrease': [0.0001,0.001]
}
- For XGBoost (optional):
param_grid={'n_estimators':np.arange(50,110,25),
'scale_pos_weight':[1,2,5],
'learning_rate':[0.01,0.1,0.05],
'gamma':[1,3],
'subsample':[0.7,0.9]
}
Tuning of the Random Forest ModelΒΆ
# defining model
rf = RandomForestClassifier(random_state=1)
# Parameter grid to pass in RandomSearchCV
param_grid = {
"n_estimators": [50,110,25],
"min_samples_leaf": np.arange(1, 4),
"max_features": [np.arange(0.3, 0.6, 0.1),'sqrt'],
"max_samples": np.arange(0.4, 0.7, 0.1)
}
# Type of scoring used to compare parameter combinations
scorer = metrics.make_scorer(metrics.recall_score)
#Calling RandomizedSearchCV
randomized_cv = RandomizedSearchCV(estimator=rf, param_distributions=param_grid, n_iter=10, n_jobs = -1, scoring=scorer, cv=5, random_state=1)
#Fitting parameters in RandomizedSearchCV
randomized_cv.fit(X_train,y_train)
print("Best parameters are {} with CV score={}:" .format(randomized_cv.best_params_,randomized_cv.best_score_))
Best parameters are {'n_estimators': 50, 'min_samples_leaf': 2, 'max_samples': 0.4, 'max_features': 'sqrt'} with CV score=0.9873935444657258:
# Set the model to the best combination of parameters
rf_tuned = RandomForestClassifier(
random_state=1,
n_estimators=50,
min_samples_leaf=2,
max_samples=0.4,
max_features='sqrt'
)
# Fit the best algorithm to the data
rf_tuned.fit(X_train, y_train)
RandomForestClassifier(max_samples=0.4, min_samples_leaf=2, n_estimators=50,
random_state=1)In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook. On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
RandomForestClassifier(max_samples=0.4, min_samples_leaf=2, n_estimators=50,
random_state=1)rf_tuned_train_perf = model_performance_classification_sklearn(rf_tuned, X_train, y_train)
print("Training performance \n",rf_tuned_train_perf)
Training performance
Accuracy Recall Precision F1
0 0.98 0.99 0.98 0.99
rf_tuned_test_perf = model_performance_classification_sklearn(rf_tuned, X_test, y_test)
print("Testing performance \n",rf_tuned_test_perf)
Training performance
Accuracy Recall Precision F1
0 0.94 0.99 0.95 0.97
Tuning of the GradientBoost ModelΒΆ
# defining model
gb_estimator = GradientBoostingClassifier(random_state=1)
# Parameter grid to pass in RandomSearchCV
param_grid = {
"init": [AdaBoostClassifier(random_state=1),DecisionTreeClassifier(random_state=1)],
"n_estimators": np.arange(50,110,25),
"learning_rate": [0.01,0.1,0.05],
"subsample":[0.7,0.9],
"max_features":[0.5,0.7,1],
}
# Type of scoring used to compare parameter combinations
scorer = metrics.make_scorer(metrics.recall_score)
#Calling RandomizedSearchCV
randomized_cv = RandomizedSearchCV(estimator=gb_estimator, param_distributions=param_grid, n_iter=10, n_jobs = -1, scoring=scorer, cv=5, random_state=1)
#Fitting parameters in RandomizedSearchCV
randomized_cv.fit(X_train,y_train)
print("Best parameters are {} with CV score={}:" .format(randomized_cv.best_params_,randomized_cv.best_score_))
Best parameters are {'subsample': 0.7, 'n_estimators': 50, 'max_features': 1, 'learning_rate': 0.05, 'init': AdaBoostClassifier(random_state=1)} with CV score=0.9996637241944718:
# Set the model to the best combination of parameters
gb_estimator_tuned = GradientBoostingClassifier(
random_state=1,
subsample=0.7,
n_estimators=50,
max_features=1,
learning_rate=0.05,
init=AdaBoostClassifier(random_state=1)
)
# Fit the best algorithm to the data
gb_estimator_tuned.fit(X_train, y_train)
GradientBoostingClassifier(init=AdaBoostClassifier(random_state=1),
learning_rate=0.05, max_features=1, n_estimators=50,
random_state=1, subsample=0.7)In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook. On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
GradientBoostingClassifier(init=AdaBoostClassifier(random_state=1),
learning_rate=0.05, max_features=1, n_estimators=50,
random_state=1, subsample=0.7)AdaBoostClassifier(random_state=1)
AdaBoostClassifier(random_state=1)
gb_tuned_train_perf = model_performance_classification_sklearn(gb_estimator_tuned, X_train, y_train)
print("Training performance \n",gb_tuned_train_perf)
Training performance
Accuracy Recall Precision F1
0 0.86 1.00 0.86 0.92
gb_tuned_test_perf = model_performance_classification_sklearn(gb_estimator_tuned, X_test, y_test)
print("Testing performance \n",gb_tuned_test_perf)
Testing performance
Accuracy Recall Precision F1
0 0.86 1.00 0.86 0.92
Tuning of the AdaBoost ModelΒΆ
# defining model
ab = AdaBoostClassifier(random_state=1)
# Parameter grid to pass in RandomSearchCV
param_grid = {
"n_estimators": np.arange(50,110,25),
"learning_rate": [0.01,0.1,0.05],
"base_estimator": [
DecisionTreeClassifier(max_depth=2, random_state=1),
DecisionTreeClassifier(max_depth=3, random_state=1),
],
}
# Type of scoring used to compare parameter combinations
scorer = metrics.make_scorer(metrics.recall_score)
#Calling RandomizedSearchCV
randomized_cv = RandomizedSearchCV(estimator=ab, param_distributions=param_grid, n_iter=10, n_jobs = -1, scoring=scorer, cv=5, random_state=1)
#Fitting parameters in RandomizedSearchCV
randomized_cv.fit(X_train,y_train)
print("Best parameters are {} with CV score={}:" .format(randomized_cv.best_params_,randomized_cv.best_score_))
Best parameters are {'n_estimators': 75, 'learning_rate': 0.1, 'base_estimator': DecisionTreeClassifier(max_depth=3, random_state=1)} with CV score=0.9889060081559957:
# Set the model to the best combination of parameters
ab_tuned = AdaBoostClassifier(
random_state=1,
n_estimators=75,
learning_rate=0.1,
base_estimator=DecisionTreeClassifier(max_depth=3, random_state=1)
)
# Fit the best algorithm to the data
ab_tuned.fit(X_train, y_train)
AdaBoostClassifier(base_estimator=DecisionTreeClassifier(max_depth=3,
random_state=1),
learning_rate=0.1, n_estimators=75, random_state=1)In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook. On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
AdaBoostClassifier(base_estimator=DecisionTreeClassifier(max_depth=3,
random_state=1),
learning_rate=0.1, n_estimators=75, random_state=1)DecisionTreeClassifier(max_depth=3, random_state=1)
DecisionTreeClassifier(max_depth=3, random_state=1)
ab_tuned_train_perf = model_performance_classification_sklearn(ab_tuned, X_train, y_train)
print("Training performance \n",ab_tuned_train_perf)
Training performance
Accuracy Recall Precision F1
0 0.98 0.99 0.98 0.99
ab_tuned_test_perf = model_performance_classification_sklearn(ab_tuned, X_test, y_test)
print("Testing performance \n",ab_tuned_test_perf)
Testing performance
Accuracy Recall Precision F1
0 0.97 0.99 0.97 0.98
Final Model Comparison of Tuned ModelsΒΆ
# training performance comparison
models_train_comp_df = pd.concat(
[rf_train_perf.T, rf_tuned_train_perf.T,
ab_regressor_train_perf.T, ab_tuned_train_perf.T,
gb_estimator_train_perf.T, gb_tuned_train_perf.T],
axis=1,
)
models_train_comp_df.columns = [
"Random Forest",
"Random Forest Tuned",
"AdaBoost",
"AdaBoost Tuned",
"GradientBoost",
"GradientBoost Tuned"
]
# Sort columns based on the 'Recall' row
#models_train_comp_df = models_train_comp_df.T.sort_values(by='Recall', ascending=False).T
print("Training performance comparison:")
models_train_comp_df
Training performance comparison:
| Random Forest | Random Forest Tuned | AdaBoost | AdaBoost Tuned | GradientBoost | GradientBoost Tuned | |
|---|---|---|---|---|---|---|
| Accuracy | 1.00 | 0.98 | 0.96 | 0.98 | 0.98 | 0.86 |
| Recall | 1.00 | 0.99 | 0.98 | 0.99 | 0.99 | 1.00 |
| Precision | 1.00 | 0.98 | 0.97 | 0.98 | 0.98 | 0.86 |
| F1 | 1.00 | 0.99 | 0.98 | 0.99 | 0.99 | 0.92 |
# testing performance comparison
models_test_comp_df = pd.concat(
[rf_test_perf.T, rf_tuned_test_perf.T,
ab_regressor_test_perf.T, ab_tuned_test_perf.T,
gb_estimator_test_perf.T, gb_tuned_test_perf.T],
axis=1,
)
models_test_comp_df.columns = [
"Random Forest",
"Random Forest Tuned",
"AdaBoost",
"AdaBoost Tuned",
"GradientBoost",
"GradientBoost Tuned"
]
# Sort columns based on the 'Recall' row
#models_test_comp_df = models_train_comp_df.T.sort_values(by='Recall', ascending=False).T
print("Testing performance comparison:")
models_test_comp_df
Testing performance comparison:
| Random Forest | Random Forest Tuned | AdaBoost | AdaBoost Tuned | GradientBoost | GradientBoost Tuned | |
|---|---|---|---|---|---|---|
| Accuracy | 0.95 | 0.94 | 0.96 | 0.97 | 0.96 | 0.86 |
| Recall | 0.99 | 0.99 | 0.98 | 0.99 | 0.99 | 1.00 |
| Precision | 0.96 | 0.95 | 0.97 | 0.97 | 0.97 | 0.86 |
| F1 | 0.97 | 0.97 | 0.98 | 0.98 | 0.98 | 0.92 |
GradientBoost Tuned appears to give the best performance with a perfect Recall score of 1 on both Test and Train data stamples. Let's look at the confusion matrices for this model along with the feature importances.
#Train Confusion Matrix
plot_confusion_matrix(gb_estimator_tuned, X_train, y_train)
confusion_matrix(y_train, gb_estimator_tuned.predict(X_train))
array([[ 164, 975],
[ 1, 5948]])
confusion_matrix(y_test, gb_estimator_tuned.predict(X_test))
array([[ 61, 427],
[ 0, 2551]])
#Test Confusion Matrix
plot_confusion_matrix(gb_estimator_tuned, X_test, y_test)
feature_names = X_train.columns
importances = gb_estimator_tuned.feature_importances_
indices = np.argsort(importances)
plt.figure(figsize=(12,12))
plt.title('Feature Importances')
plt.barh(range(len(indices)), importances[indices], color='violet', align='center')
plt.yticks(range(len(indices)), [feature_names[i] for i in indices])
plt.xlabel('Relative Importance')
plt.show()
Business Insights and ConclusionsΒΆ
The dominant factors for predicting if a customer will drop their credit cards are:
- Total_Trans_Amt: Total Transaction Amount (Last 12 months)
- Total_Revolving_Bal: The balance that carries over from one month to the next
- Total_Ct_Chng_Q4_Q1: Ratio of the total transaction count in 4th quarter and the total transaction count in the 1st quarter
- Avg_Utilization_Ratio: Represents how much of the available credit the customer spent
Customers with lower Total_Trans_Amt are more likely to drop. Customers with a lower Total_Revolving_Bal are more likely to drop. Customers with a lower Total_Ct_Chng_Q4_Q1 are more likely to drop. Customers with a low Avg_Utilization_Ratio are more likelty to drop.