Recommendation System Project Overview¶
Context¶
Online E-commerce websites like Amazon, Flipkart uses different recommendation models to provide different suggestions to different users. Amazon currently uses item-to-item collaborative filtering, which scales to massive data sets and produces high-quality recommendations in real-time.
Objective¶
Build a recommendation system to recommend products to customers based on their previous ratings for other products. Apply the concepts and techniques you have learned in the previous weeks and summarise your insights at the end.
Dataset¶
We are using the Electronics dataset from the Amazon Reviews data repository (http://jmcauley.ucsd.edu/data/amazon/), which has several datasets.
Attribute Information
- userId: Every user identified with a unique id
- productId: Every product identified with a unique id
- Rating: Rating of the corresponding product by the corresponding user
- timestamp: Time of the rating ( ignore this column for this exercise)
Import Required Libraries¶
#install library joblib
!pip install joblib
Requirement already satisfied: joblib in /usr/local/lib/python3.11/dist-packages (1.4.2)
import numpy as np
import pandas as pd
import math
import json
import time
import matplotlib.pyplot as plt
import seaborn as sns
from sklearn.metrics.pairwise import cosine_similarity
from sklearn.model_selection import train_test_split
from sklearn.neighbors import NearestNeighbors
import joblib
import scipy.sparse
from scipy.sparse import csr_matrix
import warnings; warnings.simplefilter('ignore')
from scipy.sparse.linalg import svds
%matplotlib inline
Data Import¶
Part 1. Read and explore the dataset. ( Rename column, plot histograms, find data characteristics)¶
# uncomment and run the following line if using Google Colab
from google.colab import drive
drive.mount('/content/drive')
Mounted at /content/drive
#Import the data set
df = pd.read_csv('/content/drive/MyDrive/Personal/UT Austin/Recommendation Systems/Project/ratings_Electronics.csv', header=None)
df.columns = ['user_id', 'prod_id', 'rating', 'prod_name']
df = df.drop('prod_name', axis=1)
df_copy = df.copy(deep=True)
# see few rows of the imported dataset
df.tail()
| user_id | prod_id | rating | |
|---|---|---|---|
| 7824477 | A2YZI3C9MOHC0L | BT008UKTMW | 5.0 |
| 7824478 | A322MDK0M89RHN | BT008UKTMW | 5.0 |
| 7824479 | A1MH90R0ADMIK0 | BT008UKTMW | 4.0 |
| 7824480 | A10M2KEFPEQDHN | BT008UKTMW | 4.0 |
| 7824481 | A2G81TMIOIDEQQ | BT008V9J9U | 5.0 |
# Check the number of rows and columns
rows, columns = df.shape
print("No of rows: ", rows)
print("No of columns: ", columns)
No of rows: 7824482 No of columns: 3
We have ~7.8M rows of data.
#Check Data types
df.dtypes
user_id object prod_id object rating float64 dtype: object
# Check for missing values present
print('Number of missing values across columns-\n', df.isnull().sum())
Number of missing values across columns- user_id 0 prod_id 0 rating 0 dtype: int64
There are no missing values.
# Summary statistics of 'rating' variable
df[['rating']].describe().transpose()
| count | mean | std | min | 25% | 50% | 75% | max | |
|---|---|---|---|---|---|---|---|---|
| rating | 7824482.0 | 4.012337 | 1.38091 | 1.0 | 3.0 | 5.0 | 5.0 | 5.0 |
# find minimum and maximum ratings
def find_min_max_rating():
print('The minimum rating is: %d' %(df['rating'].min()))
print('The maximum rating is: %d' %(df['rating'].max()))
find_min_max_rating()
The minimum rating is: 1 The maximum rating is: 5
Ratings are on scale of 1 - 5
# function to create labeled barplots
def labeled_barplot(data, feature, perc=False, n=None):
"""
Barplot with percentage at the top
data: dataframe
feature: dataframe column
perc: whether to display percentages instead of count (default is False)
n: displays the top n category levels (default is None, i.e., display all levels)
"""
total = len(data[feature]) # length of the column
count = data[feature].nunique()
if n is None:
plt.figure(figsize=(count + 1, 5))
else:
plt.figure(figsize=(n + 1, 5))
plt.xticks(rotation=90, fontsize=15)
ax = sns.countplot(
data=data,
x=feature,
hue=feature,
palette="Paired",
order=data[feature].value_counts().index[:n].sort_values(),
)
for p in ax.patches:
if perc == True:
label = "{:.1f}%".format(
100 * p.get_height() / total
) # percentage of each class of the category
else:
label = p.get_height() # count of each level of the category
x = p.get_x() + p.get_width() / 2 # width of the plot
y = p.get_height() # height of the plot
ax.annotate(
label,
(x, y),
ha="center",
va="center",
size=12,
xytext=(0, 5),
textcoords="offset points",
) # annotate the percentage
plt.show() # show the plot
labeled_barplot(df, "rating", perc=True)
# Number of unique user id and product id in the data
print('Number of unique USERS in Raw data = {:,}'.format(df['user_id'].nunique()))
print('Number of unique ITEMS in Raw data = {:,}'.format(df['prod_id'].nunique()))
Number of unique USERS in Raw data = 4,201,696 Number of unique ITEMS in Raw data = 476,002
user_counts = df.groupby('user_id').size().reset_index(name='row_count')
# Plot histogram with log scale on the y-axis
plt.figure(figsize=(10, 6))
sns.histplot(user_counts['row_count'], bins=30, kde=False)
plt.yscale('log') # Set y-axis to logarithmic scale
plt.xlabel('Number of Rows per User', fontsize=12)
plt.ylabel('Number of Users (Log Scale)', fontsize=12)
plt.title('Distribution of Rows per User', fontsize=14)
plt.grid(axis='y')
plt.show()
Here we can see there area huge number of users, who have given very few reviews.
Part 2. Take subset of dataset to make it less sparse/more dense. (Keep only users who gave 50 or more ratings)¶
# Top 10 users based on rating
most_rated = df.groupby('user_id').size().sort_values(ascending=False)[:10]
most_rated
| 0 | |
|---|---|
| user_id | |
| A5JLAU2ARJ0BO | 520 |
| ADLVFFE4VBT8 | 501 |
| A3OXHLG6DIBRW8 | 498 |
| A6FIAB28IS79 | 431 |
| A680RUE1FDO8B | 406 |
| A1ODOGXEYECQQ8 | 380 |
| A36K2N527TXXJN | 314 |
| A2AY4YUOX2N1BQ | 311 |
| AWPODHOB4GFWL | 308 |
| A25C2M3QF9G7OQ | 296 |
Data model preparation as per requirement on number of minimum ratings¶
counts = df['user_id'].value_counts()
df_final = df[df['user_id'].isin(counts[counts >= 50].index)]
print('Number of users who have rated 50 or more items = {:,}'.format(len(df_final)))
print('Number of unique USERS in final data = {:,}'.format(df_final['user_id'].nunique()))
print('Number of unique ITEMS in final data = {:,}'.format(df_final['prod_id'].nunique()))
Number of users who have rated 50 or more items = 125,871 Number of unique USERS in final data = 1,540 Number of unique ITEMS in final data = 48,190
df_final has users who have rated 50 or more items¶
df_final.head()
| user_id | prod_id | rating | |
|---|---|---|---|
| 94 | A3BY5KCNQZXV5U | 0594451647 | 5.0 |
| 118 | AT09WGFUM934H | 0594481813 | 3.0 |
| 177 | A32HSNCNPRUMTR | 0970407998 | 1.0 |
| 178 | A17HMM1M7T9PJ1 | 0970407998 | 4.0 |
| 492 | A3CLWR1UUZT6TG | 0972683275 | 5.0 |
Calculate the density of the rating matrix¶
final_ratings_matrix = df_final.pivot(index = 'user_id', columns ='prod_id', values = 'rating').fillna(0)
print('Shape of final_ratings_matrix: ', final_ratings_matrix.shape)
given_num_of_ratings = np.count_nonzero(final_ratings_matrix)
print('given_num_of_ratings = {:,}'.format(given_num_of_ratings))
possible_num_of_ratings = final_ratings_matrix.shape[0] * final_ratings_matrix.shape[1]
print('possible_num_of_ratings = {:,}'.format(possible_num_of_ratings))
density = (given_num_of_ratings/possible_num_of_ratings)
density *= 100
print ('density: {:4.2f}%'.format(density))
Shape of final_ratings_matrix: (1540, 48190) given_num_of_ratings = 125,871 possible_num_of_ratings = 74,212,600 density: 0.17%
Only 0.17% of the cells are populated with a rating value, so the matrix is very sparse!
final_ratings_matrix.tail()
| prod_id | 0594451647 | 0594481813 | 0970407998 | 0972683275 | 1400501466 | 1400501520 | 1400501776 | 1400532620 | 1400532655 | 140053271X | ... | B00L5YZCCG | B00L8I6SFY | B00L8QCVL6 | B00LA6T0LS | B00LBZ1Z7K | B00LED02VY | B00LGN7Y3G | B00LGQ6HL8 | B00LI4ZZO8 | B00LKG1MC8 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| user_id | |||||||||||||||||||||
| AZBXKUH4AIW3X | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | ... | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 |
| AZCE11PSTCH1L | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | ... | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 |
| AZMY6E8B52L2T | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | ... | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 |
| AZNUHQSHZHSUE | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | ... | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 |
| AZOK5STV85FBJ | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | ... | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 |
5 rows × 48190 columns
# Matrix with one row per 'Product' and one column per 'user' for Item-based CF
final_ratings_matrix_T = final_ratings_matrix.transpose()
final_ratings_matrix_T.head()
| user_id | A100UD67AHFODS | A100WO06OQR8BQ | A105S56ODHGJEK | A105TOJ6LTVMBG | A10AFVU66A79Y1 | A10H24TDLK2VDP | A10NMELR4KX0J6 | A10O7THJ2O20AG | A10PEXB6XAQ5XF | A10X9ME6R66JDX | ... | AYOTEJ617O60K | AYP0YPLSP9ISM | AZ515FFZ7I2P7 | AZ8XSDMIX04VJ | AZAC8O310IK4E | AZBXKUH4AIW3X | AZCE11PSTCH1L | AZMY6E8B52L2T | AZNUHQSHZHSUE | AZOK5STV85FBJ |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| prod_id | |||||||||||||||||||||
| 0594451647 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | ... | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 |
| 0594481813 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | ... | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 |
| 0970407998 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | ... | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 |
| 0972683275 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | ... | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 |
| 1400501466 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | ... | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 |
5 rows × 1540 columns
Part 3. Split the data randomly into train and test dataset. ( For example split it in 70/30 ratio)¶
#Split the training and test data in the ratio 70:30
train_data, test_data = train_test_split(df_final, test_size = 0.3, random_state=0)
print(train_data.head(5))
user_id prod_id rating 6595853 A2BYV7S1QP2YIG B009EAHVTA 5.0 4738241 AB094YABX21WQ B0056XCEAA 1.0 4175596 A3D0UM4ZD2CMAW B004I763AW 5.0 3753016 AATWFX0ZZSE6C B0040NPHMO 3.0 1734767 A1NNMOD9H36Q8E B0015VW3BM 4.0
def shape():
print("Test data shape: ", test_data.shape)
print("Train data shape: ", train_data.shape)
shape()
Test data shape: (37762, 3) Train data shape: (88109, 3)
Part 4. Build Popularity Recommender model. (Non-personalised)¶
#Count of user_id for each unique product as recommendation score
#The numer of users that interacted with a given product is taken as a proxy for its popularity
train_data_grouped = train_data.groupby('prod_id').agg({'user_id': 'count'}).reset_index()
train_data_grouped.rename(columns = {'user_id': 'score'},inplace=True)
train_data_grouped.head()
| prod_id | score | |
|---|---|---|
| 0 | 0594451647 | 1 |
| 1 | 0594481813 | 1 |
| 2 | 0970407998 | 1 |
| 3 | 0972683275 | 3 |
| 4 | 1400501466 | 4 |
Note: We haven't taken into account at all the rating the user gave, so another method could be to consider not only number of users that interacted with the product, but also the rating they gave. For example we could calculate the avg rating for a product and multiply that by the number of users for a ratings weighted popularity score.
#Sort the products in descending order of score (highest score first). If two or more products have the same score,
#they are further sorted by prod_id in ascending order.
train_data_sort = train_data_grouped.sort_values(['score', 'prod_id'], ascending = [0,1])
#Generate a recommendation rank based upon score
#ascending=0: Higher scores get a lower rank (i.e., the highest score gets rank 1).
train_data_sort['Rank'] = train_data_sort['score'].rank(ascending=0, method='first')
#Get the top 5 recommendations
popularity_recommendations = train_data_sort.head(5)
popularity_recommendations
| prod_id | score | Rank | |
|---|---|---|---|
| 30847 | B0088CJT4U | 133 | 1.0 |
| 30287 | B007WTAJTO | 124 | 2.0 |
| 19647 | B003ES5ZUU | 122 | 3.0 |
| 8752 | B000N99BBC | 114 | 4.0 |
| 30555 | B00829THK0 | 97 | 5.0 |
# Use popularity based recommender model to make predictions
def recommend(user_id):
user_recommendations = popularity_recommendations
#Add user_id column for which the recommendations are being generated
user_recommendations['user_id'] = user_id
#Bring user_id column to the front
cols = user_recommendations.columns.tolist()
cols = cols[-1:] + cols[:-1]
user_recommendations = user_recommendations[cols]
return user_recommendations
find_recom = [15,121,200] # This list is user choice.
for i in find_recom:
print("Here is the recommendation for the userId: %d\n" %(i))
print(recommend(i))
print("\n")
Here is the recommendation for the userId: 15
user_id prod_id score Rank
30847 15 B0088CJT4U 133 1.0
30287 15 B007WTAJTO 124 2.0
19647 15 B003ES5ZUU 122 3.0
8752 15 B000N99BBC 114 4.0
30555 15 B00829THK0 97 5.0
Here is the recommendation for the userId: 121
user_id prod_id score Rank
30847 121 B0088CJT4U 133 1.0
30287 121 B007WTAJTO 124 2.0
19647 121 B003ES5ZUU 122 3.0
8752 121 B000N99BBC 114 4.0
30555 121 B00829THK0 97 5.0
Here is the recommendation for the userId: 200
user_id prod_id score Rank
30847 200 B0088CJT4U 133 1.0
30287 200 B007WTAJTO 124 2.0
19647 200 B003ES5ZUU 122 3.0
8752 200 B000N99BBC 114 4.0
30555 200 B00829THK0 97 5.0
Since this is a popularity-based recommender model, recommendations remain the same for all users. We predict the products based on the popularity. It is not personalized to particular user
Step 5. Build Collaborative Filtering model.¶
A collaborative filtering model is a recommendation system technique that predicts a user's preferences for items (products in this case) based on their past behavior and the preferences of similar users. It operates under the assumption that users who agreed on certain items in the past are likely to agree on others in the future. Collaborative filtering can be user-based, focusing on similarities between users, or item-based, focusing on similarities between items. Unlike content-based methods, it doesn't require item features, relying instead on user-item interaction data (e.g., ratings or purchases). This makes it effective but also sensitive to sparse data or cold-start problems.
Model-based Collaborative Filtering: Singular Value Decomposition (SVD)¶
#Combine Training and Test Data
df_CF = pd.concat([train_data, test_data]).reset_index()
df_CF.tail()
| index | user_id | prod_id | rating | |
|---|---|---|---|---|
| 125866 | 621872 | A3OXHLG6DIBRW8 | B0007UQNOA | 3.0 |
| 125867 | 1942808 | A365PBEOWM7EI7 | B001DVZXC0 | 3.0 |
| 125868 | 5219963 | A3QDY9I0CNMD2W | B005WXQO3W | 5.0 |
| 125869 | 876608 | AR18DH5SL9F73 | B000EPR7AC | 5.0 |
| 125870 | 975289 | A3VL4RXCWNSR3H | B000GM7MRG | 5.0 |
#User-based Collaborative Filtering
# Matrix with row per 'user' and column per 'item'
pivot_df = df_CF.pivot(index = 'user_id', columns ='prod_id', values = 'rating').fillna(0)
print(pivot_df.shape)
pivot_df.head()
(1540, 48190)
| prod_id | 0594451647 | 0594481813 | 0970407998 | 0972683275 | 1400501466 | 1400501520 | 1400501776 | 1400532620 | 1400532655 | 140053271X | ... | B00L5YZCCG | B00L8I6SFY | B00L8QCVL6 | B00LA6T0LS | B00LBZ1Z7K | B00LED02VY | B00LGN7Y3G | B00LGQ6HL8 | B00LI4ZZO8 | B00LKG1MC8 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| user_id | |||||||||||||||||||||
| A100UD67AHFODS | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | ... | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 |
| A100WO06OQR8BQ | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | ... | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 |
| A105S56ODHGJEK | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | ... | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 |
| A105TOJ6LTVMBG | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | ... | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 |
| A10AFVU66A79Y1 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | ... | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 |
5 rows × 48190 columns
#Adds a new column user_index that assigns a unique numerical index to each user_id.
pivot_df['user_index'] = np.arange(0, pivot_df.shape[0], 1)
pivot_df.head()
| prod_id | 0594451647 | 0594481813 | 0970407998 | 0972683275 | 1400501466 | 1400501520 | 1400501776 | 1400532620 | 1400532655 | 140053271X | ... | B00L8I6SFY | B00L8QCVL6 | B00LA6T0LS | B00LBZ1Z7K | B00LED02VY | B00LGN7Y3G | B00LGQ6HL8 | B00LI4ZZO8 | B00LKG1MC8 | user_index |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| user_id | |||||||||||||||||||||
| A100UD67AHFODS | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | ... | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0 |
| A100WO06OQR8BQ | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | ... | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 1 |
| A105S56ODHGJEK | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | ... | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 2 |
| A105TOJ6LTVMBG | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | ... | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 3 |
| A10AFVU66A79Y1 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | ... | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 4 |
5 rows × 48191 columns
#Replaces the original user_id index with the user_index column to make numerical indexing easier for subsequent modeling
pivot_df.set_index(['user_index'], inplace=True)
# Actual ratings given by users
pivot_df.head()
| prod_id | 0594451647 | 0594481813 | 0970407998 | 0972683275 | 1400501466 | 1400501520 | 1400501776 | 1400532620 | 1400532655 | 140053271X | ... | B00L5YZCCG | B00L8I6SFY | B00L8QCVL6 | B00LA6T0LS | B00LBZ1Z7K | B00LED02VY | B00LGN7Y3G | B00LGQ6HL8 | B00LI4ZZO8 | B00LKG1MC8 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| user_index | |||||||||||||||||||||
| 0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | ... | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 |
| 1 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | ... | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 |
| 2 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | ... | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 |
| 3 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | ... | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 |
| 4 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | ... | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 |
5 rows × 48190 columns
# Convert pivot_df to a NumPy array
pivot_matrix = pivot_df.values
# Singular Value Decomposition
# The parameter k specifies the number of singular values and vectors to compute. It reduces the dimensionality of the matrix.
# In other words, k represents the number of latent (hidden) features used to represent each user in the data, which then allows
# identifying users that are similar to each other
# U Represents users in the latent feature space. Each row corresponds to a user, with k-dimensional features capturing their preferences.
# sigma represents the importance of each latent feature
# Vt represents the products in the latent feature space, each row corresponds to an item, with k-dimensional features capturing its characteristics
U, sigma, Vt = svds(pivot_matrix, k=50)
# Construct diagonal array in SVD
sigma = np.diag(sigma)
Perform matrix multiplication of the three components from SVD.
U: The user matrix, where rows represent users in the latent feature space.
sigma: The diagonal matrix of singular values, capturing the importance of each latent feature.
Vt: The item matrix, where columns represent items in the latent feature space.
The result, all_user_predicted_ratings, is an approximate reconstruction of the original user-item matrix (pivot_df), with predicted ratings replacing the original values (including zeros where ratings were missing).
# Reconstruct the predicted ratings matrix
all_user_predicted_ratings = np.dot(np.dot(U, sigma), Vt)
# Convert the predictions to a DataFrame, preserving user_id as the index
preds_df = pd.DataFrame(all_user_predicted_ratings, index=pivot_df.index, columns=pivot_df.columns)
# Display the first few rows
preds_df.head()
| prod_id | 0594451647 | 0594481813 | 0970407998 | 0972683275 | 1400501466 | 1400501520 | 1400501776 | 1400532620 | 1400532655 | 140053271X | ... | B00L5YZCCG | B00L8I6SFY | B00L8QCVL6 | B00LA6T0LS | B00LBZ1Z7K | B00LED02VY | B00LGN7Y3G | B00LGQ6HL8 | B00LI4ZZO8 | B00LKG1MC8 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| user_index | |||||||||||||||||||||
| 0 | 0.005086 | 0.002178 | 0.003668 | -0.040843 | 0.009640 | 0.006808 | 0.020659 | 0.000649 | 0.020331 | 0.005633 | ... | 0.000238 | -0.061477 | 0.001214 | -0.123433 | 0.028490 | 0.016109 | 0.002855 | -0.174568 | 0.011367 | -0.012997 |
| 1 | 0.002286 | -0.010898 | -0.000724 | 0.130259 | 0.007506 | -0.003350 | 0.063711 | -0.000674 | 0.016111 | -0.002433 | ... | -0.000038 | 0.013766 | 0.001473 | 0.025588 | -0.042103 | 0.004251 | 0.002177 | -0.024362 | -0.014765 | 0.038570 |
| 2 | -0.001655 | -0.002675 | -0.007355 | 0.007264 | 0.005152 | -0.003986 | -0.003480 | 0.006961 | -0.006606 | -0.002719 | ... | -0.001708 | -0.051040 | 0.000325 | -0.054867 | 0.017870 | -0.004996 | -0.002426 | 0.083928 | -0.112205 | 0.005964 |
| 3 | 0.001856 | 0.011019 | -0.005910 | -0.014134 | 0.000179 | 0.001877 | -0.005391 | -0.001709 | 0.004968 | 0.001402 | ... | 0.000582 | -0.009326 | -0.000465 | -0.048315 | 0.023302 | 0.006790 | 0.003380 | 0.005460 | -0.015263 | -0.025996 |
| 4 | 0.001115 | -0.002670 | 0.011018 | 0.014434 | 0.010319 | 0.006002 | 0.017151 | 0.003726 | 0.001404 | 0.005645 | ... | 0.000207 | 0.023761 | 0.000747 | -0.019347 | -0.012749 | 0.001026 | 0.001364 | -0.020580 | 0.011828 | 0.012770 |
5 rows × 48190 columns
# Recommend the items with the highest predicted ratings
def recommend_items(userID, pivot_df, preds_df, num_recommendations):
user_idx = userID-1 # index starts at 0
# Get and sort the user's ratings
sorted_user_ratings = pivot_df.iloc[user_idx].sort_values(ascending=False)
#sorted_user_ratings
sorted_user_predictions = preds_df.iloc[user_idx].sort_values(ascending=False)
#sorted_user_predictions
temp = pd.concat([sorted_user_ratings, sorted_user_predictions], axis=1)
temp.index.name = 'Recommended Items'
temp.columns = ['user_ratings', 'user_predictions']
temp = temp.loc[temp.user_ratings == 0]
temp = temp.sort_values('user_predictions', ascending=False)
print('\nBelow are the recommended items for user(user_id = {}):\n'.format(userID))
print(temp.head(num_recommendations))
#Enter 'userID' and 'num_recommendations' for the user #
userID = 121
num_recommendations = 5
recommend_items(userID, pivot_df, preds_df, num_recommendations)
Below are the recommended items for user(user_id = 121):
user_ratings user_predictions
Recommended Items
B000LRMS66 0.0 0.543927
B002WE4HE2 0.0 0.423175
B000KO0GY6 0.0 0.416801
B001XURP7W 0.0 0.356788
B005HMKKH4 0.0 0.352138
Step 6. Evaluate both the models¶
Evaluation of Model-based Collaborative Filtering (SVD) by calculating the Root Mean Square Error (RMSE) between the actual ratings and the predicted ratings for all items.¶
# Actual ratings given by the users
final_ratings_matrix.head()
| prod_id | 0594451647 | 0594481813 | 0970407998 | 0972683275 | 1400501466 | 1400501520 | 1400501776 | 1400532620 | 1400532655 | 140053271X | ... | B00L5YZCCG | B00L8I6SFY | B00L8QCVL6 | B00LA6T0LS | B00LBZ1Z7K | B00LED02VY | B00LGN7Y3G | B00LGQ6HL8 | B00LI4ZZO8 | B00LKG1MC8 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| user_id | |||||||||||||||||||||
| A100UD67AHFODS | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | ... | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 |
| A100WO06OQR8BQ | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | ... | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 |
| A105S56ODHGJEK | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | ... | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 |
| A105TOJ6LTVMBG | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | ... | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 |
| A10AFVU66A79Y1 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | ... | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 |
5 rows × 48190 columns
# Average ACTUAL rating for each item
final_ratings_matrix.mean().head()
| 0 | |
|---|---|
| prod_id | |
| 0594451647 | 0.003247 |
| 0594481813 | 0.001948 |
| 0970407998 | 0.003247 |
| 0972683275 | 0.012338 |
| 1400501466 | 0.012987 |
# Predicted ratings
preds_df.head()
| prod_id | 0594451647 | 0594481813 | 0970407998 | 0972683275 | 1400501466 | 1400501520 | 1400501776 | 1400532620 | 1400532655 | 140053271X | ... | B00L5YZCCG | B00L8I6SFY | B00L8QCVL6 | B00LA6T0LS | B00LBZ1Z7K | B00LED02VY | B00LGN7Y3G | B00LGQ6HL8 | B00LI4ZZO8 | B00LKG1MC8 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| user_index | |||||||||||||||||||||
| 0 | 0.005086 | 0.002178 | 0.003668 | -0.040843 | 0.009640 | 0.006808 | 0.020659 | 0.000649 | 0.020331 | 0.005633 | ... | 0.000238 | -0.061477 | 0.001214 | -0.123433 | 0.028490 | 0.016109 | 0.002855 | -0.174568 | 0.011367 | -0.012997 |
| 1 | 0.002286 | -0.010898 | -0.000724 | 0.130259 | 0.007506 | -0.003350 | 0.063711 | -0.000674 | 0.016111 | -0.002433 | ... | -0.000038 | 0.013766 | 0.001473 | 0.025588 | -0.042103 | 0.004251 | 0.002177 | -0.024362 | -0.014765 | 0.038570 |
| 2 | -0.001655 | -0.002675 | -0.007355 | 0.007264 | 0.005152 | -0.003986 | -0.003480 | 0.006961 | -0.006606 | -0.002719 | ... | -0.001708 | -0.051040 | 0.000325 | -0.054867 | 0.017870 | -0.004996 | -0.002426 | 0.083928 | -0.112205 | 0.005964 |
| 3 | 0.001856 | 0.011019 | -0.005910 | -0.014134 | 0.000179 | 0.001877 | -0.005391 | -0.001709 | 0.004968 | 0.001402 | ... | 0.000582 | -0.009326 | -0.000465 | -0.048315 | 0.023302 | 0.006790 | 0.003380 | 0.005460 | -0.015263 | -0.025996 |
| 4 | 0.001115 | -0.002670 | 0.011018 | 0.014434 | 0.010319 | 0.006002 | 0.017151 | 0.003726 | 0.001404 | 0.005645 | ... | 0.000207 | 0.023761 | 0.000747 | -0.019347 | -0.012749 | 0.001026 | 0.001364 | -0.020580 | 0.011828 | 0.012770 |
5 rows × 48190 columns
# Average PREDICTED rating for each item
preds_df.mean().head()
| 0 | |
|---|---|
| prod_id | |
| 0594451647 | 0.001953 |
| 0594481813 | 0.002875 |
| 0970407998 | 0.003355 |
| 0972683275 | 0.010343 |
| 1400501466 | 0.004871 |
# Create a DataFrame rmse_df to store both the average actual ratings and the average predicted ratings for each item in a single table.
rmse_df = pd.concat([final_ratings_matrix.mean(), preds_df.mean()], axis=1)
rmse_df.columns = ['Avg_actual_ratings', 'Avg_predicted_ratings']
print(rmse_df.shape)
rmse_df['item_index'] = np.arange(0, rmse_df.shape[0], 1)
rmse_df.head()
(48190, 2)
| Avg_actual_ratings | Avg_predicted_ratings | item_index | |
|---|---|---|---|
| prod_id | |||
| 0594451647 | 0.003247 | 0.001953 | 0 |
| 0594481813 | 0.001948 | 0.002875 | 1 |
| 0970407998 | 0.003247 | 0.003355 | 2 |
| 0972683275 | 0.012338 | 0.010343 | 3 |
| 1400501466 | 0.012987 | 0.004871 | 4 |
RMSE = round((((rmse_df.Avg_actual_ratings - rmse_df.Avg_predicted_ratings) ** 2).mean() ** 0.5), 5)
print('\nRMSE SVD Model = {} \n'.format(RMSE))
RMSE SVD Model = 0.00275
The RMSE = 0.00275 indicates that the predicted ratings are very close to the actual ratings on average.
Step 7. Get the top 5 recommendations based on the user product interactions¶
# Enter 'userID' and 'num_recommendations' for the user #
userID = 15
num_recommendations = 5
recommend_items(userID, pivot_df, preds_df, num_recommendations)
Below are the recommended items for user(user_id = 15):
user_ratings user_predictions
Recommended Items
B007WTAJTO 0.0 0.335415
B000QUUFRW 0.0 0.281862
B002WE6D44 0.0 0.228786
B00004ZCJE 0.0 0.193438
B001XURP7W 0.0 0.170882
# Enter 'userID' and 'num_recommendations' for the user #
userID = 121
num_recommendations = 5
recommend_items(userID, pivot_df, preds_df, num_recommendations)
Below are the recommended items for user(user_id = 121):
user_ratings user_predictions
Recommended Items
B000LRMS66 0.0 0.543927
B002WE4HE2 0.0 0.423175
B000KO0GY6 0.0 0.416801
B001XURP7W 0.0 0.356788
B005HMKKH4 0.0 0.352138
# Enter 'userID' and 'num_recommendations' for the user #
userID = 200
num_recommendations = 5
recommend_items(userID, pivot_df, preds_df, num_recommendations)
Below are the recommended items for user(user_id = 200):
user_ratings user_predictions
Recommended Items
B008X9Z8NE 0.0 1.141688
B0079UAT0A 0.0 1.101302
B008X9Z528 0.0 1.072300
B004CLYEFK 0.0 1.025798
B008X9Z7N0 0.0 1.002051
In the above, we can see that the recommendations are different based on the given user.
Summary¶
The Popularity-based recommender system is non-personalised and the recommendations are based on purely on frequecy counts, which may be not suitable for a given user.
Model-based Collaborative Filtering is a personalised recommender system, the recommendations are based on the past behavior of the user and it is not dependent on any additional information.
You can see the differance for the sample user id set of 15, 121, and 200. The Popularity based model recommended the same set of 5 products for all 3 users, but the Collaborative Filtering based model recommendeded a different list for each user based on the user's past interaction history.