Recommendation System Project Overview¶

Context¶

Online E-commerce websites like Amazon, Flipkart uses different recommendation models to provide different suggestions to different users. Amazon currently uses item-to-item collaborative filtering, which scales to massive data sets and produces high-quality recommendations in real-time.

Objective¶

Build a recommendation system to recommend products to customers based on their previous ratings for other products. Apply the concepts and techniques you have learned in the previous weeks and summarise your insights at the end.

Dataset¶

We are using the Electronics dataset from the Amazon Reviews data repository (http://jmcauley.ucsd.edu/data/amazon/), which has several datasets.

Attribute Information

  • userId: Every user identified with a unique id
  • productId: Every product identified with a unique id
  • Rating: Rating of the corresponding product by the corresponding user
  • timestamp: Time of the rating ( ignore this column for this exercise)

Import Required Libraries¶

In [3]:
#install library joblib
!pip install joblib
Requirement already satisfied: joblib in /usr/local/lib/python3.11/dist-packages (1.4.2)
In [45]:
import numpy as np
import pandas as pd
import math
import json
import time
import matplotlib.pyplot as plt
import seaborn as sns
from sklearn.metrics.pairwise import cosine_similarity
from sklearn.model_selection import train_test_split
from sklearn.neighbors import NearestNeighbors
import joblib
import scipy.sparse
from scipy.sparse import csr_matrix
import warnings; warnings.simplefilter('ignore')
from scipy.sparse.linalg import svds
%matplotlib inline

Data Import¶

Part 1. Read and explore the dataset. ( Rename column, plot histograms, find data characteristics)¶

In [1]:
# uncomment and run the following line if using Google Colab
from google.colab import drive
drive.mount('/content/drive')
Mounted at /content/drive
In [5]:
#Import the data set
df = pd.read_csv('/content/drive/MyDrive/Personal/UT Austin/Recommendation Systems/Project/ratings_Electronics.csv', header=None)
df.columns = ['user_id', 'prod_id', 'rating', 'prod_name']
df = df.drop('prod_name', axis=1)
df_copy = df.copy(deep=True)
In [6]:
# see few rows of the imported dataset
df.tail()
Out[6]:
user_id prod_id rating
7824477 A2YZI3C9MOHC0L BT008UKTMW 5.0
7824478 A322MDK0M89RHN BT008UKTMW 5.0
7824479 A1MH90R0ADMIK0 BT008UKTMW 4.0
7824480 A10M2KEFPEQDHN BT008UKTMW 4.0
7824481 A2G81TMIOIDEQQ BT008V9J9U 5.0
In [7]:
# Check the number of rows and columns
rows, columns = df.shape
print("No of rows: ", rows)
print("No of columns: ", columns)
No of rows:  7824482
No of columns:  3

We have ~7.8M rows of data.

In [ ]:
#Check Data types
df.dtypes
Out[ ]:
user_id     object
prod_id     object
rating     float64
dtype: object
In [8]:
# Check for missing values present
print('Number of missing values across columns-\n', df.isnull().sum())
Number of missing values across columns-
 user_id    0
prod_id    0
rating     0
dtype: int64

There are no missing values.

In [9]:
# Summary statistics of 'rating' variable
df[['rating']].describe().transpose()
Out[9]:
count mean std min 25% 50% 75% max
rating 7824482.0 4.012337 1.38091 1.0 3.0 5.0 5.0 5.0
In [10]:
# find minimum and maximum ratings

def find_min_max_rating():
    print('The minimum rating is: %d' %(df['rating'].min()))
    print('The maximum rating is: %d' %(df['rating'].max()))

find_min_max_rating()
The minimum rating is: 1
The maximum rating is: 5

Ratings are on scale of 1 - 5

In [13]:
# function to create labeled barplots

def labeled_barplot(data, feature, perc=False, n=None):
    """
    Barplot with percentage at the top

    data: dataframe
    feature: dataframe column
    perc: whether to display percentages instead of count (default is False)
    n: displays the top n category levels (default is None, i.e., display all levels)
    """

    total = len(data[feature])  # length of the column
    count = data[feature].nunique()
    if n is None:
        plt.figure(figsize=(count + 1, 5))
    else:
        plt.figure(figsize=(n + 1, 5))

    plt.xticks(rotation=90, fontsize=15)
    ax = sns.countplot(
        data=data,
        x=feature,
        hue=feature,
        palette="Paired",
        order=data[feature].value_counts().index[:n].sort_values(),
    )

    for p in ax.patches:
        if perc == True:
            label = "{:.1f}%".format(
                100 * p.get_height() / total
            )  # percentage of each class of the category
        else:
            label = p.get_height()  # count of each level of the category

        x = p.get_x() + p.get_width() / 2  # width of the plot
        y = p.get_height()  # height of the plot

        ax.annotate(
            label,
            (x, y),
            ha="center",
            va="center",
            size=12,
            xytext=(0, 5),
            textcoords="offset points",
        )  # annotate the percentage

    plt.show()  # show the plot
In [14]:
labeled_barplot(df, "rating", perc=True)
No description has been provided for this image
In [16]:
# Number of unique user id and product id in the data
print('Number of unique USERS in Raw data = {:,}'.format(df['user_id'].nunique()))
print('Number of unique ITEMS in Raw data = {:,}'.format(df['prod_id'].nunique()))
Number of unique USERS in Raw data = 4,201,696
Number of unique ITEMS in Raw data = 476,002
In [21]:
user_counts = df.groupby('user_id').size().reset_index(name='row_count')
In [24]:
# Plot histogram with log scale on the y-axis
plt.figure(figsize=(10, 6))
sns.histplot(user_counts['row_count'], bins=30, kde=False)
plt.yscale('log')  # Set y-axis to logarithmic scale
plt.xlabel('Number of Rows per User', fontsize=12)
plt.ylabel('Number of Users (Log Scale)', fontsize=12)
plt.title('Distribution of Rows per User', fontsize=14)
plt.grid(axis='y')
plt.show()
No description has been provided for this image

Here we can see there area huge number of users, who have given very few reviews.

Part 2. Take subset of dataset to make it less sparse/more dense. (Keep only users who gave 50 or more ratings)¶

In [25]:
# Top 10 users based on rating
most_rated = df.groupby('user_id').size().sort_values(ascending=False)[:10]
most_rated
Out[25]:
0
user_id
A5JLAU2ARJ0BO 520
ADLVFFE4VBT8 501
A3OXHLG6DIBRW8 498
A6FIAB28IS79 431
A680RUE1FDO8B 406
A1ODOGXEYECQQ8 380
A36K2N527TXXJN 314
A2AY4YUOX2N1BQ 311
AWPODHOB4GFWL 308
A25C2M3QF9G7OQ 296

Data model preparation as per requirement on number of minimum ratings¶

In [26]:
counts = df['user_id'].value_counts()
df_final = df[df['user_id'].isin(counts[counts >= 50].index)]
In [27]:
print('Number of users who have rated 50 or more items = {:,}'.format(len(df_final)))
print('Number of unique USERS in final data = {:,}'.format(df_final['user_id'].nunique()))
print('Number of unique ITEMS in final data = {:,}'.format(df_final['prod_id'].nunique()))
Number of users who have rated 50 or more items = 125,871
Number of unique USERS in final data = 1,540
Number of unique ITEMS in final data = 48,190

df_final has users who have rated 50 or more items¶

In [31]:
df_final.head()
Out[31]:
user_id prod_id rating
94 A3BY5KCNQZXV5U 0594451647 5.0
118 AT09WGFUM934H 0594481813 3.0
177 A32HSNCNPRUMTR 0970407998 1.0
178 A17HMM1M7T9PJ1 0970407998 4.0
492 A3CLWR1UUZT6TG 0972683275 5.0

Calculate the density of the rating matrix¶

In [30]:
final_ratings_matrix = df_final.pivot(index = 'user_id', columns ='prod_id', values = 'rating').fillna(0)
print('Shape of final_ratings_matrix: ', final_ratings_matrix.shape)

given_num_of_ratings = np.count_nonzero(final_ratings_matrix)
print('given_num_of_ratings =  {:,}'.format(given_num_of_ratings))
possible_num_of_ratings = final_ratings_matrix.shape[0] * final_ratings_matrix.shape[1]
print('possible_num_of_ratings = {:,}'.format(possible_num_of_ratings))
density = (given_num_of_ratings/possible_num_of_ratings)
density *= 100
print ('density: {:4.2f}%'.format(density))
Shape of final_ratings_matrix:  (1540, 48190)
given_num_of_ratings =  125,871
possible_num_of_ratings = 74,212,600
density: 0.17%

Only 0.17% of the cells are populated with a rating value, so the matrix is very sparse!

In [32]:
final_ratings_matrix.tail()
Out[32]:
prod_id 0594451647 0594481813 0970407998 0972683275 1400501466 1400501520 1400501776 1400532620 1400532655 140053271X ... B00L5YZCCG B00L8I6SFY B00L8QCVL6 B00LA6T0LS B00LBZ1Z7K B00LED02VY B00LGN7Y3G B00LGQ6HL8 B00LI4ZZO8 B00LKG1MC8
user_id
AZBXKUH4AIW3X 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 ... 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0
AZCE11PSTCH1L 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 ... 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0
AZMY6E8B52L2T 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 ... 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0
AZNUHQSHZHSUE 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 ... 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0
AZOK5STV85FBJ 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 ... 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0

5 rows × 48190 columns

In [33]:
# Matrix with one row per 'Product' and one column per 'user' for Item-based CF
final_ratings_matrix_T = final_ratings_matrix.transpose()
final_ratings_matrix_T.head()
Out[33]:
user_id A100UD67AHFODS A100WO06OQR8BQ A105S56ODHGJEK A105TOJ6LTVMBG A10AFVU66A79Y1 A10H24TDLK2VDP A10NMELR4KX0J6 A10O7THJ2O20AG A10PEXB6XAQ5XF A10X9ME6R66JDX ... AYOTEJ617O60K AYP0YPLSP9ISM AZ515FFZ7I2P7 AZ8XSDMIX04VJ AZAC8O310IK4E AZBXKUH4AIW3X AZCE11PSTCH1L AZMY6E8B52L2T AZNUHQSHZHSUE AZOK5STV85FBJ
prod_id
0594451647 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 ... 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0
0594481813 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 ... 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0
0970407998 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 ... 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0
0972683275 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 ... 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0
1400501466 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 ... 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0

5 rows × 1540 columns

Part 3. Split the data randomly into train and test dataset. ( For example split it in 70/30 ratio)¶

In [34]:
#Split the training and test data in the ratio 70:30
train_data, test_data = train_test_split(df_final, test_size = 0.3, random_state=0)

print(train_data.head(5))
                user_id     prod_id  rating
6595853  A2BYV7S1QP2YIG  B009EAHVTA     5.0
4738241   AB094YABX21WQ  B0056XCEAA     1.0
4175596  A3D0UM4ZD2CMAW  B004I763AW     5.0
3753016   AATWFX0ZZSE6C  B0040NPHMO     3.0
1734767  A1NNMOD9H36Q8E  B0015VW3BM     4.0
In [35]:
def shape():
    print("Test data shape: ", test_data.shape)
    print("Train data shape: ", train_data.shape)
shape()
Test data shape:  (37762, 3)
Train data shape:  (88109, 3)

Part 4. Build Popularity Recommender model. (Non-personalised)¶

In [36]:
#Count of user_id for each unique product as recommendation score
#The numer of users that interacted with a given product is taken as a proxy for its popularity

train_data_grouped = train_data.groupby('prod_id').agg({'user_id': 'count'}).reset_index()
train_data_grouped.rename(columns = {'user_id': 'score'},inplace=True)
train_data_grouped.head()
Out[36]:
prod_id score
0 0594451647 1
1 0594481813 1
2 0970407998 1
3 0972683275 3
4 1400501466 4

Note: We haven't taken into account at all the rating the user gave, so another method could be to consider not only number of users that interacted with the product, but also the rating they gave. For example we could calculate the avg rating for a product and multiply that by the number of users for a ratings weighted popularity score.

In [37]:
#Sort the products in descending order of score (highest score first). If two or more products have the same score,
#they are further sorted by prod_id in ascending order.
train_data_sort = train_data_grouped.sort_values(['score', 'prod_id'], ascending = [0,1])

#Generate a recommendation rank based upon score
#ascending=0: Higher scores get a lower rank (i.e., the highest score gets rank 1).
train_data_sort['Rank'] = train_data_sort['score'].rank(ascending=0, method='first')

#Get the top 5 recommendations
popularity_recommendations = train_data_sort.head(5)
popularity_recommendations
Out[37]:
prod_id score Rank
30847 B0088CJT4U 133 1.0
30287 B007WTAJTO 124 2.0
19647 B003ES5ZUU 122 3.0
8752 B000N99BBC 114 4.0
30555 B00829THK0 97 5.0
In [38]:
# Use popularity based recommender model to make predictions
def recommend(user_id):
    user_recommendations = popularity_recommendations

    #Add user_id column for which the recommendations are being generated
    user_recommendations['user_id'] = user_id

    #Bring user_id column to the front
    cols = user_recommendations.columns.tolist()
    cols = cols[-1:] + cols[:-1]
    user_recommendations = user_recommendations[cols]

    return user_recommendations
In [39]:
find_recom = [15,121,200]   # This list is user choice.
for i in find_recom:
    print("Here is the recommendation for the userId: %d\n" %(i))
    print(recommend(i))
    print("\n")
Here is the recommendation for the userId: 15

       user_id     prod_id  score  Rank
30847       15  B0088CJT4U    133   1.0
30287       15  B007WTAJTO    124   2.0
19647       15  B003ES5ZUU    122   3.0
8752        15  B000N99BBC    114   4.0
30555       15  B00829THK0     97   5.0


Here is the recommendation for the userId: 121

       user_id     prod_id  score  Rank
30847      121  B0088CJT4U    133   1.0
30287      121  B007WTAJTO    124   2.0
19647      121  B003ES5ZUU    122   3.0
8752       121  B000N99BBC    114   4.0
30555      121  B00829THK0     97   5.0


Here is the recommendation for the userId: 200

       user_id     prod_id  score  Rank
30847      200  B0088CJT4U    133   1.0
30287      200  B007WTAJTO    124   2.0
19647      200  B003ES5ZUU    122   3.0
8752       200  B000N99BBC    114   4.0
30555      200  B00829THK0     97   5.0


Since this is a popularity-based recommender model, recommendations remain the same for all users. We predict the products based on the popularity. It is not personalized to particular user

Step 5. Build Collaborative Filtering model.¶

A collaborative filtering model is a recommendation system technique that predicts a user's preferences for items (products in this case) based on their past behavior and the preferences of similar users. It operates under the assumption that users who agreed on certain items in the past are likely to agree on others in the future. Collaborative filtering can be user-based, focusing on similarities between users, or item-based, focusing on similarities between items. Unlike content-based methods, it doesn't require item features, relying instead on user-item interaction data (e.g., ratings or purchases). This makes it effective but also sensitive to sparse data or cold-start problems.

Model-based Collaborative Filtering: Singular Value Decomposition (SVD)¶

In [40]:
#Combine Training and Test Data
df_CF = pd.concat([train_data, test_data]).reset_index()
df_CF.tail()
Out[40]:
index user_id prod_id rating
125866 621872 A3OXHLG6DIBRW8 B0007UQNOA 3.0
125867 1942808 A365PBEOWM7EI7 B001DVZXC0 3.0
125868 5219963 A3QDY9I0CNMD2W B005WXQO3W 5.0
125869 876608 AR18DH5SL9F73 B000EPR7AC 5.0
125870 975289 A3VL4RXCWNSR3H B000GM7MRG 5.0
In [41]:
#User-based Collaborative Filtering
# Matrix with row per 'user' and column per 'item'
pivot_df = df_CF.pivot(index = 'user_id', columns ='prod_id', values = 'rating').fillna(0)
print(pivot_df.shape)
pivot_df.head()
(1540, 48190)
Out[41]:
prod_id 0594451647 0594481813 0970407998 0972683275 1400501466 1400501520 1400501776 1400532620 1400532655 140053271X ... B00L5YZCCG B00L8I6SFY B00L8QCVL6 B00LA6T0LS B00LBZ1Z7K B00LED02VY B00LGN7Y3G B00LGQ6HL8 B00LI4ZZO8 B00LKG1MC8
user_id
A100UD67AHFODS 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 ... 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0
A100WO06OQR8BQ 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 ... 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0
A105S56ODHGJEK 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 ... 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0
A105TOJ6LTVMBG 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 ... 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0
A10AFVU66A79Y1 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 ... 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0

5 rows × 48190 columns

In [42]:
#Adds a new column user_index that assigns a unique numerical index to each user_id.
pivot_df['user_index'] = np.arange(0, pivot_df.shape[0], 1)
pivot_df.head()
Out[42]:
prod_id 0594451647 0594481813 0970407998 0972683275 1400501466 1400501520 1400501776 1400532620 1400532655 140053271X ... B00L8I6SFY B00L8QCVL6 B00LA6T0LS B00LBZ1Z7K B00LED02VY B00LGN7Y3G B00LGQ6HL8 B00LI4ZZO8 B00LKG1MC8 user_index
user_id
A100UD67AHFODS 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 ... 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0
A100WO06OQR8BQ 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 ... 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 1
A105S56ODHGJEK 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 ... 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 2
A105TOJ6LTVMBG 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 ... 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 3
A10AFVU66A79Y1 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 ... 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 4

5 rows × 48191 columns

In [43]:
#Replaces the original user_id index with the user_index column to make numerical indexing easier for subsequent modeling
pivot_df.set_index(['user_index'], inplace=True)

# Actual ratings given by users
pivot_df.head()
Out[43]:
prod_id 0594451647 0594481813 0970407998 0972683275 1400501466 1400501520 1400501776 1400532620 1400532655 140053271X ... B00L5YZCCG B00L8I6SFY B00L8QCVL6 B00LA6T0LS B00LBZ1Z7K B00LED02VY B00LGN7Y3G B00LGQ6HL8 B00LI4ZZO8 B00LKG1MC8
user_index
0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 ... 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0
1 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 ... 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0
2 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 ... 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0
3 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 ... 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0
4 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 ... 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0

5 rows × 48190 columns

In [46]:
# Convert pivot_df to a NumPy array
pivot_matrix = pivot_df.values
In [48]:
# Singular Value Decomposition
# The parameter k specifies the number of singular values and vectors to compute. It reduces the dimensionality of the matrix.
# In other words, k represents the number of latent (hidden) features used to represent each user in the data, which then allows
# identifying users that are similar to each other
# U Represents users in the latent feature space. Each row corresponds to a user, with k-dimensional features capturing their preferences.
# sigma represents the importance of each latent feature
# Vt represents the products in the latent feature space, each row corresponds to an item, with k-dimensional features capturing its characteristics
U, sigma, Vt = svds(pivot_matrix, k=50)
# Construct diagonal array in SVD
sigma = np.diag(sigma)

Perform matrix multiplication of the three components from SVD.

U: The user matrix, where rows represent users in the latent feature space.

sigma: The diagonal matrix of singular values, capturing the importance of each latent feature.

Vt: The item matrix, where columns represent items in the latent feature space.

The result, all_user_predicted_ratings, is an approximate reconstruction of the original user-item matrix (pivot_df), with predicted ratings replacing the original values (including zeros where ratings were missing).

In [52]:
# Reconstruct the predicted ratings matrix
all_user_predicted_ratings = np.dot(np.dot(U, sigma), Vt)

# Convert the predictions to a DataFrame, preserving user_id as the index
preds_df = pd.DataFrame(all_user_predicted_ratings, index=pivot_df.index, columns=pivot_df.columns)

# Display the first few rows
preds_df.head()
Out[52]:
prod_id 0594451647 0594481813 0970407998 0972683275 1400501466 1400501520 1400501776 1400532620 1400532655 140053271X ... B00L5YZCCG B00L8I6SFY B00L8QCVL6 B00LA6T0LS B00LBZ1Z7K B00LED02VY B00LGN7Y3G B00LGQ6HL8 B00LI4ZZO8 B00LKG1MC8
user_index
0 0.005086 0.002178 0.003668 -0.040843 0.009640 0.006808 0.020659 0.000649 0.020331 0.005633 ... 0.000238 -0.061477 0.001214 -0.123433 0.028490 0.016109 0.002855 -0.174568 0.011367 -0.012997
1 0.002286 -0.010898 -0.000724 0.130259 0.007506 -0.003350 0.063711 -0.000674 0.016111 -0.002433 ... -0.000038 0.013766 0.001473 0.025588 -0.042103 0.004251 0.002177 -0.024362 -0.014765 0.038570
2 -0.001655 -0.002675 -0.007355 0.007264 0.005152 -0.003986 -0.003480 0.006961 -0.006606 -0.002719 ... -0.001708 -0.051040 0.000325 -0.054867 0.017870 -0.004996 -0.002426 0.083928 -0.112205 0.005964
3 0.001856 0.011019 -0.005910 -0.014134 0.000179 0.001877 -0.005391 -0.001709 0.004968 0.001402 ... 0.000582 -0.009326 -0.000465 -0.048315 0.023302 0.006790 0.003380 0.005460 -0.015263 -0.025996
4 0.001115 -0.002670 0.011018 0.014434 0.010319 0.006002 0.017151 0.003726 0.001404 0.005645 ... 0.000207 0.023761 0.000747 -0.019347 -0.012749 0.001026 0.001364 -0.020580 0.011828 0.012770

5 rows × 48190 columns

In [53]:
# Recommend the items with the highest predicted ratings

def recommend_items(userID, pivot_df, preds_df, num_recommendations):

    user_idx = userID-1 # index starts at 0

    # Get and sort the user's ratings
    sorted_user_ratings = pivot_df.iloc[user_idx].sort_values(ascending=False)
    #sorted_user_ratings
    sorted_user_predictions = preds_df.iloc[user_idx].sort_values(ascending=False)
    #sorted_user_predictions

    temp = pd.concat([sorted_user_ratings, sorted_user_predictions], axis=1)
    temp.index.name = 'Recommended Items'
    temp.columns = ['user_ratings', 'user_predictions']

    temp = temp.loc[temp.user_ratings == 0]
    temp = temp.sort_values('user_predictions', ascending=False)
    print('\nBelow are the recommended items for user(user_id = {}):\n'.format(userID))
    print(temp.head(num_recommendations))
In [54]:
#Enter 'userID' and 'num_recommendations' for the user #
userID = 121
num_recommendations = 5
recommend_items(userID, pivot_df, preds_df, num_recommendations)
Below are the recommended items for user(user_id = 121):

                   user_ratings  user_predictions
Recommended Items                                
B000LRMS66                  0.0          0.543927
B002WE4HE2                  0.0          0.423175
B000KO0GY6                  0.0          0.416801
B001XURP7W                  0.0          0.356788
B005HMKKH4                  0.0          0.352138

Step 6. Evaluate both the models¶

Evaluation of Model-based Collaborative Filtering (SVD) by calculating the Root Mean Square Error (RMSE) between the actual ratings and the predicted ratings for all items.¶

In [55]:
# Actual ratings given by the users
final_ratings_matrix.head()
Out[55]:
prod_id 0594451647 0594481813 0970407998 0972683275 1400501466 1400501520 1400501776 1400532620 1400532655 140053271X ... B00L5YZCCG B00L8I6SFY B00L8QCVL6 B00LA6T0LS B00LBZ1Z7K B00LED02VY B00LGN7Y3G B00LGQ6HL8 B00LI4ZZO8 B00LKG1MC8
user_id
A100UD67AHFODS 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 ... 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0
A100WO06OQR8BQ 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 ... 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0
A105S56ODHGJEK 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 ... 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0
A105TOJ6LTVMBG 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 ... 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0
A10AFVU66A79Y1 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 ... 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0 0.0

5 rows × 48190 columns

In [56]:
# Average ACTUAL rating for each item
final_ratings_matrix.mean().head()
Out[56]:
0
prod_id
0594451647 0.003247
0594481813 0.001948
0970407998 0.003247
0972683275 0.012338
1400501466 0.012987

In [57]:
# Predicted ratings
preds_df.head()
Out[57]:
prod_id 0594451647 0594481813 0970407998 0972683275 1400501466 1400501520 1400501776 1400532620 1400532655 140053271X ... B00L5YZCCG B00L8I6SFY B00L8QCVL6 B00LA6T0LS B00LBZ1Z7K B00LED02VY B00LGN7Y3G B00LGQ6HL8 B00LI4ZZO8 B00LKG1MC8
user_index
0 0.005086 0.002178 0.003668 -0.040843 0.009640 0.006808 0.020659 0.000649 0.020331 0.005633 ... 0.000238 -0.061477 0.001214 -0.123433 0.028490 0.016109 0.002855 -0.174568 0.011367 -0.012997
1 0.002286 -0.010898 -0.000724 0.130259 0.007506 -0.003350 0.063711 -0.000674 0.016111 -0.002433 ... -0.000038 0.013766 0.001473 0.025588 -0.042103 0.004251 0.002177 -0.024362 -0.014765 0.038570
2 -0.001655 -0.002675 -0.007355 0.007264 0.005152 -0.003986 -0.003480 0.006961 -0.006606 -0.002719 ... -0.001708 -0.051040 0.000325 -0.054867 0.017870 -0.004996 -0.002426 0.083928 -0.112205 0.005964
3 0.001856 0.011019 -0.005910 -0.014134 0.000179 0.001877 -0.005391 -0.001709 0.004968 0.001402 ... 0.000582 -0.009326 -0.000465 -0.048315 0.023302 0.006790 0.003380 0.005460 -0.015263 -0.025996
4 0.001115 -0.002670 0.011018 0.014434 0.010319 0.006002 0.017151 0.003726 0.001404 0.005645 ... 0.000207 0.023761 0.000747 -0.019347 -0.012749 0.001026 0.001364 -0.020580 0.011828 0.012770

5 rows × 48190 columns

In [58]:
# Average PREDICTED rating for each item
preds_df.mean().head()
Out[58]:
0
prod_id
0594451647 0.001953
0594481813 0.002875
0970407998 0.003355
0972683275 0.010343
1400501466 0.004871

In [61]:
# Create a DataFrame rmse_df to store both the average actual ratings and the average predicted ratings for each item in a single table.
rmse_df = pd.concat([final_ratings_matrix.mean(), preds_df.mean()], axis=1)
rmse_df.columns = ['Avg_actual_ratings', 'Avg_predicted_ratings']
print(rmse_df.shape)
rmse_df['item_index'] = np.arange(0, rmse_df.shape[0], 1)
rmse_df.head()
(48190, 2)
Out[61]:
Avg_actual_ratings Avg_predicted_ratings item_index
prod_id
0594451647 0.003247 0.001953 0
0594481813 0.001948 0.002875 1
0970407998 0.003247 0.003355 2
0972683275 0.012338 0.010343 3
1400501466 0.012987 0.004871 4
In [ ]:
 
In [60]:
RMSE = round((((rmse_df.Avg_actual_ratings - rmse_df.Avg_predicted_ratings) ** 2).mean() ** 0.5), 5)
print('\nRMSE SVD Model = {} \n'.format(RMSE))
RMSE SVD Model = 0.00275 

The RMSE = 0.00275 indicates that the predicted ratings are very close to the actual ratings on average.

Step 7. Get the top 5 recommendations based on the user product interactions¶

In [64]:
# Enter 'userID' and 'num_recommendations' for the user #
userID = 15
num_recommendations = 5
recommend_items(userID, pivot_df, preds_df, num_recommendations)
Below are the recommended items for user(user_id = 15):

                   user_ratings  user_predictions
Recommended Items                                
B007WTAJTO                  0.0          0.335415
B000QUUFRW                  0.0          0.281862
B002WE6D44                  0.0          0.228786
B00004ZCJE                  0.0          0.193438
B001XURP7W                  0.0          0.170882
In [63]:
# Enter 'userID' and 'num_recommendations' for the user #
userID = 121
num_recommendations = 5
recommend_items(userID, pivot_df, preds_df, num_recommendations)
Below are the recommended items for user(user_id = 121):

                   user_ratings  user_predictions
Recommended Items                                
B000LRMS66                  0.0          0.543927
B002WE4HE2                  0.0          0.423175
B000KO0GY6                  0.0          0.416801
B001XURP7W                  0.0          0.356788
B005HMKKH4                  0.0          0.352138
In [62]:
# Enter 'userID' and 'num_recommendations' for the user #
userID = 200
num_recommendations = 5
recommend_items(userID, pivot_df, preds_df, num_recommendations)
Below are the recommended items for user(user_id = 200):

                   user_ratings  user_predictions
Recommended Items                                
B008X9Z8NE                  0.0          1.141688
B0079UAT0A                  0.0          1.101302
B008X9Z528                  0.0          1.072300
B004CLYEFK                  0.0          1.025798
B008X9Z7N0                  0.0          1.002051

In the above, we can see that the recommendations are different based on the given user.

Summary¶

The Popularity-based recommender system is non-personalised and the recommendations are based on purely on frequecy counts, which may be not suitable for a given user.

Model-based Collaborative Filtering is a personalised recommender system, the recommendations are based on the past behavior of the user and it is not dependent on any additional information.

You can see the differance for the sample user id set of 15, 121, and 200. The Popularity based model recommended the same set of 5 products for all 3 users, but the Collaborative Filtering based model recommendeded a different list for each user based on the user's past interaction history.