Problem Statement¶
Business Context¶
The prices of the stocks of companies listed under a global exchange are influenced by a variety of factors, with the company's financial performance, innovations and collaborations, and market sentiment being factors that play a significant role. News and media reports can rapidly affect investor perceptions and, consequently, stock prices in the highly competitive financial industry. With the sheer volume of news and opinions from a wide variety of sources, investors and financial analysts often struggle to stay updated and accurately interpret its impact on the market. As a result, investment firms need sophisticated tools to analyze market sentiment and integrate this information into their investment strategies.
Problem Definition¶
With an ever-rising number of news articles and opinions, an investment startup aims to leverage artificial intelligence to address the challenge of interpreting stock-related news and its impact on stock prices. They have collected historical daily news for a specific company listed under NASDAQ, along with data on its daily stock price and trade volumes.
As a member of the Data Science and AI team in the startup, you have been tasked with analyzing the data, developing an AI-driven sentiment analysis system that will automatically process and analyze news articles to gauge market sentiment, and summarizing the news at a weekly level to enhance the accuracy of their stock price predictions and optimize investment strategies. This will empower their financial analysts with actionable insights, leading to more informed investment decisions and improved client outcomes.
Data Dictionary¶
Date: The date the news was releasedNews: The content of news articles that could potentially affect the company's stock priceOpen: The stock price (in $) at the beginning of the dayHigh: The highest stock price (in $) reached during the dayLow: The lowest stock price (in $) reached during the dayClose: The adjusted stock price (in $) at the end of the dayVolume: The number of shares traded during the dayLabel: The sentiment polarity of the news content- 1: positive
- 0: neutral
- -1: negative
Installing and Importing Necessary Libraries¶
# installing the sentence-transformers library
!pip install -U sentence-transformers -q
# Import the library for doing candlestick charts
!pip install mplfinance
import mplfinance as mpf
Collecting mplfinance Downloading mplfinance-0.12.10b0-py3-none-any.whl.metadata (19 kB) Requirement already satisfied: matplotlib in /usr/local/lib/python3.10/dist-packages (from mplfinance) (3.8.0) Requirement already satisfied: pandas in /usr/local/lib/python3.10/dist-packages (from mplfinance) (2.2.2) Requirement already satisfied: contourpy>=1.0.1 in /usr/local/lib/python3.10/dist-packages (from matplotlib->mplfinance) (1.3.1) Requirement already satisfied: cycler>=0.10 in /usr/local/lib/python3.10/dist-packages (from matplotlib->mplfinance) (0.12.1) Requirement already satisfied: fonttools>=4.22.0 in /usr/local/lib/python3.10/dist-packages (from matplotlib->mplfinance) (4.55.3) Requirement already satisfied: kiwisolver>=1.0.1 in /usr/local/lib/python3.10/dist-packages (from matplotlib->mplfinance) (1.4.7) Requirement already satisfied: numpy<2,>=1.21 in /usr/local/lib/python3.10/dist-packages (from matplotlib->mplfinance) (1.26.4) Requirement already satisfied: packaging>=20.0 in /usr/local/lib/python3.10/dist-packages (from matplotlib->mplfinance) (24.2) Requirement already satisfied: pillow>=6.2.0 in /usr/local/lib/python3.10/dist-packages (from matplotlib->mplfinance) (11.0.0) Requirement already satisfied: pyparsing>=2.3.1 in /usr/local/lib/python3.10/dist-packages (from matplotlib->mplfinance) (3.2.0) Requirement already satisfied: python-dateutil>=2.7 in /usr/local/lib/python3.10/dist-packages (from matplotlib->mplfinance) (2.8.2) Requirement already satisfied: pytz>=2020.1 in /usr/local/lib/python3.10/dist-packages (from pandas->mplfinance) (2024.2) Requirement already satisfied: tzdata>=2022.7 in /usr/local/lib/python3.10/dist-packages (from pandas->mplfinance) (2024.2) Requirement already satisfied: six>=1.5 in /usr/local/lib/python3.10/dist-packages (from python-dateutil>=2.7->matplotlib->mplfinance) (1.17.0) Downloading mplfinance-0.12.10b0-py3-none-any.whl (75 kB) ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 75.0/75.0 kB 6.2 MB/s eta 0:00:00 Installing collected packages: mplfinance Successfully installed mplfinance-0.12.10b0
# to read and manipulate the data
import pandas as pd
import numpy as np
pd.set_option('max_colwidth', None) # setting column to the maximum column width as per the data
import os
from google.colab import userdata
import json
from imblearn.over_sampling import SMOTE
from sklearn.decomposition import PCA
# to visualise data
import matplotlib.pyplot as plt
import seaborn as sns
# to use regular expressions for manipulating text data
import re
# to manipulate string data
import string
# to load the natural language toolkit
import nltk
nltk.download('stopwords') # loading the stopwords
nltk.download('wordnet') # loading the wordnet module that is used in stemming
# Deep Learning library
import torch
# to load transformer models
from sentence_transformers import SentenceTransformer
from transformers import T5Tokenizer, T5ForConditionalGeneration, pipeline
# to remove common stop words
from nltk.corpus import stopwords
# to perform stemming
from nltk.stem.porter import PorterStemmer
# To tune the model
from sklearn.model_selection import GridSearchCV
# to import Word2Vec
from gensim.models import Word2Vec
# to create Bag of Words
from sklearn.feature_extraction.text import CountVectorizer
# to split data into train and test sets
from sklearn.model_selection import train_test_split
# to build a Random Forest model
from sklearn.ensemble import RandomForestClassifier
# To compute metrics to evaluate the model
from sklearn.metrics import confusion_matrix, accuracy_score, f1_score, precision_score, recall_score, classification_report
import warnings
warnings.filterwarnings("ignore")
[nltk_data] Downloading package stopwords to /root/nltk_data... [nltk_data] Unzipping corpora/stopwords.zip. [nltk_data] Downloading package wordnet to /root/nltk_data...
Loading the dataset¶
from google.colab import drive
drive.mount('/content/drive')
Mounted at /content/drive
# loading data into a pandas dataframe
df = pd.read_csv("/content/drive/MyDrive/Personal/UT Austin/Project 6 - NLP/stock_news.csv")
data = df.copy()
#Add a string readable version of the Labael
label_mapping = {-1: 'negative', 0: 'neutral', 1: 'positive'}
data['Label_String'] = data['Label'].map(label_mapping)
# Add a column for the character count of the 'News' column
data['News_Char_Count'] = data['News'].str.len()
# Convert the Date column to a datetime type
data['Date'] = pd.to_datetime(data['Date'])
#Calculate the Open - Close value
data['Open_Close'] = data['Open'] - data['Close']
#Calculate the High - Low value
data['High_Low'] = data['High'] - data['Low']
# Identify the set of numeric values
num_features = ['Open','High','Low','Close','Volume','News_Char_Count','Open_Close','High_Low']
# Create a summarized data frame where we are just interested in the stock related values
# Grouping by Date
data_unique = (
data.groupby('Date')
.agg({
'Open': 'first',
'High': 'first',
'Low': 'first',
'Close': 'first',
'Volume': 'first',
'Open_Close': 'first',
'High_Low': 'first',
'Label': 'sum', # Sum the values in Label column
})
.rename(columns={'Label': 'Label_Sum'}) # Rename Label to Label_Sum
.reset_index() # Reset index to make Date a column
)
# Add Row_Count column
data_unique['Row_Count'] = data.groupby('Date').size().values
Data Overview¶
Checking the first 5 rows¶
data.head(5)
| Date | News | Open | High | Low | Close | Volume | Label | Label_String | News_Char_Count | Open_Close | High_Low | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 0 | 2019-01-02 | The tech sector experienced a significant decline in the aftermarket following Apple's Q1 revenue warning. Notable suppliers, including Skyworks, Broadcom, Lumentum, Qorvo, and TSMC, saw their stocks drop in response to Apple's downward revision of its revenue expectations for the quarter, previously announced in January. | 41.74 | 42.24 | 41.48 | 40.25 | 130672400 | -1 | negative | 324 | 1.49 | 0.76 |
| 1 | 2019-01-02 | Apple lowered its fiscal Q1 revenue guidance to $84 billion from earlier estimates of $89-$93 billion due to weaker than expected iPhone sales. The announcement caused a significant drop in Apple's stock price and negatively impacted related suppliers, leading to broader market declines for tech indices such as Nasdaq 10 | 41.74 | 42.24 | 41.48 | 40.25 | 130672400 | -1 | negative | 323 | 1.49 | 0.76 |
| 2 | 2019-01-02 | Apple cut its fiscal first quarter revenue forecast from $89-$93 billion to $84 billion due to weaker demand in China and fewer iPhone upgrades. CEO Tim Cook also mentioned constrained sales of Airpods and Macbooks. Apple's shares fell 8.5% in post market trading, while Asian suppliers like Hon | 41.74 | 42.24 | 41.48 | 40.25 | 130672400 | -1 | negative | 296 | 1.49 | 0.76 |
| 3 | 2019-01-02 | This news article reports that yields on long-dated U.S. Treasury securities hit their lowest levels in nearly a year on January 2, 2019, due to concerns about the health of the global economy following weak economic data from China and Europe, as well as the partial U.S. government shutdown. Apple | 41.74 | 42.24 | 41.48 | 40.25 | 130672400 | -1 | negative | 300 | 1.49 | 0.76 |
| 4 | 2019-01-02 | Apple's revenue warning led to a decline in USD JPY pair and a gain in Japanese yen, as investors sought safety in the highly liquid currency. Apple's underperformance in Q1, with forecasted revenue of $84 billion compared to analyst expectations of $91.5 billion, triggered risk aversion mood in markets | 41.74 | 42.24 | 41.48 | 40.25 | 130672400 | -1 | negative | 305 | 1.49 | 0.76 |
data_unique.head(5)
| Date | Open | High | Low | Close | Volume | Open_Close | High_Low | Label_Sum | Row_Count | |
|---|---|---|---|---|---|---|---|---|---|---|
| 0 | 2019-01-02 | 41.74 | 42.24 | 41.48 | 40.25 | 130672400 | 1.49 | 0.76 | -9 | 14 |
| 1 | 2019-01-03 | 43.57 | 43.79 | 43.22 | 42.47 | 103544800 | 1.10 | 0.56 | -10 | 28 |
| 2 | 2019-01-04 | 47.91 | 47.92 | 47.10 | 46.42 | 111448000 | 1.49 | 0.82 | 0 | 10 |
| 3 | 2019-01-07 | 50.79 | 51.12 | 50.16 | 49.11 | 109012000 | 1.68 | 0.96 | 3 | 10 |
| 4 | 2019-01-08 | 53.47 | 54.51 | 51.69 | 50.79 | 216071600 | 2.69 | 2.82 | 3 | 9 |
Checking the shape of the data¶
data.shape
(349, 12)
data_unique.shape
(71, 10)
Checking for missing values¶
data.isnull().sum()
| 0 | |
|---|---|
| Date | 0 |
| News | 0 |
| Open | 0 |
| High | 0 |
| Low | 0 |
| Close | 0 |
| Volume | 0 |
| Label | 0 |
| Label_String | 0 |
| News_Char_Count | 0 |
| Open_Close | 0 |
| High_Low | 0 |
- There are no missing values in the data
Checking for duplicate values¶
# checking for duplicate values
data.duplicated().sum()
0
- There are no duplicate values in the data, however, we do know that for a given value of date, the stock related values of Open, High, Low, Close, and Volume are all the same.
Checking the statistical summary¶
# Suppress scientific notation
pd.options.display.float_format = '{:,.2f}'.format
# Display the statistical summary
data_unique.describe().T
| count | mean | min | 25% | 50% | 75% | max | std | |
|---|---|---|---|---|---|---|---|---|
| Date | 71 | 2019-03-01 00:40:33.802817024 | 2019-01-02 00:00:00 | 2019-01-29 12:00:00 | 2019-02-27 00:00:00 | 2019-03-28 12:00:00 | 2019-04-30 00:00:00 | NaN |
| Open | 71.00 | 47.20 | 37.57 | 42.61 | 47.19 | 50.79 | 66.82 | 6.82 |
| High | 71.00 | 47.66 | 37.82 | 42.92 | 47.39 | 51.12 | 67.06 | 6.85 |
| Low | 71.00 | 46.79 | 37.30 | 42.41 | 46.48 | 50.47 | 65.86 | 6.79 |
| Close | 71.00 | 45.93 | 36.25 | 41.42 | 45.67 | 49.54 | 64.81 | 6.80 |
| Volume | 71.00 | 115,724,822.54 | 45,448,000.00 | 89,389,800.00 | 109,012,000.00 | 132,407,200.00 | 244,439,200.00 | 39,094,252.34 |
| Open_Close | 71.00 | 1.28 | -0.19 | 0.92 | 1.30 | 1.52 | 2.69 | 0.51 |
| High_Low | 71.00 | 0.86 | 0.36 | 0.53 | 0.77 | 1.05 | 2.82 | 0.45 |
| Label_Sum | 71.00 | -0.27 | -10.00 | -1.00 | 0.00 | 1.00 | 3.00 | 2.09 |
| Row_Count | 71.00 | 4.92 | 1.00 | 2.00 | 4.00 | 6.00 | 28.00 | 4.39 |
Checking the datatype¶
data.info()
<class 'pandas.core.frame.DataFrame'> RangeIndex: 349 entries, 0 to 348 Data columns (total 12 columns): # Column Non-Null Count Dtype --- ------ -------------- ----- 0 Date 349 non-null datetime64[ns] 1 News 349 non-null object 2 Open 349 non-null float64 3 High 349 non-null float64 4 Low 349 non-null float64 5 Close 349 non-null float64 6 Volume 349 non-null int64 7 Label 349 non-null int64 8 Label_String 349 non-null object 9 News_Char_Count 349 non-null int64 10 Open_Close 349 non-null float64 11 High_Low 349 non-null float64 dtypes: datetime64[ns](1), float64(6), int64(3), object(2) memory usage: 32.8+ KB
Exploratory Data Analysis¶
# function to create labeled barplots
def labeled_barplot(data, feature, perc=False, n=None):
"""
Barplot with percentage at the top
data: dataframe
feature: dataframe column
perc: whether to display percentages instead of count (default is False)
n: displays the top n category levels (default is None, i.e., display all levels)
"""
total = len(data[feature]) # length of the column
count = data[feature].nunique()
if n is None:
plt.figure(figsize=(count + 1, 5))
else:
plt.figure(figsize=(n + 1, 5))
plt.xticks(rotation=90, fontsize=15)
ax = sns.countplot(
data=data,
x=feature,
hue=feature,
palette="Paired",
order=data[feature].value_counts().index[:n].sort_values(),
)
for p in ax.patches:
if perc == True:
label = "{:.1f}%".format(
100 * p.get_height() / total
) # percentage of each class of the category
else:
label = p.get_height() # count of each level of the category
x = p.get_x() + p.get_width() / 2 # width of the plot
y = p.get_height() # height of the plot
ax.annotate(
label,
(x, y),
ha="center",
va="center",
size=12,
xytext=(0, 5),
textcoords="offset points",
) # annotate the percentage
plt.show() # show the plot
def histogram_boxplot(data, feature, figsize=(15, 10), kde=False, bins=None):
"""
Boxplot and histogram combined
data: dataframe
feature: dataframe column
figsize: size of figure (default (15,10))
kde: whether to show the density curve (default False)
bins: number of bins for histogram (default None)
"""
f2, (ax_box2, ax_hist2) = plt.subplots(
nrows=2, # Number of rows of the subplot grid= 2
sharex=True, # x-axis will be shared among all subplots
gridspec_kw={"height_ratios": (0.25, 0.75)},
figsize=figsize,
) # creating the 2 subplots
sns.boxplot(
data=data, x=feature, ax=ax_box2, showmeans=True, color="violet"
) # boxplot will be created and a triangle will indicate the mean value of the column
sns.histplot(
data=data, x=feature, kde=kde, ax=ax_hist2, bins=bins
) if bins else sns.histplot(
data=data, x=feature, kde=kde, ax=ax_hist2
) # For histogram
ax_hist2.axvline(
data[feature].mean(), color="green", linestyle="--"
) # Add mean to the histogram
ax_hist2.axvline(
data[feature].median(), color="black", linestyle="-"
) # Add median to the h
Univariate Analysis¶
- Distribution of individual variables
- Compute and check the distribution of the length of news content
Distribution of sentiments¶
labeled_barplot(data, "Label_String", perc=True)
For each numerical feature, let's do a histogram and boxplot.¶
histogram_boxplot(data_unique, 'Open')
histogram_boxplot(data_unique, 'High')
histogram_boxplot(data_unique, 'Low')
histogram_boxplot(data_unique, 'Close')
histogram_boxplot(data_unique, 'Volume')
histogram_boxplot(data_unique, 'Open_Close')
histogram_boxplot(data_unique, 'High_Low')
histogram_boxplot(data, 'News_Char_Count')
# For easier comparison of the Open, Close, High, and Low values plot together on the same chart
# List of selected numeric columns
selected_columns = ['Open', 'Close', 'High', 'Low']
# Melt the DataFrame for easier boxplot handling
melted_data = data_unique[selected_columns].melt(var_name='Variable', value_name='Value')
# Create the boxplot
plt.figure(figsize=(10, 6))
sns.boxplot(x='Variable', y='Value', data=melted_data)
plt.title('Boxplot of Selected Numeric Columns')
plt.show()
Bivariate Analysis¶
- Correlation
- Sentiment Polarity vs Price
- Date vs Price
Note: The above points are listed to provide guidance on how to approach bivariate analysis. Analysis has to be done beyond the above listed points to get maximum scores.
Let's see what the correlation looks like for the numeric data and colored based on the sentiment label.¶
# Define the color mapping
custom_palette = {-1: 'red', 0: 'gray', 1: 'green'}
# Scatter plot matrix
plt.figure(figsize=(12, 8))
sns.pairplot(data, vars=num_features, hue='Label', diag_kind='kde', palette=custom_palette);
<Figure size 1200x800 with 0 Axes>
From the above we can see that the Open, High, Low, and Close valies are all highly correlated. There is no discernible correlation with the Volume or New_Chart_Count value.
Let's look at how many news articles there are per week.¶
# Extract the year and ISO week number
data['Year-Week'] = data['Date'].dt.isocalendar().year.astype(str) + '-' + data['Date'].dt.isocalendar().week.astype(str).str.zfill(2)
# Group by Year-Week and count the number of records
weekly_counts = data.groupby('Year-Week').size().reset_index(name='Count')
# Plot the counts
plt.figure(figsize=(10, 6))
plt.bar(weekly_counts['Year-Week'], weekly_counts['Count'], color='skyblue')
plt.title('Number of News Articles Per Week')
plt.xlabel('Year-Week')
plt.ylabel('Record Count')
plt.xticks(rotation=45)
plt.grid(axis='y')
plt.show()
We can see that the reviews range from Week #1-18 in 2019.
Correlation between Sentiment Polarity and Close Price¶
fig, ax1 = plt.subplots(figsize=(12, 6))
# Plot Close price
ax1.plot(data_unique['Date'], data_unique['Close'], label='Close Price', color='blue')
ax1.set_xlabel('Date', fontsize=14)
ax1.set_ylabel('Close Price', color='blue', fontsize=14)
# Add trend line for Close price
x_close = np.arange(len(data_unique['Date'])) # Convert Date to a numeric range for regression
y_close = data_unique['Close']
coef_close = np.polyfit(x_close, y_close, 1) # Fit a linear regression line
trend_close = np.polyval(coef_close, x_close)
ax1.plot(data_unique['Date'], trend_close, linestyle='--', color='blue', label='Trend Close Price')
# Plot Label_Sum on a secondary y-axis
ax2 = ax1.twinx()
ax2.plot(data_unique['Date'], data_unique['Label_Sum'], label='Sentiment Polarity', color='red')
ax2.set_ylabel('Sentiment Polarity', color='red', fontsize=14)
# Add trend line for Label_Sum
y_label_sum = data_unique['Label_Sum']
coef_label_sum = np.polyfit(x_close, y_label_sum, 1) # Fit a linear regression line
trend_label_sum = np.polyval(coef_label_sum, x_close)
ax2.plot(data_unique['Date'], trend_label_sum, linestyle='--', color='red', label='Trend Sentiment Polarity')
# Add grid and title
plt.title('Close Price and Sentiment Polarity Over Time with Trend Lines', fontsize=16)
plt.grid()
plt.show()
While the data is quite eractic from day to day, we can see that there is a slight improvement per the trend line in the Sentiment Polarity for the day which corresponds in an upward trend line of the Close price.
Finance "candlestick" style chart to show the stock Open, High, Low, and Close price for each date.¶
# Subset the DataFrame
data_unique_sub = data_unique[['Open', 'High', 'Low', 'Close']]
# Plot the candlestick chart
mpf.plot(data_unique_sub, type='candle', title='Candlestick Chart', style='yahoo', figsize=(16, 8))
In the above chart the bars represent the difference between the Open and Close price (Red if Close < Open or Green if Close > Open). The upper wick represents the difference between the High to the higher of the Open or Close, while the lower wick represents the difference from the Low to the lower of Open or Close. There is a problem with the data where in most cases the Close value is lower than the Low value which should not be possible.
Data Preprocessing¶
Removing special characters from the text¶
# defining a function to remove special characters
def remove_special_characters(text):
# Defining the regex pattern to match non-alphanumeric characters
pattern = '[^A-Za-z0-9]+'
# Finding the specified pattern and replacing non-alphanumeric characters with a blank string
new_text = ''.join(re.sub(pattern, ' ', text))
return new_text
# Applying the function to remove special characters
data['Cleaned_News'] = data['News'].apply(remove_special_characters)
# checking a couple of instances of cleaned data
data.loc[0:3, ['News','Cleaned_News']]
| News | Cleaned_News | |
|---|---|---|
| 0 | The tech sector experienced a significant decline in the aftermarket following Apple's Q1 revenue warning. Notable suppliers, including Skyworks, Broadcom, Lumentum, Qorvo, and TSMC, saw their stocks drop in response to Apple's downward revision of its revenue expectations for the quarter, previously announced in January. | The tech sector experienced a significant decline in the aftermarket following Apple s Q1 revenue warning Notable suppliers including Skyworks Broadcom Lumentum Qorvo and TSMC saw their stocks drop in response to Apple s downward revision of its revenue expectations for the quarter previously announced in January |
| 1 | Apple lowered its fiscal Q1 revenue guidance to $84 billion from earlier estimates of $89-$93 billion due to weaker than expected iPhone sales. The announcement caused a significant drop in Apple's stock price and negatively impacted related suppliers, leading to broader market declines for tech indices such as Nasdaq 10 | Apple lowered its fiscal Q1 revenue guidance to 84 billion from earlier estimates of 89 93 billion due to weaker than expected iPhone sales The announcement caused a significant drop in Apple s stock price and negatively impacted related suppliers leading to broader market declines for tech indices such as Nasdaq 10 |
| 2 | Apple cut its fiscal first quarter revenue forecast from $89-$93 billion to $84 billion due to weaker demand in China and fewer iPhone upgrades. CEO Tim Cook also mentioned constrained sales of Airpods and Macbooks. Apple's shares fell 8.5% in post market trading, while Asian suppliers like Hon | Apple cut its fiscal first quarter revenue forecast from 89 93 billion to 84 billion due to weaker demand in China and fewer iPhone upgrades CEO Tim Cook also mentioned constrained sales of Airpods and Macbooks Apple s shares fell 8 5 in post market trading while Asian suppliers like Hon |
| 3 | This news article reports that yields on long-dated U.S. Treasury securities hit their lowest levels in nearly a year on January 2, 2019, due to concerns about the health of the global economy following weak economic data from China and Europe, as well as the partial U.S. government shutdown. Apple | This news article reports that yields on long dated U S Treasury securities hit their lowest levels in nearly a year on January 2 2019 due to concerns about the health of the global economy following weak economic data from China and Europe as well as the partial U S government shutdown Apple |
Lowercasing¶
# changing the case of the text data to lower case
data['Cleaned_News'] = data['Cleaned_News'].str.lower()
# checking a couple of instances of cleaned data
data.loc[0:3, ['News','Cleaned_News']]
| News | Cleaned_News | |
|---|---|---|
| 0 | The tech sector experienced a significant decline in the aftermarket following Apple's Q1 revenue warning. Notable suppliers, including Skyworks, Broadcom, Lumentum, Qorvo, and TSMC, saw their stocks drop in response to Apple's downward revision of its revenue expectations for the quarter, previously announced in January. | the tech sector experienced a significant decline in the aftermarket following apple s q1 revenue warning notable suppliers including skyworks broadcom lumentum qorvo and tsmc saw their stocks drop in response to apple s downward revision of its revenue expectations for the quarter previously announced in january |
| 1 | Apple lowered its fiscal Q1 revenue guidance to $84 billion from earlier estimates of $89-$93 billion due to weaker than expected iPhone sales. The announcement caused a significant drop in Apple's stock price and negatively impacted related suppliers, leading to broader market declines for tech indices such as Nasdaq 10 | apple lowered its fiscal q1 revenue guidance to 84 billion from earlier estimates of 89 93 billion due to weaker than expected iphone sales the announcement caused a significant drop in apple s stock price and negatively impacted related suppliers leading to broader market declines for tech indices such as nasdaq 10 |
| 2 | Apple cut its fiscal first quarter revenue forecast from $89-$93 billion to $84 billion due to weaker demand in China and fewer iPhone upgrades. CEO Tim Cook also mentioned constrained sales of Airpods and Macbooks. Apple's shares fell 8.5% in post market trading, while Asian suppliers like Hon | apple cut its fiscal first quarter revenue forecast from 89 93 billion to 84 billion due to weaker demand in china and fewer iphone upgrades ceo tim cook also mentioned constrained sales of airpods and macbooks apple s shares fell 8 5 in post market trading while asian suppliers like hon |
| 3 | This news article reports that yields on long-dated U.S. Treasury securities hit their lowest levels in nearly a year on January 2, 2019, due to concerns about the health of the global economy following weak economic data from China and Europe, as well as the partial U.S. government shutdown. Apple | this news article reports that yields on long dated u s treasury securities hit their lowest levels in nearly a year on january 2 2019 due to concerns about the health of the global economy following weak economic data from china and europe as well as the partial u s government shutdown apple |
Removing extra whitespace¶
# removing extra whitespaces from the text
data['Cleaned_News'] = data['Cleaned_News'].str.strip()
# checking a couple of instances of cleaned data
data.loc[0:3, ['News','Cleaned_News']]
| News | Cleaned_News | |
|---|---|---|
| 0 | The tech sector experienced a significant decline in the aftermarket following Apple's Q1 revenue warning. Notable suppliers, including Skyworks, Broadcom, Lumentum, Qorvo, and TSMC, saw their stocks drop in response to Apple's downward revision of its revenue expectations for the quarter, previously announced in January. | the tech sector experienced a significant decline in the aftermarket following apple s q1 revenue warning notable suppliers including skyworks broadcom lumentum qorvo and tsmc saw their stocks drop in response to apple s downward revision of its revenue expectations for the quarter previously announced in january |
| 1 | Apple lowered its fiscal Q1 revenue guidance to $84 billion from earlier estimates of $89-$93 billion due to weaker than expected iPhone sales. The announcement caused a significant drop in Apple's stock price and negatively impacted related suppliers, leading to broader market declines for tech indices such as Nasdaq 10 | apple lowered its fiscal q1 revenue guidance to 84 billion from earlier estimates of 89 93 billion due to weaker than expected iphone sales the announcement caused a significant drop in apple s stock price and negatively impacted related suppliers leading to broader market declines for tech indices such as nasdaq 10 |
| 2 | Apple cut its fiscal first quarter revenue forecast from $89-$93 billion to $84 billion due to weaker demand in China and fewer iPhone upgrades. CEO Tim Cook also mentioned constrained sales of Airpods and Macbooks. Apple's shares fell 8.5% in post market trading, while Asian suppliers like Hon | apple cut its fiscal first quarter revenue forecast from 89 93 billion to 84 billion due to weaker demand in china and fewer iphone upgrades ceo tim cook also mentioned constrained sales of airpods and macbooks apple s shares fell 8 5 in post market trading while asian suppliers like hon |
| 3 | This news article reports that yields on long-dated U.S. Treasury securities hit their lowest levels in nearly a year on January 2, 2019, due to concerns about the health of the global economy following weak economic data from China and Europe, as well as the partial U.S. government shutdown. Apple | this news article reports that yields on long dated u s treasury securities hit their lowest levels in nearly a year on january 2 2019 due to concerns about the health of the global economy following weak economic data from china and europe as well as the partial u s government shutdown apple |
Removing stopwords¶
- The idea with stop word removal is to exclude words that appear frequently throughout all the documents in the corpus.
- Pronouns and articles are typically categorized as stop words.
- The
NLTKlibrary has an in-built list of stop words and it can utilize that list to remove the stop words from a dataset.
# defining a function to remove stop words using the NLTK library
def remove_stopwords(text):
# Split text into separate words
words = text.split()
# Removing English language stopwords
new_text = ' '.join([word for word in words if word not in stopwords.words('english')])
return new_text
# Applying the function to remove stop words using the NLTK library
data['Cleaned_News_without_stopwords'] = data['Cleaned_News'].apply(remove_stopwords)
# checking a couple of instances of cleaned data
data.loc[0:3, ['Cleaned_News','Cleaned_News_without_stopwords']]
| Cleaned_News | Cleaned_News_without_stopwords | |
|---|---|---|
| 0 | the tech sector experienced a significant decline in the aftermarket following apple s q1 revenue warning notable suppliers including skyworks broadcom lumentum qorvo and tsmc saw their stocks drop in response to apple s downward revision of its revenue expectations for the quarter previously announced in january | tech sector experienced significant decline aftermarket following apple q1 revenue warning notable suppliers including skyworks broadcom lumentum qorvo tsmc saw stocks drop response apple downward revision revenue expectations quarter previously announced january |
| 1 | apple lowered its fiscal q1 revenue guidance to 84 billion from earlier estimates of 89 93 billion due to weaker than expected iphone sales the announcement caused a significant drop in apple s stock price and negatively impacted related suppliers leading to broader market declines for tech indices such as nasdaq 10 | apple lowered fiscal q1 revenue guidance 84 billion earlier estimates 89 93 billion due weaker expected iphone sales announcement caused significant drop apple stock price negatively impacted related suppliers leading broader market declines tech indices nasdaq 10 |
| 2 | apple cut its fiscal first quarter revenue forecast from 89 93 billion to 84 billion due to weaker demand in china and fewer iphone upgrades ceo tim cook also mentioned constrained sales of airpods and macbooks apple s shares fell 8 5 in post market trading while asian suppliers like hon | apple cut fiscal first quarter revenue forecast 89 93 billion 84 billion due weaker demand china fewer iphone upgrades ceo tim cook also mentioned constrained sales airpods macbooks apple shares fell 8 5 post market trading asian suppliers like hon |
| 3 | this news article reports that yields on long dated u s treasury securities hit their lowest levels in nearly a year on january 2 2019 due to concerns about the health of the global economy following weak economic data from china and europe as well as the partial u s government shutdown apple | news article reports yields long dated u treasury securities hit lowest levels nearly year january 2 2019 due concerns health global economy following weak economic data china europe well partial u government shutdown apple |
Stemming¶
Stemming is a language processing method that chops off word endings to find the root or base form of words.
For example,
- Original Word: Jumping, Stemmed Word: Jump
- Original Word: Running, Stemmed Word: Run
The Porter Stemmer is one of the widely-used algorithms for stemming, and it shorten words to their root form by removing suffixes.
# Loading the Porter Stemmer
ps = PorterStemmer()
# defining a function to perform stemming
def apply_porter_stemmer(text):
# Split text into separate words
words = text.split()
# Applying the Porter Stemmer on every word of a message and joining the stemmed words back into a single string
new_text = ' '.join([ps.stem(word) for word in words])
return new_text
# Applying the function to perform stemming
data['Final_Cleaned_News'] = data['Cleaned_News_without_stopwords'].apply(apply_porter_stemmer)
# checking a couple of instances of cleaned data
data.loc[0:2,['Cleaned_News_without_stopwords','Final_Cleaned_News']]
| Cleaned_News_without_stopwords | Final_Cleaned_News | |
|---|---|---|
| 0 | tech sector experienced significant decline aftermarket following apple q1 revenue warning notable suppliers including skyworks broadcom lumentum qorvo tsmc saw stocks drop response apple downward revision revenue expectations quarter previously announced january | tech sector experienc signific declin aftermarket follow appl q1 revenu warn notabl supplier includ skywork broadcom lumentum qorvo tsmc saw stock drop respons appl downward revis revenu expect quarter previous announc januari |
| 1 | apple lowered fiscal q1 revenue guidance 84 billion earlier estimates 89 93 billion due weaker expected iphone sales announcement caused significant drop apple stock price negatively impacted related suppliers leading broader market declines tech indices nasdaq 10 | appl lower fiscal q1 revenu guidanc 84 billion earlier estim 89 93 billion due weaker expect iphon sale announc caus signific drop appl stock price neg impact relat supplier lead broader market declin tech indic nasdaq 10 |
| 2 | apple cut fiscal first quarter revenue forecast 89 93 billion 84 billion due weaker demand china fewer iphone upgrades ceo tim cook also mentioned constrained sales airpods macbooks apple shares fell 8 5 post market trading asian suppliers like hon | appl cut fiscal first quarter revenu forecast 89 93 billion 84 billion due weaker demand china fewer iphon upgrad ceo tim cook also mention constrain sale airpod macbook appl share fell 8 5 post market trade asian supplier like hon |
Word Embeddings¶
Word2Vec¶
Word2Vec is a popular natural language processing (NLP) technique for learning word embeddings, introduced by Google in 2013. It uses shallow neural networks to transform words into dense vector representations that capture semantic relationships based on their context in a corpus.
# creating a list of all words in our data
words_list = [item.split(" ") for item in data['Final_Cleaned_News'].values]
# Checking the words from the first 3 news articles
words_list[0:3]
[['tech', 'sector', 'experienc', 'signific', 'declin', 'aftermarket', 'follow', 'appl', 'q1', 'revenu', 'warn', 'notabl', 'supplier', 'includ', 'skywork', 'broadcom', 'lumentum', 'qorvo', 'tsmc', 'saw', 'stock', 'drop', 'respons', 'appl', 'downward', 'revis', 'revenu', 'expect', 'quarter', 'previous', 'announc', 'januari'], ['appl', 'lower', 'fiscal', 'q1', 'revenu', 'guidanc', '84', 'billion', 'earlier', 'estim', '89', '93', 'billion', 'due', 'weaker', 'expect', 'iphon', 'sale', 'announc', 'caus', 'signific', 'drop', 'appl', 'stock', 'price', 'neg', 'impact', 'relat', 'supplier', 'lead', 'broader', 'market', 'declin', 'tech', 'indic', 'nasdaq', '10'], ['appl', 'cut', 'fiscal', 'first', 'quarter', 'revenu', 'forecast', '89', '93', 'billion', '84', 'billion', 'due', 'weaker', 'demand', 'china', 'fewer', 'iphon', 'upgrad', 'ceo', 'tim', 'cook', 'also', 'mention', 'constrain', 'sale', 'airpod', 'macbook', 'appl', 'share', 'fell', '8', '5', 'post', 'market', 'trade', 'asian', 'supplier', 'like', 'hon']]
#Defining the dimension of the embedded vector.
vec_size=100
# creating an instance of Word2Vec, by default CBOW is used, rather than Skip-gram
model_W2V = Word2Vec(words_list, vector_size = vec_size, min_count = 1, workers = 6)
# Checking the size of the vocabulary
print("Length of the vocabulary is", len(list(model_W2V.wv.key_to_index)))
Length of the vocabulary is 2580
# Checking the word embedding of a random word
word = "fiscal"
model_W2V.wv[word]
array([ 0.00490734, -0.00088611, -0.00297797, 0.00540985, 0.00958347,
-0.00958432, 0.0087942 , 0.00288539, 0.00693005, -0.00199828,
-0.01003374, 0.00461714, 0.00915809, 0.00602206, 0.00624163,
-0.0118016 , 0.00548494, -0.00525224, -0.0039441 , -0.00395565,
0.00702446, 0.00123063, 0.01044334, 0.00692605, 0.00948722,
0.00873028, 0.00320285, -0.00910078, 0.00420447, 0.00765262,
0.00203086, -0.00644444, -0.00218457, -0.00504377, -0.00048371,
0.00873844, 0.00349679, -0.01009597, -0.00693369, -0.01403526,
-0.00762316, -0.0105751 , -0.00968795, 0.00787044, -0.00304691,
0.00749471, 0.00202889, -0.00592009, 0.01045779, 0.01113912,
0.0061313 , -0.01104943, -0.01002257, 0.00712905, 0.00377165,
-0.00171765, -0.00742769, 0.00599708, -0.00299191, -0.00708066,
-0.00474398, -0.00724669, 0.00273155, -0.00099778, 0.00251359,
0.00928195, 0.00541441, -0.00122063, -0.01252228, -0.00165958,
-0.0115237 , 0.00184164, 0.01280665, -0.00330248, 0.00129604,
-0.00801472, 0.00819776, -0.00314116, -0.00142288, 0.0065364 ,
0.00311128, -0.00884774, -0.00268911, 0.01406473, -0.010445 ,
-0.00077327, 0.00391806, 0.01015147, 0.00497957, 0.01039966,
-0.00512604, 0.00764369, 0.00073153, 0.00263202, -0.00017564,
0.00256066, -0.00492771, -0.01124426, 0.00847635, 0.00699234],
dtype=float32)
# Checking top 5 similar words to the word 'price'
similar = model_W2V.wv.similar_by_word('fiscal', topn=5)
print(similar)
[('logist', 0.4777337312698364), ('talk', 0.47714221477508545), ('app', 0.46340253949165344), ('trade', 0.4600534737110138), ('new', 0.4576209485530853)]
# Dictionary with key as words and the value as the embedding vector.
words = model_W2V.wv.key_to_index
# Preview the first 5 entries in the dictionary
preview = list(words.items())[:5]
print(preview)
[('appl', 0), ('china', 1), ('stock', 2), ('report', 3), ('compani', 4)]
def average_vectorizer_Word2Vec(doc):
# Initializing a feature vector for the sentence
feature_vector = np.zeros(vec_size, dtype="float64")
# Creating a list of words in the sentence that are present in the model vocabulary
words_in_vocab = [word for word in doc.split() if word in model_W2V.wv]
# Add the vector representations of the words
for word in words_in_vocab:
feature_vector += model_W2V.wv[word]
# Dividing by the number of words to get the average vector
if len(words_in_vocab) > 0:
feature_vector /= len(words_in_vocab)
return feature_vector
# creating a dataframe of the vectorized documents
df_word2vec = pd.DataFrame(data['Final_Cleaned_News'].apply(average_vectorizer_Word2Vec).tolist(), columns=['Feature '+str(i) for i in range(vec_size)])
df_word2vec
| Feature 0 | Feature 1 | Feature 2 | Feature 3 | Feature 4 | Feature 5 | Feature 6 | Feature 7 | Feature 8 | Feature 9 | ... | Feature 90 | Feature 91 | Feature 92 | Feature 93 | Feature 94 | Feature 95 | Feature 96 | Feature 97 | Feature 98 | Feature 99 | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 0 | -0.003392 | 0.012078 | 0.003880 | 0.001852 | 0.001697 | -0.019270 | 0.002941 | 0.027341 | -0.006952 | -0.007268 | ... | 0.015694 | 0.005844 | 0.001833 | 0.002261 | 0.019891 | 0.010866 | 0.005425 | -0.012500 | 0.002716 | -0.001451 |
| 1 | -0.005132 | 0.012922 | 0.004624 | 0.002074 | 0.000740 | -0.023828 | 0.003841 | 0.030416 | -0.005991 | -0.007827 | ... | 0.018081 | 0.006465 | 0.001994 | 0.004069 | 0.023411 | 0.012184 | 0.006233 | -0.014573 | 0.002863 | -0.003039 |
| 2 | -0.004184 | 0.012508 | 0.004161 | 0.001998 | 0.002608 | -0.022112 | 0.003211 | 0.028248 | -0.005493 | -0.006948 | ... | 0.017009 | 0.006072 | 0.000625 | 0.003125 | 0.020047 | 0.012189 | 0.005308 | -0.014774 | 0.000637 | -0.001744 |
| 3 | -0.005275 | 0.012699 | 0.005840 | 0.001705 | 0.000236 | -0.020950 | 0.003786 | 0.028516 | -0.007097 | -0.007338 | ... | 0.016906 | 0.002901 | -0.000092 | 0.001967 | 0.020378 | 0.011036 | 0.004492 | -0.012530 | 0.002682 | -0.002250 |
| 4 | -0.003924 | 0.011533 | 0.003834 | 0.002347 | 0.001770 | -0.017197 | 0.002016 | 0.024016 | -0.004102 | -0.005053 | ... | 0.014025 | 0.006180 | 0.001340 | 0.002914 | 0.018628 | 0.008251 | 0.002922 | -0.010100 | 0.001309 | -0.001879 |
| ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... |
| 344 | -0.003708 | 0.004882 | 0.003544 | -0.001492 | 0.001838 | -0.007881 | 0.000769 | 0.012026 | -0.002401 | -0.001645 | ... | 0.005878 | 0.000466 | -0.000457 | -0.000188 | 0.006205 | 0.005683 | 0.003361 | -0.004775 | 0.001840 | 0.000246 |
| 345 | -0.001175 | 0.008862 | 0.003189 | 0.000598 | 0.001407 | -0.012676 | 0.003323 | 0.017915 | -0.005944 | -0.004987 | ... | 0.009941 | 0.003640 | 0.000348 | 0.002338 | 0.013408 | 0.006861 | 0.003428 | -0.010040 | 0.001554 | -0.000693 |
| 346 | -0.005371 | 0.010723 | 0.002088 | 0.001753 | 0.003583 | -0.014902 | 0.003426 | 0.023352 | -0.004235 | -0.004497 | ... | 0.013923 | 0.004020 | 0.001818 | 0.002295 | 0.015907 | 0.007772 | 0.006056 | -0.009076 | 0.002524 | -0.000945 |
| 347 | -0.004412 | 0.011307 | 0.002666 | 0.000813 | 0.001143 | -0.018848 | 0.003425 | 0.025135 | -0.004535 | -0.007266 | ... | 0.015246 | 0.003580 | 0.000887 | 0.002229 | 0.019492 | 0.011115 | 0.004311 | -0.012893 | 0.000184 | -0.004159 |
| 348 | -0.004898 | 0.010505 | 0.002135 | 0.002024 | 0.001285 | -0.018407 | 0.002769 | 0.024588 | -0.003441 | -0.006206 | ... | 0.015973 | 0.001951 | 0.000962 | 0.004091 | 0.017758 | 0.010830 | 0.004712 | -0.010964 | 0.000681 | -0.002843 |
349 rows × 100 columns
GloVe¶
from gensim.models import KeyedVectors
# load the Stanford GloVe model
filename = '/content/drive/MyDrive/Personal/UT Austin/Project 6 - NLP/glove.6B.100d.txt.word2vec'
glove_model = KeyedVectors.load_word2vec_format(filename, binary=False)
# Checking the size of the vocabulary
print("Length of the vocabulary is", len(glove_model.index_to_key))
Length of the vocabulary is 400000
# Checking the word embedding of a random word
word = "fiscal"
glove_model[word]
array([ 0.037566 , -0.3505 , 0.6053 , -0.88877 , 0.26267 ,
-0.85035 , -0.41264 , 0.15612 , -0.38347 , 0.75271 ,
0.14319 , 0.036477 , 0.21866 , 0.17531 , -0.046873 ,
-0.16092 , -0.47602 , 0.43815 , -0.0029819, -0.54908 ,
0.29455 , -0.69913 , -0.78756 , 1.0152 , -1.1138 ,
-0.22557 , 0.16026 , -0.7821 , -1.0417 , -0.73787 ,
0.4298 , 0.38303 , -0.5378 , -0.52836 , -0.68083 ,
0.62297 , -0.36233 , -0.14177 , 0.32653 , -0.014921 ,
-0.19752 , -0.82104 , 0.38737 , 0.83035 , -0.50677 ,
-0.030901 , 0.62872 , -0.75123 , -1.0036 , -0.76945 ,
-0.23083 , -0.56152 , -0.19614 , 0.79836 , -0.21336 ,
-2.2264 , 1.2042 , -0.81874 , 1.3231 , 0.41051 ,
-0.016212 , -0.88813 , -1.8089 , 0.046677 , -0.10258 ,
-0.02645 , -1.0452 , -0.59149 , 1.0489 , -0.75717 ,
0.7852 , 0.58732 , -0.746 , 0.38022 , -0.54294 ,
0.31535 , -1.1958 , 1.2508 , -0.44103 , 0.076059 ,
0.16927 , 0.42416 , -0.17511 , 0.58678 , -0.51208 ,
-0.30556 , -0.33275 , -0.53697 , -0.25197 , -0.57803 ,
0.39421 , 0.87378 , -0.50057 , -0.17725 , -0.25021 ,
0.11751 , 0.82264 , -0.0396 , 0.83154 , -0.26962 ],
dtype=float32)
#Returning the top 5 similar words.
result = glove_model.most_similar("fiscal", topn=5)
print(result)
[('budget', 0.7522984743118286), ('budgetary', 0.7011541724205017), ('spending', 0.7007430791854858), ('deficit', 0.6957388520240784), ('economic', 0.671556293964386)]
#Defining the dimension of the embedded vector.
vec_size=100
def average_vectorizer_GloVe(doc):
# Initializing a feature vector for the sentence
feature_vector = np.zeros(vec_size, dtype="float64")
# Creating a list of words in the sentence that are present in the model vocabulary
words_in_vocab = [word for word in doc.split() if word in glove_model]
# Add vectors for the words in the vocabulary
for word in words_in_vocab:
feature_vector += model[word]
# Compute the average vector if there are valid words
if len(words_in_vocab) > 0:
feature_vector /= len(words_in_vocab)
return feature_vector
# Create a DataFrame of vectorized documents
df_glove = pd.DataFrame(
data['Final_Cleaned_News'].apply(average_vectorizer_GloVe).tolist(),
columns=[f'Feature {i}' for i in range(vec_size)]
)
df_glove
| Feature 0 | Feature 1 | Feature 2 | Feature 3 | Feature 4 | Feature 5 | Feature 6 | Feature 7 | Feature 8 | Feature 9 | ... | Feature 90 | Feature 91 | Feature 92 | Feature 93 | Feature 94 | Feature 95 | Feature 96 | Feature 97 | Feature 98 | Feature 99 | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 0 | 0.021671 | 0.096785 | -0.047464 | -0.066726 | -0.215078 | -0.603108 | -0.089054 | -0.001979 | 0.114239 | -0.132121 | ... | -0.016118 | 0.282037 | -0.146394 | -0.217284 | -0.137730 | 0.162514 | 0.208869 | 0.056668 | 0.186156 | -0.025540 |
| 1 | 0.171825 | 0.341351 | 0.234650 | -0.042054 | -0.082840 | -0.600302 | -0.060674 | -0.101788 | -0.159387 | 0.023604 | ... | 0.001762 | 0.271578 | -0.190244 | -0.136070 | -0.442046 | 0.225178 | 0.175980 | -0.036701 | 0.410090 | -0.159631 |
| 2 | 0.010512 | 0.270341 | 0.301482 | -0.087113 | 0.075485 | -0.476547 | -0.039534 | -0.014161 | -0.123561 | -0.054536 | ... | 0.113562 | 0.199301 | -0.095302 | -0.220799 | -0.559601 | 0.130351 | 0.039756 | -0.088995 | 0.484577 | -0.204774 |
| 3 | -0.147407 | 0.226970 | 0.377158 | 0.184659 | -0.110523 | -0.447424 | -0.103314 | 0.033698 | -0.021951 | -0.041905 | ... | 0.031969 | 0.314802 | -0.276519 | 0.034750 | -0.347012 | 0.131816 | 0.201519 | -0.220721 | 0.328415 | -0.101222 |
| 4 | 0.040798 | 0.198312 | 0.071460 | 0.037332 | -0.098294 | -0.407203 | -0.078844 | -0.102327 | -0.132380 | -0.031861 | ... | 0.041134 | 0.188611 | -0.051261 | -0.364269 | -0.216124 | 0.242409 | 0.243280 | -0.071136 | 0.173813 | -0.131392 |
| ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... |
| 344 | -0.133417 | 0.058861 | 0.415255 | -0.291097 | 0.051498 | 0.079182 | 0.057609 | 0.127537 | -0.074177 | -0.070428 | ... | 0.306373 | -0.166262 | 0.083386 | -0.120771 | -0.347820 | 0.049615 | -0.054991 | -0.299478 | 0.437251 | 0.188692 |
| 345 | 0.159967 | 0.246612 | 0.288849 | 0.093632 | 0.043483 | -0.284527 | -0.167188 | -0.049947 | -0.246812 | -0.088806 | ... | 0.042174 | 0.246746 | -0.054319 | -0.047361 | -0.572373 | 0.368322 | 0.077437 | 0.020181 | 0.441181 | 0.015956 |
| 346 | 0.043433 | 0.127144 | 0.151791 | -0.053902 | -0.067352 | -0.187623 | -0.020754 | 0.138385 | -0.354622 | -0.084654 | ... | -0.066027 | 0.240859 | -0.119884 | -0.026659 | -0.428011 | 0.054745 | -0.025141 | -0.129436 | 0.258558 | -0.027775 |
| 347 | -0.098248 | 0.090049 | 0.181622 | -0.028409 | -0.128997 | -0.497643 | -0.247336 | -0.034014 | -0.103773 | -0.195073 | ... | 0.012848 | 0.070933 | -0.053851 | -0.265540 | -0.415342 | 0.111974 | 0.295720 | -0.207213 | 0.378713 | 0.026234 |
| 348 | 0.066789 | 0.339016 | 0.296032 | -0.123497 | -0.115996 | -0.460525 | 0.035026 | 0.121192 | -0.158255 | -0.070437 | ... | -0.013131 | 0.204805 | -0.259545 | 0.037068 | -0.419796 | 0.014846 | 0.085902 | -0.183811 | 0.634333 | -0.208610 |
349 rows × 100 columns
Sentence Transformer¶
The all-MiniLM-L6-v2 is a sentence-transformers model and is intended to be used as a sentence and short paragraph encoder. It maps sentences & paragraphs to a 384 dimensional dense vector space and can be used for tasks like clustering or semantic search. By default, input text longer than 256 word pieces is truncated. See: https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2
# defining the model
model = SentenceTransformer('sentence-transformers/all-MiniLM-L6-v2')
modules.json: 0%| | 0.00/349 [00:00<?, ?B/s]
config_sentence_transformers.json: 0%| | 0.00/116 [00:00<?, ?B/s]
README.md: 0%| | 0.00/10.7k [00:00<?, ?B/s]
sentence_bert_config.json: 0%| | 0.00/53.0 [00:00<?, ?B/s]
config.json: 0%| | 0.00/612 [00:00<?, ?B/s]
model.safetensors: 0%| | 0.00/90.9M [00:00<?, ?B/s]
tokenizer_config.json: 0%| | 0.00/350 [00:00<?, ?B/s]
vocab.txt: 0%| | 0.00/232k [00:00<?, ?B/s]
tokenizer.json: 0%| | 0.00/466k [00:00<?, ?B/s]
special_tokens_map.json: 0%| | 0.00/112 [00:00<?, ?B/s]
1_Pooling/config.json: 0%| | 0.00/190 [00:00<?, ?B/s]
# setting the device to GPU if available, else CPU
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
# encoding the dataset
# Note: in the case of the sentence transformer we encode the original News text
embedding_matrix = model.encode(data['News'], device=device, show_progress_bar=True)
Batches: 0%| | 0/11 [00:00<?, ?it/s]
# printing the shape of the embedding matrix
embedding_matrix.shape
(349, 384)
- Each news article has been converted to a 384-dimensional vector
# printing the embedding vector of the first review in the dataset
embedding_matrix[0,:]
array([-2.02308688e-03, -3.67735364e-02, 7.73542598e-02, 4.67134081e-02,
3.25521305e-02, 2.10231193e-03, 4.32834737e-02, 3.95344906e-02,
5.82279935e-02, 8.87513440e-03, 7.09636807e-02, 4.99076843e-02,
6.46608099e-02, -4.97970264e-03, -1.30518479e-02, -2.98355557e-02,
-8.91322084e-03, -7.82000348e-02, -2.17109658e-02, -5.24822623e-02,
-5.14276549e-02, -3.30719985e-02, -3.32051516e-02, 4.18125615e-02,
7.99547583e-02, 1.54092284e-02, -2.15781443e-02, 5.19438535e-02,
-4.65799086e-02, -3.71372141e-02, -1.04225606e-01, 9.86078903e-02,
5.21786399e-02, 3.46578993e-02, 1.48809813e-02, -4.47350414e-03,
5.70117794e-02, -2.41722930e-02, 2.14048829e-02, -6.52144775e-02,
-3.30645032e-02, 1.61960777e-02, -6.63141608e-02, 4.39943299e-02,
3.82153168e-02, -4.86519411e-02, 1.62651073e-02, -4.02665772e-02,
-3.34568229e-03, 3.20955738e-02, -3.91200650e-03, -1.26830470e-02,
4.49698754e-02, 3.39314081e-02, -5.17101809e-02, 5.32578826e-02,
-4.53628600e-02, -2.19655819e-02, 7.14351833e-02, 3.10435817e-02,
4.33289856e-02, -9.26511064e-02, 1.73018910e-02, -8.11785087e-03,
6.26880452e-02, -1.42353158e-02, -4.14227284e-02, 5.68226501e-02,
-8.62419307e-02, 6.02035522e-02, 4.75493930e-02, -7.17870072e-02,
1.17750419e-02, 4.39625569e-02, -3.44558358e-02, 3.67184021e-02,
5.21757454e-02, -2.14000400e-02, -5.96954022e-03, 8.72079656e-03,
3.53135467e-02, -3.95235457e-02, -9.36869308e-02, 1.56855937e-02,
-3.84228751e-02, 2.47944966e-02, 3.62802781e-02, -1.32978177e-02,
-4.67082486e-02, -5.69416471e-02, 6.84317723e-02, -2.25550658e-03,
1.43187074e-02, -1.19597111e-02, 5.61927594e-02, -5.40325232e-02,
1.84007827e-02, -9.87436026e-02, 1.36832474e-02, 1.76982135e-02,
4.83816527e-02, 3.39467935e-02, 5.20425402e-02, -3.35446149e-02,
-4.39126827e-02, -9.73524675e-02, 4.13890406e-02, 2.77670706e-03,
3.94195206e-02, 1.73664875e-02, -2.37191021e-02, 1.02214050e-02,
-5.20212874e-02, -4.82435077e-02, -6.77878410e-02, 7.58207291e-02,
-1.03107914e-02, 6.10160008e-02, -4.92668599e-02, -1.84614286e-02,
-5.64168356e-02, 6.09038062e-02, -4.93265828e-03, 1.02207344e-02,
3.61768417e-02, 1.62944254e-02, -1.01151675e-01, -9.37196024e-34,
1.30963568e-02, 2.90462077e-02, -5.68401664e-02, -1.07232621e-02,
3.43215726e-02, -5.42571954e-02, 6.12289645e-02, 6.73203496e-03,
-6.39356300e-02, -3.71847749e-02, -4.51456308e-02, 5.50827645e-02,
-5.48340753e-02, -9.50715691e-02, 7.80199915e-02, -1.12085916e-01,
-5.90517521e-02, -8.70311912e-03, 6.69128373e-02, -2.09151953e-02,
-4.48345281e-02, -6.54890388e-02, -5.31035922e-02, 6.96217045e-02,
8.78782757e-03, 7.72807980e-04, 1.89882703e-03, 3.55078094e-02,
1.91084370e-02, 2.86549870e-02, -5.54806180e-02, 6.67085126e-02,
2.76120994e-02, -1.35227054e-01, -2.72631440e-02, -8.76491703e-03,
-6.54005632e-02, 2.73343958e-02, 8.32479000e-02, -4.15787622e-02,
-6.82980940e-02, 6.94951788e-02, -5.72324954e-02, 5.30299847e-04,
-4.17730864e-03, -3.16040106e-02, 4.05772962e-02, -4.03363183e-02,
-3.90767045e-02, -3.31442282e-02, -6.78732619e-02, 1.14780087e-02,
2.73937657e-02, -2.94763129e-02, 5.07603735e-02, -1.22914165e-02,
4.15568054e-02, -5.26329800e-02, -1.52790388e-02, 6.56478256e-02,
7.96086937e-02, 7.34061301e-02, -4.55296673e-02, 4.26630601e-02,
-7.28153586e-02, 8.61650407e-02, 9.41243321e-02, 3.60597968e-02,
-1.42589822e-01, 1.10695466e-01, -2.01496780e-02, -4.56183329e-02,
-2.93876100e-02, 6.98730955e-03, 5.36932424e-02, -1.35457525e-02,
-8.81077200e-02, -2.48123445e-02, -4.63539083e-03, -9.48696285e-02,
6.09889068e-02, -3.14058773e-02, 3.15860510e-02, 2.20643748e-02,
3.89253758e-02, -7.49614369e-03, 8.55575129e-02, -4.86009046e-02,
-1.11074734e-03, 5.33099240e-03, -5.82354516e-02, -4.99342568e-02,
-2.81568989e-02, 5.94845489e-02, -6.10694103e-03, -7.56994760e-34,
-4.94262390e-02, 1.74969826e-02, -2.44862102e-02, 2.48290002e-02,
-8.59438926e-02, -1.83840487e-02, 5.83668575e-02, 6.86286390e-03,
1.76914260e-02, 6.17085546e-02, 4.79522906e-02, 4.48756404e-02,
-9.16801393e-02, 1.66509878e-02, -3.22471485e-02, -3.55337863e-03,
6.30989298e-02, -1.44052148e-01, 4.32584174e-02, -7.65528381e-02,
6.65068254e-02, -3.25963572e-02, -2.85959151e-02, -2.90052872e-02,
-5.33847362e-02, 6.87323511e-02, -5.31523395e-03, 7.13094622e-02,
-1.95374023e-02, -4.22353158e-03, -7.53196999e-02, -7.40312785e-02,
1.34283388e-02, 6.54501915e-02, 7.09960088e-02, 3.40943178e-03,
-3.15817855e-02, 1.96611229e-03, -2.51394920e-02, -2.93941796e-02,
5.40001728e-02, 1.96978673e-02, 2.19473094e-02, 1.68580730e-02,
2.46497244e-02, -4.93298136e-02, 3.24860141e-02, 1.54560497e-02,
9.43609327e-02, -1.62996035e-02, 9.13450867e-03, 4.94888797e-02,
2.44886857e-02, 7.70643204e-02, -6.84276223e-02, 3.09738442e-02,
-2.02454627e-03, 6.57486171e-02, -9.38693434e-02, 1.92349646e-02,
4.56197129e-04, -5.13908081e-02, 3.38722169e-02, -3.80237140e-02,
2.85909586e-02, -2.24880148e-02, 8.34821463e-02, 4.39710878e-02,
2.28685364e-02, -4.97476235e-02, 1.22772634e-01, -3.04845572e-02,
1.06069371e-02, -7.15398118e-02, -3.84447798e-02, 5.95832765e-02,
-9.02277008e-02, 1.52782211e-02, -5.60677387e-02, -1.62651632e-02,
1.27833351e-01, 8.89350697e-02, 1.75961980e-03, 4.24433500e-02,
-3.15954722e-02, 7.07846731e-02, 2.89797056e-02, 2.44632196e-02,
-1.99160688e-02, 5.09029515e-02, -1.44732865e-02, -7.35410452e-02,
3.10671497e-02, 7.79150426e-02, -1.30469561e-01, -3.58579761e-08,
-1.97357070e-02, -7.96370581e-02, 2.42430381e-02, -1.87578034e-02,
6.77263886e-02, -2.62168534e-02, 1.00161145e-02, 4.15350460e-02,
1.08690090e-01, 3.81879359e-02, -7.81213790e-02, -1.09489430e-02,
-1.17989458e-01, 7.97276795e-02, 2.46654917e-02, 9.32490744e-04,
-7.25726709e-02, 8.93294718e-03, 4.05104943e-02, -1.03700548e-01,
-6.20276993e-03, 3.49569283e-02, 7.82294050e-02, -2.42466442e-02,
-3.33974361e-02, 3.95032912e-02, 1.34983344e-03, -7.48429075e-02,
5.03437705e-02, 5.74472127e-03, -2.39702631e-02, 1.32117537e-03,
6.57341480e-02, -5.62656298e-02, -6.76571950e-02, -4.22494002e-02,
-2.64710542e-02, 2.83033047e-02, 4.93412018e-02, 3.14059108e-02,
-3.22604328e-02, -1.48835257e-02, -7.07752928e-02, -9.63545218e-03,
-3.38620879e-02, -1.05421888e-02, -5.47044948e-02, 2.66298596e-02,
5.15929535e-02, -1.93037149e-02, 1.74268726e-02, -2.24287678e-02,
5.11563558e-04, 2.79130321e-02, -6.66390806e-02, -5.92727289e-02,
3.25671141e-03, -5.93598583e-04, -5.50961010e-02, -2.71336753e-02,
-3.57511383e-03, -1.31502107e-01, 7.41634518e-02, 5.75100295e-02],
dtype=float32)
Sentiment Analysis¶
Helper Functions¶
# creating a function to plot the confusion matrix
def plot_confusion_matrix(actual, predicted):
cm = confusion_matrix(actual, predicted)
plt.figure(figsize = (5, 4))
label_list = ['Negative', 'Neutral', 'Positive']
sns.heatmap(cm, annot = True, fmt = '.0f', cmap='Blues', xticklabels = label_list, yticklabels = label_list)
plt.ylabel('Actual')
plt.xlabel('Predicted')
plt.show()
# defining a function to compute different metrics to check performance of a classification model built using sklearn
def model_performance_classification_sklearn(model, predictors, target):
"""
Function to compute different metrics to check classification model performance
model: classifier
predictors: independent variables
target: dependent variable
"""
# predicting using the independent variables
pred = model.predict(predictors)
acc = accuracy_score(target, pred) # to compute Accuracy
recall = recall_score(target, pred, average='weighted') # to compute Recall
precision = precision_score(target, pred, average='weighted') # to compute Precision
f1 = f1_score(target, pred, average='weighted') # to compute F1-score
# creating a dataframe of metrics
df_perf = pd.DataFrame(
{
"Accuracy": acc,
"Recall": recall,
"Precision": precision,
"F1": f1,
},
index=[0],
)
return df_perf
Random Forest Model with Word2Vec¶
# Storing independent variable
X = df_word2vec.copy()
# Storing target variable
y = data['Label']
# Split data into training and testing set.
X_train, X_test, y_train, y_test = train_test_split(X ,y, test_size = 0.25, random_state = 42, stratify=y)
# Apply SMOTE to balance the training data
smote = SMOTE(random_state=42)
X_train, y_train = smote.fit_resample(X_train, y_train)
RF Base Model¶
# Building the model
rf_word2vec_base = RandomForestClassifier(class_weight= "balanced",random_state = 42)
# Fitting on train data
rf_word2vec_base.fit(X_train, y_train)
RandomForestClassifier(class_weight='balanced', random_state=42)In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
RandomForestClassifier(class_weight='balanced', random_state=42)
# Predicting on train data
y_pred_train_base_w2v = rf_word2vec_base.predict(X_train)
# Predicting on test data
y_pred_test_base_w2v = rf_word2vec_base.predict(X_test)
#Confusion Matrix for Train
plot_confusion_matrix(y_train, y_pred_train_base_w2v)
#Calculating the metrics on training data
word2vec_base_train=model_performance_classification_sklearn(rf_word2vec_base, X_train,y_train)
print("Training performance:\n", word2vec_base_train)
Training performance:
Accuracy Recall Precision F1
0 1.0 1.0 1.0 1.0
Here we can see that we have overfitting, given the pefect 1.0 score on the training data.
#Confusion Matrix for Test
plot_confusion_matrix(y_test, y_pred_test_base_w2v)
#Calculating different metrics on testing data
word2vec_base_test=model_performance_classification_sklearn(rf_word2vec_base, X_test,y_test)
print("Testing performance:\n", word2vec_base_test)
Testing performance:
Accuracy Recall Precision F1
0 0.477273 0.477273 0.467157 0.453467
print(classification_report(y_test, y_pred_test_base_w2v))
precision recall f1-score support
-1 0.39 0.44 0.42 25
0 0.53 0.65 0.58 43
1 0.43 0.15 0.22 20
accuracy 0.48 88
macro avg 0.45 0.41 0.41 88
weighted avg 0.47 0.48 0.45 88
Results on Test data are significantly worse than Training data indicating overfitting.
RF Model with Grid Search¶
# Choose the type of classifier.
word2vec_rf_tuned = RandomForestClassifier(class_weight= "balanced",random_state=42,bootstrap=True)
parameters = {
'max_depth': list(np.arange(5,10,2)),
'n_estimators': np.arange(50,110,25),
'max_features': [0.3,0.4]
}
# Run the grid search
grid_obj = GridSearchCV(word2vec_rf_tuned, parameters, scoring='recall',cv=5,n_jobs=-1)
grid_obj = grid_obj.fit(X_train, y_train)
# Set the clf to the best combination of parameters
word2vec_rf_tuned = grid_obj.best_estimator_
# Fit the best algorithm to the data.
word2vec_rf_tuned.fit(X_train, y_train)
RandomForestClassifier(class_weight='balanced', max_depth=5, max_features=0.3,
n_estimators=50, random_state=42)In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook. On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
RandomForestClassifier(class_weight='balanced', max_depth=5, max_features=0.3,
n_estimators=50, random_state=42)# Predicting on train data
y_pred_train_tuned_w2v = word2vec_rf_tuned.predict(X_train)
# Predicting on test data
y_pred_test_tuned_w2v = word2vec_rf_tuned.predict(X_test)
plot_confusion_matrix(y_train, y_pred_train_tuned_w2v)
#Calculating different metrics on training data
word2vec_tuned_train=model_performance_classification_sklearn(word2vec_rf_tuned, X_train,y_train)
print("Training performance:\n", word2vec_tuned_train)
Training performance:
Accuracy Recall Precision F1
0 0.965879 0.965879 0.965857 0.965858
plot_confusion_matrix(y_test, y_pred_test_tuned_w2v)
#Calculating different metrics on testing data
word2vec_tuned_test=model_performance_classification_sklearn(word2vec_rf_tuned, X_test,y_test)
print("Testing performance:\n", word2vec_tuned_test)
Testing performance:
Accuracy Recall Precision F1
0 0.454545 0.454545 0.460147 0.454951
print(classification_report(y_test, y_pred_test_tuned_w2v))
precision recall f1-score support
-1 0.35 0.44 0.39 25
0 0.56 0.53 0.55 43
1 0.38 0.30 0.33 20
accuracy 0.45 88
macro avg 0.43 0.42 0.42 88
weighted avg 0.46 0.45 0.45 88
Random Forest Model with GloVe¶
# Storing independent variable
X = df_glove.copy()
# Storing target variable
y = data['Label']
# Split data into training and testing set.
X_train_glove, X_test_glove, y_train_glove, y_test_glove = train_test_split(X ,y, test_size = 0.25, random_state = 42, stratify=y)
# Apply SMOTE to balance the training data
smote = SMOTE(random_state=42)
X_train_glove, y_train_glove = smote.fit_resample(X_train_glove, y_train_glove)
RF Base model¶
# Building the model
rf_glovec_base = RandomForestClassifier(class_weight= "balanced",random_state = 42)
# Fitting on train data
rf_glovec_base.fit(X_train_glove, y_train_glove)
RandomForestClassifier(class_weight='balanced', random_state=42)In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
RandomForestClassifier(class_weight='balanced', random_state=42)
# Predicting on train data
y_pred_train_base_gl = rf_glovec_base.predict(X_train_glove)
# Predicting on test data
y_pred_test_base_gl = rf_glovec_base.predict(X_test_glove)
plot_confusion_matrix(y_train_glove, y_pred_train_base_gl)
#Calculating different metrics on training data
glove_base_train=model_performance_classification_sklearn(rf_glovec_base, X_train_glove,y_train_glove)
print("Training performance:\n", glove_base_train)
Training performance:
Accuracy Recall Precision F1
0 1.0 1.0 1.0 1.0
plot_confusion_matrix(y_test_glove, y_pred_test_base_gl)
#Calculating different metrics on testing data
glove_base_test=model_performance_classification_sklearn(rf_glovec_base, X_test_glove,y_test_glove)
print("Testing performance:\n", glove_base_test)
Testing performance:
Accuracy Recall Precision F1
0 0.420455 0.420455 0.409526 0.414106
RF Model with Grid Search¶
# Choose the type of classifier.
glove_rf_tuned = RandomForestClassifier(class_weight= "balanced",random_state=1,bootstrap=True)
parameters = {
'max_depth': list(np.arange(5,10,2)),
'n_estimators': np.arange(50,110,25),
'max_features': [0.3,0.4]
}
# Run the grid search
grid_obj = GridSearchCV(glove_rf_tuned, parameters, scoring='recall',cv=5,n_jobs=-1)
grid_obj = grid_obj.fit(X_train_glove, y_train_glove)
# Set the clf to the best combination of parameters
glove_rf_tuned = grid_obj.best_estimator_
# Fit the best algorithm to the data.
glove_rf_tuned.fit(X_train_glove, y_train_glove)
RandomForestClassifier(class_weight='balanced', max_depth=5, max_features=0.3,
n_estimators=50, random_state=1)In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook. On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
RandomForestClassifier(class_weight='balanced', max_depth=5, max_features=0.3,
n_estimators=50, random_state=1)# Predicting on train data
y_pred_train_tuned_gl = glove_rf_tuned.predict(X_train_glove)
# Predicting on test data
y_pred_test_tuned_gl = glove_rf_tuned.predict(X_test_glove)
plot_confusion_matrix(y_train_glove, y_pred_train_tuned_gl)
#Calculating different metrics on training data
glove_tuned_train=model_performance_classification_sklearn(glove_rf_tuned, X_train_glove,y_train_glove)
print("Training performance:\n", glove_tuned_train)
Training performance:
Accuracy Recall Precision F1
0 0.981627 0.981627 0.981863 0.981583
print(classification_report(y_train_glove, y_pred_train_tuned_gl))
precision recall f1-score support
-1 0.98 0.98 0.98 127
0 0.99 0.96 0.98 127
1 0.97 1.00 0.98 127
accuracy 0.98 381
macro avg 0.98 0.98 0.98 381
weighted avg 0.98 0.98 0.98 381
plot_confusion_matrix(y_test_glove, y_pred_test_tuned_gl)
#Calculating different metrics on testing data
glove_tuned_test=model_performance_classification_sklearn(glove_rf_tuned, X_test_glove,y_test_glove)
print("Testing performance:\n", glove_tuned_test)
Testing performance:
Accuracy Recall Precision F1
0 0.397727 0.397727 0.412581 0.401627
print(classification_report(y_test_glove, y_pred_test_tuned_gl))
precision recall f1-score support
-1 0.39 0.44 0.42 25
0 0.49 0.40 0.44 43
1 0.28 0.35 0.31 20
accuracy 0.40 88
macro avg 0.39 0.40 0.39 88
weighted avg 0.41 0.40 0.40 88
Random Forest Model with Sentence Transformer¶
#Split the data into Train and Test
X = embedding_matrix
y = data["Label"]
X_train_transformer, X_test_transformer, y_train_transformer, y_test_transformer = train_test_split(X, y, test_size=0.25, random_state=42, stratify=y)
# Try reducing to ~80 dimensions to help with overfitting
pca = PCA(n_components=80, random_state=1)
X_train_transformer = pca.fit_transform(X_train_transformer)
X_test_transformer = pca.transform(X_test_transformer)
# Apply SMOTE to balance the training data
smote = SMOTE(random_state=42)
X_train_transformer, y_train_transformer = smote.fit_resample(X_train_transformer, y_train_transformer)
rf_transformer = RandomForestClassifier(
class_weight="balanced",
random_state=1,
bootstrap=True,
max_depth=4, # Start with shallow trees
min_samples_split=5, # Require more samples per split
min_samples_leaf=4, # Require more samples per leaf
n_estimators=100, # More trees for stability
max_features='sqrt' # Restrict features per split
)
print(X_train_transformer.shape, X_test_transformer.shape)
(381, 80) (88, 80)
print(y_train_transformer.shape, y_test_transformer.shape)
(381,) (88,)
RF Base Model¶
# Building the model
rf_transformer = RandomForestClassifier(n_estimators = 100, max_depth = 7, random_state = 42)
# Fitting on train data
rf_transformer.fit(X_train_transformer, y_train_transformer)
RandomForestClassifier(max_depth=7, random_state=42)In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
RandomForestClassifier(max_depth=7, random_state=42)
# Predicting on train data
y_pred_train_transformer = rf_transformer.predict(X_train_transformer)
# Predicting on test data
y_pred_test_transformer = rf_transformer.predict(X_test_transformer)
#Confusion Matrix for Train
plot_confusion_matrix(y_train_transformer, y_pred_train_transformer)
#Calculating different metrics on training data
transformer_base_train=model_performance_classification_sklearn(rf_transformer, X_train_transformer,y_train_transformer)
print("Training performance:\n", transformer_base_train)
Training performance:
Accuracy Recall Precision F1
0 1.0 1.0 1.0 1.0
print(classification_report(y_train_transformer, y_pred_train_transformer))
precision recall f1-score support
-1 1.00 1.00 1.00 127
0 1.00 1.00 1.00 127
1 1.00 1.00 1.00 127
accuracy 1.00 381
macro avg 1.00 1.00 1.00 381
weighted avg 1.00 1.00 1.00 381
#Confusion Matrix for Test
plot_confusion_matrix(y_test_transformer, y_pred_test_transformer)
#Calculating different metrics on testing data
transformer_base_test=model_performance_classification_sklearn(rf_transformer, X_test_transformer,y_test_transformer)
print("Testing performance:\n", transformer_base_test)
Testing performance:
Accuracy Recall Precision F1
0 0.511364 0.511364 0.509144 0.509204
print(classification_report(y_test_transformer, y_pred_test_transformer))
precision recall f1-score support
-1 0.52 0.48 0.50 25
0 0.53 0.58 0.56 43
1 0.44 0.40 0.42 20
accuracy 0.51 88
macro avg 0.50 0.49 0.49 88
weighted avg 0.51 0.51 0.51 88
RF model with Grid Search¶
from sklearn.model_selection import StratifiedKFold
# Choose the type of classifier.
rf_transformer_tuned = RandomForestClassifier(class_weight= "balanced",random_state=1,bootstrap=True, ccp_alpha=0.01)
# Use stratified k-fold with more folds given small dataset
cv = StratifiedKFold(n_splits=10, shuffle=True, random_state=1)
parameters = {
'max_depth': [3, 4, 5],
'min_samples_split': [4, 5, 6],
'min_samples_leaf': [3, 4, 5],
'n_estimators': [100],
'max_features': ['sqrt', 0.2, 0.3]
}
#parameters = {
# 'max_depth': list(np.arange(5,10,2)),
# 'n_estimators': np.arange(50,110,25),
# 'max_features': [0.3,0.4]
#}
# Run the grid search
grid_obj = GridSearchCV(rf_transformer_tuned, parameters, scoring='recall',cv=cv,n_jobs=-1)
grid_obj = grid_obj.fit(X_train_transformer, y_train_transformer)
# Set the clf to the best combination of parameters
rf_transformer_tuned = grid_obj.best_estimator_
# Fit the best algorithm to the data.
rf_transformer_tuned.fit(X_train_transformer, y_train_transformer)
RandomForestClassifier(ccp_alpha=0.01, class_weight='balanced', max_depth=3,
min_samples_leaf=3, min_samples_split=4, random_state=1)In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook. On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
RandomForestClassifier(ccp_alpha=0.01, class_weight='balanced', max_depth=3,
min_samples_leaf=3, min_samples_split=4, random_state=1)# Predicting on train data
y_pred_train_tuned_transformer = rf_transformer_tuned.predict(X_train_transformer)
# Predicting on test data
y_pred_test_tuned_transformer = rf_transformer_tuned.predict(X_test_transformer)
#Confusion Matrix for Train
plot_confusion_matrix(y_train_transformer, y_pred_train_tuned_transformer)
#Calculating different metrics on training data
transformer_tuned_train=model_performance_classification_sklearn(rf_transformer_tuned, X_train_transformer,y_train_transformer)
print("Training performance:\n", transformer_tuned_train)
Training performance:
Accuracy Recall Precision F1
0 0.913386 0.913386 0.916088 0.913113
print(classification_report(y_train_transformer, y_pred_train_tuned_transformer))
precision recall f1-score support
-1 0.90 0.90 0.90 127
0 0.96 0.87 0.91 127
1 0.89 0.98 0.93 127
accuracy 0.91 381
macro avg 0.92 0.91 0.91 381
weighted avg 0.92 0.91 0.91 381
#Confusion Matrix for Test
plot_confusion_matrix(y_test_transformer, y_pred_test_tuned_transformer)
#Calculating different metrics on testing data
transformer_tuned_test=model_performance_classification_sklearn(rf_transformer_tuned, X_test_transformer,y_test_transformer)
print("Testing performance:\n", transformer_tuned_test)
Testing performance:
Accuracy Recall Precision F1
0 0.420455 0.420455 0.442859 0.425578
print(classification_report(y_test_transformer, y_pred_test_tuned_transformer))
precision recall f1-score support
-1 0.40 0.40 0.40 25
0 0.53 0.42 0.47 43
1 0.31 0.45 0.37 20
accuracy 0.42 88
macro avg 0.41 0.42 0.41 88
weighted avg 0.44 0.42 0.43 88
Weekly News Summarization using LLM API¶
Instructions¶
Note:
The model is expected to summarize the news from the week by identifying the top three positive and negative events that are most likely to impact the price of the stock.
As an output, the model is expected to return a JSON containing two keys, one for Positive Events and one for Negative Events.
For the project, we need to define the prompt to be fed to the LLM to help it understand the task to perform. The following should be the components of the prompt:
- Role: Specifies the role the LLM will be taking up to perform the specified task, along with any specific details regarding the role
- Example:
You are an expert data analyst specializing in news content analysis.
- Task: Specifies the task to be performed and outlines what needs to be accomplished, clearly defining the objective
- Example:
Analyze the provided news headline and return the main topics contained within it.
- Instructions: Provides detailed guidelines on how to perform the task, which includes steps, rules, and criteria to ensure the task is executed correctly
- Example:
Instructions:
1. Read the news headline carefully.
2. Identify the main subjects or entities mentioned in the headline.
3. Determine the key events or actions described in the headline.
4. Extract relevant keywords that represent the topics.
5. List the topics in a concise manner.
- Output Format: Specifies the format in which the final response should be structured, ensuring consistency and clarity in the generated output
- Example:
Return the output in JSON format with keys as the topic number and values as the actual topic.
Full Prompt Example:
You are an expert data analyst specializing in news content analysis.
Task: Analyze the provided news headline and return the main topics contained within it.
Instructions:
1. Read the news headline carefully.
2. Identify the main subjects or entities mentioned in the headline.
3. Determine the key events or actions described in the headline.
4. Extract relevant keywords that represent the topics.
5. List the topics in a concise manner.
Return the output in JSON format with keys as the topic number and values as the actual topic.
Sample Output:
{"1": "Politics", "2": "Economy", "3": "Health" }
Note: Attempting to use a locally downloaded LLM proved impractically slow to process results, so instead using an API to access the LLM, in thise case OpenAI's gpt-4o model.
Setup OpenAI API¶
#Retrieve API Key
os.environ["OPENAI_API_KEY"] = userdata.get('openai_key')
from openai import OpenAI
#Initialize the API client
client = OpenAI(
api_key=os.environ.get("OPENAI_API_KEY"), # This is the default and can be omitted
)
#Perform a basic check to very API is Working
chat_completion = client.chat.completions.create(
messages=[
{
"role": "user",
"content": "Say this is a test",
}
],
model="gpt-4o",
)
print(chat_completion.choices[0].message.content)
This is a test. How can I assist you further?
Load the data and aggregate weekly¶
from google.colab import drive
drive.mount('/content/drive')
Mounted at /content/drive
# loading data into a pandas dataframe
df = pd.read_csv("/content/drive/MyDrive/Personal/UT Austin/Project 6 - NLP/stock_news.csv")
data = df.copy()
data["Date"] = pd.to_datetime(data['Date']) # Convert the 'Date' column to datetime format.
# Group the data by week using the 'Date' column.
weekly_grouped = data.groupby(pd.Grouper(key='Date', freq='W'))
weekly_grouped = weekly_grouped.agg(
{
'News': lambda x: ' || '.join(x) # Join the news values with ' || ' separator.
}
).reset_index()
print(weekly_grouped.shape)
(18, 2)
#Verify columns present
weekly_grouped.columns
Index(['Date', 'News'], dtype='object')
#Show row and column count
weekly_grouped.shape
(18, 2)
#Values for the Date column
weekly_grouped['Date']
| Date | |
|---|---|
| 0 | 2019-01-06 |
| 1 | 2019-01-13 |
| 2 | 2019-01-20 |
| 3 | 2019-01-27 |
| 4 | 2019-02-03 |
| 5 | 2019-02-10 |
| 6 | 2019-02-17 |
| 7 | 2019-02-24 |
| 8 | 2019-03-03 |
| 9 | 2019-03-10 |
| 10 | 2019-03-17 |
| 11 | 2019-03-24 |
| 12 | 2019-03-31 |
| 13 | 2019-04-07 |
| 14 | 2019-04-14 |
| 15 | 2019-04-21 |
| 16 | 2019-04-28 |
| 17 | 2019-05-05 |
#Preview first two rows
weekly_grouped.head(2)
| Date | News | |
|---|---|---|
| 0 | 2019-01-06 | The tech sector experienced a significant decline in the aftermarket following Apple's Q1 revenue warning. Notable suppliers, including Skyworks, Broadcom, Lumentum, Qorvo, and TSMC, saw their stocks drop in response to Apple's downward revision of its revenue expectations for the quarter, previously announced in January. || Apple lowered its fiscal Q1 revenue guidance to $84 billion from earlier estimates of $89-$93 billion due to weaker than expected iPhone sales. The announcement caused a significant drop in Apple's stock price and negatively impacted related suppliers, leading to broader market declines for tech indices such as Nasdaq 10 || Apple cut its fiscal first quarter revenue forecast from $89-$93 billion to $84 billion due to weaker demand in China and fewer iPhone upgrades. CEO Tim Cook also mentioned constrained sales of Airpods and Macbooks. Apple's shares fell 8.5% in post market trading, while Asian suppliers like Hon || This news article reports that yields on long-dated U.S. Treasury securities hit their lowest levels in nearly a year on January 2, 2019, due to concerns about the health of the global economy following weak economic data from China and Europe, as well as the partial U.S. government shutdown. Apple || Apple's revenue warning led to a decline in USD JPY pair and a gain in Japanese yen, as investors sought safety in the highly liquid currency. Apple's underperformance in Q1, with forecasted revenue of $84 billion compared to analyst expectations of $91.5 billion, triggered risk aversion mood in markets || Apple CEO Tim Cook discussed the company's Q1 warning on CNBC, attributing US-China trade tensions as a factor. Despite not mentioning iPhone unit sales specifically, Cook indicated Apple may comment on them again. Services revenue is projected to exceed $10.8 billion in Q1. Cook also addressed the lack of || Roku Inc has announced plans to offer premium video channels on a subscription basis through its free streaming service, The Roku Channel. Partners include CBS Corp's Showtime, Lionsgate's Starz, and Viacom Inc's Noggin. This model follows Amazon's successful Channels business, which generated an estimated || Wall Street saw modest gains on Wednesday but were threatened by fears of a global economic slowdown following Apple's shocking revenue forecast cut, blaming weak demand in China. The tech giant's suppliers and S&P 500 futures also suffered losses. Reports of decelerating factory activity in China and the euro zone || Apple's fiscal first quarter revenue came in below analysts' estimates at around $84 billion, a significant drop from the forecasted range of $89-$93 billion. The tech giant attributed the shortfall to lower iPhone revenue and upgrades, as well as weakness in emerging markets. Several brokerages had already reduced their production estimates || Apple Inc. lowered its quarterly sales forecast for the fiscal first quarter, underperforming analysts' expectations due to slowing Chinese economy and trade tensions. The news sent Apple shares tumbling and affected Asia-listed suppliers like Hon Hai Precision Industry Co Ltd, Taiwan Semiconductor Manufacturing Company, and LG Innot || The Australian dollar experienced significant volatility on Thursday, plunging to multi-year lows against major currencies due to automated selling, liquidity issues, and a drought of trades. The largest intra-day falls in the Aussie's history occurred amid violent movements in AUD/JPY and AUD/ || In early Asian trading on Thursday, the Japanese yen surged as the U.S. dollar and Australian dollar collapsed in thin markets due to massive stop loss sales triggered by Apple's earnings warning of sluggish iPhone sales in China and risk aversion. The yen reached its lowest levels against the U.S. dollar since March || The dollar fell from above 109 to 106.67 after Apple's revenue warning, while the 10-year Treasury yield also dropped to 2.61%. This followed money flowing into US government paper. Apple's shares and U.S. stock index futures declined, with the NAS || RBC Capital maintains its bullish stance on Apple, keeping its Outperform rating and $220 price target. However, analyst Amit Daryanani warns of ongoing iPhone demand concerns, which could impact pricing power and segmentation efforts if severe. He suggests potential capital allocation adjustments if the stock underperforms for several quarters || Oil prices dropped on Thursday as investor sentiment remained affected by China's economic slowdown and turmoil in stock and currency markets. US WTI Crude Oil fell by $2.10 to $45.56 a barrel, while International Brent Oil was down $1.20 at $54.26 || In this news article, investors' concerns about a slowing Chinese and global economy, amplified by Apple's revenue warning, led to a significant surge in the Japanese yen. The yen reached its biggest one-day rise in 20 months, with gains of over 4% versus the dollar. This trend was driven by automated || In Asia, gold prices rose to over six-month highs on concerns of a global economic slowdown and stock market volatility. Apple lowered its revenue forecast for the first quarter, leading Asian stocks to decline and safe haven assets like gold and Japanese yen to gain. Data showed weakened factory activity in Asia, particularly China, adding to || Fears of a global economic slowdown led to a decline in the US dollar on Thursday, as the yen gained ground due to its status as a safe haven currency. The USD index slipped below 96, and USD JPY dropped to 107.61, while the yen strengthened by 4.4%. || In Thursday trading, long-term US Treasury yields dropped significantly below 2.6%, reaching levels not seen in over a year, as investors shifted funds from stocks to bonds following Apple's warning of decreased revenue due to emerging markets and China's impact on corporate profits, with the White House advisor adding to concerns of earnings down || Gold prices have reached their highest level since mid-June, with the yellow metal hitting $1,291.40 per ounce due to investor concerns over a slowing economy and Apple's bearish revenue outlook. Saxo Bank analyst Ole Hansen predicts gold may reach $1,300 sooner || Wedbush analyst Daniel Ives lowered his price target for Apple from $275 to $200 due to concerns over potential iPhone sales stagnation, with an estimated 750 million active iPhones worldwide that could cease growing or even decline. He maintains an Outperform rating and remains bullish on the long || Oil prices rebounded on Thursday due to dollar weakness, signs of output cuts by Saudi Arabia, and weaker fuel oil margins leading Riyadh to lower February prices for heavier crude grades sold to Asia. The Organization of the Petroleum Exporting Countries (OPEC) led by Saudi Arabia and other producers || This news article reports on the impact of Apple's Q1 revenue warning on several tech and biotech stocks. Sesen Bio (SESN) and Prana Biotechnology (PRAN) saw their stock prices drop by 28% and 11%, respectively, following the announcement. Mellanox Technologies (ML || Gold prices reached within $5 of $1,300 on Thursday as weak stock markets and a slumping dollar drove investors towards safe-haven assets. The U.S. stock market fell about 2%, with Apple's rare profit warning adding to investor unease. COMEX gold futures settled at $1 || The FDIC Chair, Jelena McWilliams, expressed no concern over market volatility affecting the U.S banking system due to banks' ample capital. She also mentioned a review of the CAMELS rating system used to evaluate bank health for potential inconsistencies and concerns regarding forum shopping. This review comes from industry || Apple cut its quarterly revenue forecast for the first time in over 15 years due to weak iPhone sales in China, representing around 20% of Apple's revenue. This marks a significant downturn during Tim Cook's tenure and reflects broader economic concerns in China exacerbated by trade tensions with the US. U || Goldman analyst Rod Hall lowered his price target for Apple from $182 to $140, citing potential risks to the tech giant's 2019 numbers due to uncertainties in Chinese demand. He reduced his revenue estimate for the year by $6 billion and EPS forecast by $1.54 || Delta Air Lines lowered its fourth-quarter revenue growth forecast to a range of 3% from the previous estimate of 3% to 5%. Earnings per share are now expected to be $1.25 to $1.30. The slower pace of improvement in late December was unexpected, and Delta cited this as || Apple's profit warning has significantly impacted the stock market and changed the outlook for interest rates. The chance of a rate cut in May has increased to 15-16% from just 3%, according to Investing com's Fed Rate Monitor Tool. There is even a 1% chance of two cuts in May. || The White House advisor, Kevin Hassett, stated that a decline in Chinese economic growth would negatively impact U.S. firm profits but recover once a trade deal is reached between Washington and Beijing. He also noted that Asian economies, including China, have been experiencing significant slowdowns since last spring due to U.S. tariffs || The White House economic adviser, Kevin Hassett, warned that more companies could face earnings downgrades due to ongoing trade negotiations between the U.S. and China, leading to a decline in oil prices on Thursday. WTI crude fell 44 cents to $44.97 a barrel, while Brent crude inched || Japanese stocks suffered significant losses on the first trading day of 2019, with the Nikkei 225 and Topix indices both falling over 3 percent. Apple's revenue forecast cut, citing weak iPhone sales in China, triggered global growth concerns and sent technology shares tumbling. The S&P 50 || Investors withdrew a record $98 billion from U.S. stock funds in December, with fears of aggressive monetary policy and an economic slowdown driving risk reduction. The S&P 500 fell 9% last month, with some seeing declines as a buying opportunity. Apple's warning of weak iPhone sales added || Apple's Q1 revenue guidance cut, resulting from weaker demand in China, led to an estimated $3.8 billion paper loss for Berkshire Hathaway due to its $252 million stake in Apple. This news, coupled with broad market declines, caused a significant $21.4 billion decrease in Berk || This news article reports that a cybersecurity researcher, Wish Wu, planned to present at the Black Hat Asia hacking conference on how to bypass Apple's Face ID biometric security on iPhones. However, his employer, Ant Financial, which operates Alipay and uses facial recognition technologies including Face ID, asked him to withdraw || OPEC's production cuts faced uncertainty as oil prices were influenced by volatile stock markets, specifically due to Apple's lowered revenue forecast and global economic slowdown fears. US WTI and Brent crude both saw gains, but these were checked by stock market declines. Shale production is expected to continue impacting the oil market in || Warren Buffett's Berkshire Hathaway suffered significant losses in the fourth quarter due to declines in Apple, its largest common stock investment. Apple cut its revenue forecast, causing a 5-6% decrease in Berkshire's Class A shares. The decline resulted in potential unrealized investment losses and could push Berk || This news article reports that on Thursday, the two-year Treasury note yield dropped below the Federal Reserve's effective rate for the first time since 2008. The market move suggests investors believe the Fed will not be able to continue tightening monetary policy. The drop in yields was attributed to a significant decline in U.S || The U.S. and China will hold their first face-to-face trade talks since agreeing to a 90-day truce in their trade war last month. Deputy U.S. Trade Representative Jeffrey Gerrish will lead the U.S. delegation for negotiations on Jan. 7 and 8, || Investors bought gold in large quantities due to concerns over a global economic slowdown, increased uncertainty in the stock market, and potential Fed rate hikes. The precious metal reached its highest price since June, with gold ETF holdings also seeing significant increases. Factors contributing to this demand include economic downturn, central bank policy mistakes, and || Delta Air Lines Inc reported lower-than-expected fourth quarter unit revenue growth, citing weaker than anticipated late bookings and increased competition. The carrier now expects total revenue per available seat mile to rise about 3 percent in the period, down from its earlier forecast of 3.5 percent growth. Fuel prices are also expected to || U.S. stocks experienced significant declines on Thursday as the S&P 500 dropped over 2%, the Dow Jones Industrial Average fell nearly 3%, and the Nasdaq Composite lost approximately 3% following a warning of weak revenue from Apple and indications of slowing U.S. factory activity, raising concerns || President Trump expressed optimism over potential trade talks with China, citing China's current economic weakness as a potential advantage for the US. This sentiment was echoed by recent reports of weakened demand for Apple iPhones in China, raising concerns about the overall health of the Chinese economy. The White House is expected to take a strong stance in || Qualcomm secured a court order in Germany banning the sale of some iPhone models due to patent infringement, leading Apple to potentially remove these devices from its stores. However, third-party resellers like Gravis continue selling the affected iPhones. This is the third major effort by Qualcomm to ban Apple's iPhones glob || Oil prices rose on Friday in Asia as China confirmed trade talks with the U.S., with WTI gaining 0.7% to $47.48 and Brent increasing 0.7% to $56.38 a barrel. The gains came after China's Commerce Ministry announced that deputy U.S. Trade || Gold prices surged past the psychologically significant level of $1,300 per ounce in Asia on Friday due to growing concerns over a potential global economic downturn. The rise in gold was attributed to weak PMI data from China and Apple's reduced quarterly sales forecast. Investors viewed gold as a safe haven asset amidst || In an internal memo, Huawei's Chen Lifang reprimanded two employees for sending a New Year greeting on the company's official Twitter account using an iPhone instead of a Huawei device. The incident caused damage to the brand and was described as a "blunder" in the memo. The mistake occurred due to || This news article reports on the positive impact of trade war talks between Beijing and Washington on European stock markets, specifically sectors sensitive to the trade war such as carmakers, industrials, mining companies, and banking. Stocks rallied with mining companies leading the gains due to copper price recovery. Bayer shares climbed despite a potential ruling restricting || Amazon has sold over 100 million devices with its Alexa digital assistant, according to The Verge. The company is cautious about releasing hardware sales figures and did not disclose holiday numbers for the Echo Dot. Over 150 products feature Alexa, and more than 28,000 smart home || The Supreme Court will review Broadcom's appeal in a shareholder lawsuit over the 2015 acquisition of Emulex. The case hinges on whether intent to defraud is required for such lawsuits, and the decision could extend beyond the Broadcom suit. An Emulex investor filed a class action lawsuit || The Chinese central bank announced a fifth reduction in the required reserve ratio (RRR) for banks, freeing up approximately 116.5 billion yuan for new lending. This follows mounting concerns about China's economic health amid slowing domestic demand and U.S. tariffs on exports. Premier Li Keqiang || The stock market rebounded strongly on Friday following positive news about US-China trade talks, a better-than-expected jobs report, and dovish comments from Federal Reserve Chairman Jerome Powell. The Dow Jones Industrial Average rose over 746 points, with the S&P 500 and Nasdaq Com |
| 1 | 2019-01-13 | Sprint and Samsung plan to release 5G smartphones in nine U.S. cities this summer, with Atlanta, Chicago, Dallas, Houston, Kansas City, Los Angeles, New York City, Phoenix, and Washington D.C. being the initial locations. Rival Verizon also announced similar plans for the first half of 20 || AMS, an Austrian tech company listed in Switzerland and a major supplier to Apple, has developed a light and infrared proximity sensor that can be placed behind a smartphone's screen. This allows for a larger display area by reducing the required space for sensors. AMS provides optical sensors for 3D facial recognition features on Apple || Deutsche Bank upgraded Vivendi's Universal Music Group valuation from €20 billion to €29 billion, surpassing the market cap of Vivendi at €28.3 billion. The bank anticipates music streaming revenue to reach €21 billion in 2023 and identifies potential suitors for || Amazon's stock is predicted to surge by over 20% by the end of this year, according to a new report from Pivotal Research. Senior analyst Brian Wieser initiated coverage on the stock with a buy rating and a year-end price target of $1,920. The growth potential for Amazon lies primarily in || AMS, an Austrian sensor specialist, is partnering with Chinese software maker Face to develop new 3D facial recognition features for smartphones. This move comes as AMS aims to reduce its dependence on Apple and boost its battered shares. AMS provides optical sensors for Apple's 3D facial recognition feature on iPhones, || Geely, China's most successful carmaker, forecasts flat sales for 2019 due to economic slowdown and cautious consumers. In 2018, it posted a 20% sales growth, but missed its target of 1.58 million cars by around 5%. Sales dropped 44 || China is making sincere efforts to address U.S. concerns and resolve the ongoing trade war, including lowering taxes on automobile imports and implementing a law banning forced technology transfers. However, Beijing cannot and should not dismantle its governance model as some in Trump's administration have demanded. Former Goldman Sachs China || Stock index futures indicate a slightly lower open on Wall Street Monday, as investors remain cautious amid lack of progress in U.S.-China trade talks and political risks from the ongoing government shutdown. Dow futures were flat, S&P 500 dipped 0.13%, while Nasdaq 10 || Qualcomm, a leading chipmaker, has announced an expansion of its lineup of car computing chips into three tiers - entry-level, Performance, Premiere, and Paramount. This move is aimed at catering to various price points in the automotive market, similar to its smartphone offerings. The company has reported a backlog || The stock market showed minimal changes at the open as investors await trade talks progress between the U.S. and China. The S&P 500 dropped 0.04%, Dow lost 0.23%, but Nasdaq gained 0.2%. The ISM services index, expected to be released at 1 || The article suggests that some economists believe the US economy may have reached its peak growth rate, making the euro a potentially bullish investment. The EUR/USD exchange rate has held steady despite weak Eurozone data due to dollar weakness and stagnant interest rate expectations in Europe. However, concerns over economic growth are emerging due to sell || The Chinese smartphone market, the world's largest, saw a decline of 12-15.5 percent in shipments last year with December experiencing a 17 percent slump, according to China Academy of Information and Communications Technology (CAICT) and market research firm Canalys. This follows a 4 percent drop in ship || Austrian tech firm AT S lowered its revenue growth forecast for 2018/19 due to weak demand from smartphone makers and the automotive industry. The company now anticipates a 3% increase in sales from last year's €991.8 million, down from its previous projection of a 6- || The stock markets in Asia surged during morning trade on Wednesday, following reports of progress in U. S - China trade talks. Negotiators extended talks for a third day and reportedly made strides on purchases of U. S goods and services. However, structural issues such as intellectual property rights remain unresolved. President Trump is eager to strike || Mercedes Benz sold over 2.31 million passenger cars in 2018, making it the top selling premium automotive brand for the third year in a row. However, analysts question how long German manufacturers can dominate the luxury car industry due to the shift towards electric and self-driving cars. Tesla, with || The S&P 500 reached a three-week high on Tuesday, driven by gains in Apple, Amazon, Facebook, and industrial shares. Investors are optimistic about a potential deal between the US and China to end their trade war. The S&P 500 has rallied over 9% since late December || The stock market continued its rally on Tuesday, with the Dow Jones Industrial Average, S&P 500, and Nasdaq Composite all posting gains. Optimism over progress in trade talks between the US and China was a major contributor to the market's positive sentiment, with reports suggesting that both parties are narrowing their differences ahead || Roku's stock dropped by 5% on Tuesday following Citron Research's reversal of its long position, labeling the company as uninvestable. This change in stance came after Apple announced partnerships with Samsung to offer iTunes services on some Samsung TVs, potentially impacting Roku's user base growth. || The Chinese authorities are expected to release a statement following the conclusion of U. S. trade talks in Beijing, with both sides signaling progress toward resolving the conflict that has roiled markets. Chinese Vice Premier Liu He, who is also the chief economic adviser to Chinese President Xi Jinping, made an appearance at the negotiations and is || Xiaomi Co-founder Lei Jun remains optimistic about the future of his smartphone company despite a recent share slump that erased $6 billion in market value. The Chinese tech firm is shifting its focus to the high end and expanding into Europe, while shunning the US market. Xiaomi aims to elevate its Red || The European Commission has launched an investigation into Nike's tax treatment in the Netherlands, expressing concerns that the company may have received an unfair advantage through royalty payment structures. The EU executive has previously probed tax schemes in Belgium, Gibraltar, Luxembourg, Ireland, and the Netherlands, with countries ordered to recover taxes from benefici || Taiwan's Foxconn, a major Apple supplier, reported an 8.3% decline in December revenue to TWD 619.3 billion ($20.1 billion), marking its first monthly revenue dip since February. The fall was due to weak demand for consumer electronics. In 2018, Foxconn || Starting tomorrow, JD.com will offer reduced prices on some Apple iPhone 8 and 8 Plus models by approximately 600 yuan and 800 yuan respectively. These price drops, amounting to a savings of around $92-$130 per unit, are in line with earlier rumors suggesting price redu || Cummins, Inc. (CMI) announced that Pat Ward, its long-term Chief Financial Officer (CFO), will retire after 31 years of service on March 31, 2019. Mark Smith, who has served as Vice President of Financial Operations since 2014, will succeed Ward, || The Federal Reserve Chairman, Jerome Powell, maintained his patient stance on monetary policy but raised concerns about the balance sheet reduction. He indicated that the Fed's balance sheet would be substantially smaller, indicating the continuation of the balance sheet wind down operation. Despite this, Powell reassured investors of a slower pace on interest rate h || Wall Street experienced a decline after the opening bell on Friday, following five consecutive days of gains. The S&P 500 dropped 13 points or 0.54%, with the Dow decreasing 128 points or 0.54%, and the Nasdaq Composite losing 37 points or || Several Chinese retailers, including Alibaba-backed Suning and JD.com, have drastically reduced iPhone prices due to weak sales in China, which prompted Apple's recent revenue warning. Discounts for the latest XR model range from 800 to 1,200 yuan. These price || Green Dot, GDOT, is a bank holding company with a wide distribution network and impressive growth. Its product offerings include bank accounts, debit and credit cards, with a focus on perks. The firm's platform business, "banking as a service," powers offerings for partners such as Apple Pay Cash, Walmart Money || US stock futures declined on Friday as disappointing holiday sales and revenue cuts from various companies raised concerns about a potential recession. The S&P 500, Dow Jones Industrial Average, and Nasdaq 100 fell, with the Fed's possible policy pause and optimistic trade talks failing to offset these negative factors. || Apple's NASDAQ AAPL stock declined by 0.52% in premarket trade Friday due to price cuts of iPhone models in China, but the company is set to launch three new iPhone models this year. Johnson & Johnson's NYSE JNJ stock edged forward after raising prescription drug prices. Starbucks || Apple is reportedly set to release three new iPhone models this year, featuring new camera setups including triple rear cameras for the premium model and dual cameras for the others. The move comes after weak sales, particularly in China, led retailers to cut prices on the XR model. Amid sluggish sales, Apple opted to stick with |
#Export weekly_grouped to xlsx
#weekly_grouped.to_excel("/content/drive/MyDrive/Personal/UT Austin/Project 6 - NLP/weekly_grouped.xlsx", index=False)
Defining the Instruction Prompt and Response Function¶
# defining the instructions for the model
prompt = """
Task: Analyze the provided news headlines that could affect the company's stock price, where headlines are delimited by ||
Instructions:
1. Read the news headlines carefully
2. Identify the top 3 Positive Events
3. Identify the top 3 Negative Events
Return the output in JSON format, strictly using this structure:
{
"Positive Events": [
"Summarized Positive Event 1",
"Summarized Positive Event 2",
"Summarized Positive Event 3"
],
"Negative Events": [
"Summarized Negative Event 1",
"Summarized Negative Event 2",
"Summarized Negative Event 3"
]
}
Do NOT include any extra text outside the JSON structure.
"""
#Create a function to call the API with the instruction prompt and news articles as input
def generate_response(prompt, news):
# Combine instruction prompt and news articles to create a combined prompt
combined_prompt = f"{news}\n{prompt}"
try:
# Generate a response using the OpenAI API
response = client.chat.completions.create(
model="gpt-4o", # Specify the desired model
messages=[
{"role": "system", "content": "You are an expert data analyst specializing in news content analysis."},
{"role": "user", "content": combined_prompt},
],
max_tokens=512,
temperature=0,
top_p=0.95,
)
# Extract the response content
response_text = response.choices[0].message.content.strip()
# Clean up the response text to remove Markdown formatting
response_text = re.sub(r"```json|```", "", response_text).strip()
# Post-process the response to ensure valid JSON
try:
json_response = json.loads(response_text)
return json_response
except json.JSONDecodeError:
print("Warning: Response is not valid JSON.")
return {"error": "Invalid JSON response", "raw_response": response_text}
except client.error.OpenAIError as e:
print(f"OpenAI API error: {e}")
return {"error": "API error", "details": str(e)}
Test the Response Function on Sample Data¶
news = weekly_grouped['News'].iloc[0]
generate_response(prompt, news)
{'Positive Events': ["Roku Inc plans to offer premium video channels through its free streaming service, The Roku Channel, following Amazon's successful Channels business model.",
'The stock market rebounded strongly on Friday following positive news about US-China trade talks, a better-than-expected jobs report, and dovish comments from Federal Reserve Chairman Jerome Powell.',
'Oil prices rose on Friday in Asia as China confirmed trade talks with the U.S., with WTI and Brent both gaining.'],
'Negative Events': ["Apple's Q1 revenue warning led to a significant decline in the tech sector, affecting suppliers like Skyworks, Broadcom, Lumentum, Qorvo, and TSMC.",
"Apple cut its fiscal first quarter revenue forecast due to weaker demand in China and fewer iPhone upgrades, causing an 8.5% drop in Apple's stock price.",
'Goldman analyst Rod Hall lowered his price target for Apple from $182 to $140, citing potential risks due to uncertainties in Chinese demand.']}
Run the Model on the Full Data Set and Store Results in a Dataframe¶
# Initialize a list to store results
results = []
# Loop through each row of weekly_grouped
for _, row in weekly_grouped.iterrows():
date = row['Date']
news_content = row['News']
# Call the API to analyze the news content
analysis_results = generate_response(prompt, news_content)
#print(analysis_results)
# Extract results from the API response
if 'error' not in analysis_results:
pos_events = analysis_results.get('Positive Events', [""] * 3) # Default to empty strings
neg_events = analysis_results.get('Negative Events', [""] * 3) # Default to empty strings
else:
# Handle invalid JSON or errors
pos_events = ["Error"] * 3
neg_events = ["Error"] * 3
# Append the results to the list
results.append({
'Date': date,
'PosEvent1': pos_events[0],
'PosEvent2': pos_events[1],
'PosEvent3': pos_events[2],
'NegEvent1': neg_events[0],
'NegEvent2': neg_events[1],
'NegEvent3': neg_events[2],
})
# Create the results DataFrame from the list
news_analysis_results = pd.DataFrame(results)
news_analysis_results.columns
Index(['Date', 'PosEvent1', 'PosEvent2', 'PosEvent3', 'NegEvent1', 'NegEvent2',
'NegEvent3'],
dtype='object')
#Row and column count of news_analysis_results
print(news_analysis_results.shape)
(18, 7)
# Preview the first 3 rows
print(news_analysis_results.head(3))
Date \
0 2019-01-06
1 2019-01-13
2 2019-01-20
PosEvent1 \
0 Roku Inc plans to offer premium video channels through its free streaming service, The Roku Channel, following Amazon's successful Channels business model.
1 Sprint and Samsung plan to release 5G smartphones in nine U.S. cities, enhancing their market presence.
2 Dialog Semiconductor shares rose 4% as investors praised its resilience amid other Apple suppliers missing targets.
PosEvent2 \
0 The stock market rebounded strongly on Friday following positive news about US-China trade talks, a better-than-expected jobs report, and dovish comments from Federal Reserve Chairman Jerome Powell.
1 Deutsche Bank upgraded Vivendi's Universal Music Group valuation, indicating strong growth potential in music streaming.
2 U.S. stocks rose on Tuesday with technology and internet companies leading gains after Netflix announced a price increase for U.S subscribers.
PosEvent3 \
0 Oil prices rose on Friday in Asia as China confirmed trade talks with the U.S., with WTI and Brent both gaining.
1 The S&P 500 reached a three-week high driven by gains in major tech stocks and optimism over US-China trade talks.
2 Verizon announced that it will offer free Apple Music subscriptions with some of its top tier data plans, deepening its partnership with Apple.
NegEvent1 \
0 Apple's Q1 revenue warning led to a significant decline in the tech sector, affecting suppliers like Skyworks, Broadcom, Lumentum, Qorvo, and TSMC.
1 Roku's stock dropped by 5% after Citron Research labeled the company as uninvestable due to Apple's partnerships with Samsung.
2 U.S. stock market declined on Monday as concerns over a global economic slowdown intensified following unexpected drops in China's exports and imports.
NegEvent2 \
0 Apple cut its fiscal first quarter revenue forecast due to weaker demand in China and fewer iPhone upgrades, causing an 8.5% drop in Apple's stock price.
1 Foxconn reported an 8.3% decline in December revenue due to weak demand for consumer electronics.
2 Apple Inc. shares declined in postmarket trading after announcing it would cut back on hiring for some divisions due to fewer iPhone sales and missing revenue forecasts.
NegEvent3
0 Goldman analyst Rod Hall lowered his price target for Apple from $182 to $140, citing potential risks to the tech giant's 2019 numbers due to uncertainties in Chinese demand.
1 Several Chinese retailers drastically reduced iPhone prices due to weak sales in China, impacting Apple's revenue.
2 Foxconn, Apple's biggest iPhone assembler, has let go around 50,000 contract workers in China earlier than usual this year.
Weekly News Summarization using Locally Downloaded LLM¶
In this section we show how to load an LLM (Mistral), however, we don't actually run this model as the time to process results proved to be unacceptably slow on the compute platform available in Google Colab (T4 GPU). A search on this question yielded the following answer: A Mistral 7B model running very slowly on Google Colab is likely due to the large size of the model, which puts significant demands on the limited memory and processing power available on the free Colab tier, making it impractical to run efficiently without utilizing optimization techniques like model loading strategies or upgrading to a paid Colab plan with more powerful GPUs; essentially, the free Colab environment might not have enough resources to handle the computational needs of a 7 billion parameter model like Mistral 7B.
#Updated recommended by GL for llama_cpp
!CMAKE_ARGS="-DLLAMA_CUBLAS=on" FORCE_CMAKE=1 pip install llama-cpp-python==0.1.85 --force-reinstall --no-cache-dir -q
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 1.8/1.8 MB 40.2 MB/s eta 0:00:00a 0:00:01 Installing build dependencies ... done Getting requirements to build wheel ... done Preparing metadata (pyproject.toml) ... done ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 62.0/62.0 kB 266.0 MB/s eta 0:00:00 ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 45.5/45.5 kB 255.7 MB/s eta 0:00:00 ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 16.4/16.4 MB 266.4 MB/s eta 0:00:00 Building wheel for llama-cpp-python (pyproject.toml) ... done ERROR: pip's dependency resolver does not currently take into account all the packages that are installed. This behaviour is the source of the following dependency conflicts. cupy-cuda12x 12.2.0 requires numpy<1.27,>=1.20, but you have numpy 2.2.1 which is incompatible. gensim 4.3.3 requires numpy<2.0,>=1.18.5, but you have numpy 2.2.1 which is incompatible. langchain 0.3.12 requires numpy<2,>=1.22.4; python_version < "3.12", but you have numpy 2.2.1 which is incompatible. matplotlib 3.8.0 requires numpy<2,>=1.21, but you have numpy 2.2.1 which is incompatible. numba 0.60.0 requires numpy<2.1,>=1.22, but you have numpy 2.2.1 which is incompatible. pytensor 2.26.4 requires numpy<2,>=1.17.0, but you have numpy 2.2.1 which is incompatible. tensorflow 2.17.1 requires numpy<2.0.0,>=1.23.5; python_version <= "3.11", but you have numpy 2.2.1 which is incompatible. thinc 8.2.5 requires numpy<2.0.0,>=1.19.0; python_version >= "3.9", but you have numpy 2.2.1 which is incompatible.
# Function to download the model from the Hugging Face model hub
from huggingface_hub import hf_hub_download
# Importing the Llama class from the llama_cpp module
from llama_cpp import Llama
# Importing the library for data manipulation
import pandas as pd
from tqdm import tqdm # For progress bar related functionalities
tqdm.pandas()
Utility Functions¶
# defining a function to parse the JSON output from the model
def extract_json_data(json_str):
import json
try:
# Find the indices of the opening and closing curly braces
json_start = json_str.find('{')
json_end = json_str.rfind('}')
if json_start != -1 and json_end != -1:
extracted_category = json_str[json_start:json_end + 1] # Extract the JSON object
data_dict = json.loads(extracted_category)
return data_dict
else:
print(f"Warning: JSON object not found in response: {json_str}")
return {}
except json.JSONDecodeError as e:
print(f"Error parsing JSON: {e}")
return {}
#Defining the response function
def response_mistral(prompt, news):
model_output = llm(
f"""
[INST]
{prompt}
News Articles: {news}
[/INST]
""",
max_tokens=32, #Complete the code to set the maximum number of tokens the model should generate for this task.
temperature=0.7, #Determines the randomness of the model's output.
top_p=0.9, #Nucleus sampling controls the diversity of tokens considered during generation.
top_k=50, #Restricts token sampling to the top k tokens based on probability.
echo=False,
)
final_output = model_output["choices"][0]["text"]
return final_output
Loading the model (Mistral)¶
model_name_or_path = "TheBloke/Mistral-7B-Instruct-v0.2-GGUF"
model_basename = "mistral-7b-instruct-v0.2.Q6_K.gguf"
model_path = hf_hub_download(
repo_id=model_name_or_path,
filename=model_basename
)
mistral-7b-instruct-v0.2.Q6_K.gguf: 0%| | 0.00/5.94G [00:00<?, ?B/s]
# Load the model
llm = Llama(
model_path=model_path,
n_threads=4,
n_batch=512,
n_gpu_layers=30,
n_ctx=4096 # Adjust context window as needed
)
AVX = 1 | AVX2 = 1 | AVX512 = 0 | AVX512_VBMI = 0 | AVX512_VNNI = 0 | FMA = 1 | NEON = 0 | ARM_FMA = 0 | F16C = 1 | FP16_VA = 0 | WASM_SIMD = 0 | BLAS = 1 | SSE3 = 1 | SSSE3 = 1 | VSX = 0 |
Conclusions and Recommendations¶
In the Sentiment Analysis Exercise, all 3 approaches, Word2Vec, GloVe, and Sentence Transformer tended to lead to overfitting. Based on some research I did the impression seems to be that having only 349 samples along with the very high dimensionality of the encodings (100 features for Word2Vec and GloVe and 384 features for all-MiniLM-L6-v2) is problematic for machine learning models like Random Forests, namely severe overfitting. There were also struggles with getting the GridSearch to achieve better results than just the base Random Forest models.
In the Summarization Exercise, very good results and performance were achieved using the Open AI API with gpt-4o model. Unfortunately I was not able to test out using a locally downloaded LLM from HuggingFace, because in each case I could not get reasonable reponse times. E.g., when trying to use the Llama 7B model it was taking 50 minutes to get a response from just a simple test prompt. Not sure if this is something to do with limitations in the Google Colab environment even when using GPU (T4).