In this project, sentiment analysis was done using natural language processing on the online reviews prevalant for various items on amazon,yelp and imdb which were lablelled. Using the spacy package of python to preprocess the data before, each individual review has been tokenized, lemmatized, filtered for stop words and vectorized inorder to prepare the data viable for the machine learning model. A pipeline was created which vectorized the preprocessed data using count vectorization or tfidf vectorizer, which is then split into training and testing datasets and were then used to train the machine learning model (support vector machine) and evaluate. **conclusion
Natural language processing (NLP) is an area of computer science and artificial intelligence concerned with the interactions between computers and human (natural) languages, in particular how to program computers to process and analyze large amounts of natural language data.
Human language is astoundingly complex and diverse. We express ourselves in infinite ways, both verbally and in writing. Not only are there hundreds of languages and dialects, but within each language is a unique set of grammar and syntax rules, terms and slang. When we write, we often misspell or abbreviate words, or omit punctuation. When we speak, we have regional accents, and we mumble, stutter and borrow terms from other languages.
While supervised and unsupervised learning, and specifically deep learning, are now widely used for modeling human language, there’s also a need for syntactic and semantic understanding and domain expertise that are not necessarily present in these machine learning approaches. NLP is important because it helps resolve ambiguity in language and adds useful numeric structure to the data for many downstream applications, such as speech recognition or text analytics.
Challenges in natural language processing frequently involve speech recognition, natural language understanding, and natural language generation.
Natural language processing includes many different techniques for interpreting human language, ranging from statistical and machine learning methods to rules-based and algorithmic approaches. We need a broad array of approaches because the text- and voice-based data varies widely, as do the practical applications.
Basic NLP tasks include tokenization and parsing, lemmatization/stemming, part-of-speech tagging, language detection and identification of semantic relationships. If you ever diagramed sentences in grade school, you’ve done these tasks manually before.
In general terms, NLP tasks break down language into shorter, elemental pieces, try to understand relationships between the pieces and explore how the pieces work together to create meaning.
These underlying tasks are often used in higher-level NLP capabilities, such as:
Content categorization A linguistic-based document summary, including search and indexing, content alerts and duplication detection.
Topic discovery and modeling Accurately capture the meaning and themes in text collections, and apply advanced analytics to text, like optimization and forecasting.
Contextual extraction Automatically pull structured information from text-based sources.
Sentiment analysis Identifying the mood or subjective opinions within large amounts of text, including average sentiment and opinion mining.
Speech-to-text and text-to-speech conversion Transforming voice commands into written text, and vice versa.
Document summarization Automatically generating synopses of large bodies of text.
Machine translation Automatic translation of text or speech from one language to another.
Spacy is written in cython language, (C extension of Python designed to give C like performance to the python program). Hence is a quite fast library. spaCy provides a concise API to access its methods and properties governed by trained machine (and deep) learning models.
Implementation of spacy and access to different properties is initiated by creating pipelines. A pipeline is created by loading the models. There are different type of models provided in the package which contains the information about language – vocabularies, trained vectors, syntaxes and entities.
These pipelines outputs a wide range of document properties such as – tokens, token’s reference index, part of speech tags, entities, vectors, sentiment, vocabulary etc. Let’s explore some of these properties.
Tokenization: Every spaCy document is tokenized into sentences and further into tokens which can be accessed by iterating the document.
Part of Speech Tagging: Part-of-speech tags are the properties of the word that are defined by the usage of the word in the grammatically correct sentence. These tags can be used as the text features in information filtering, statistical models, and rule based parsing.
Entity Detection Spacy consists of a fast entity recognition model which is capable of identifying entitiy phrases from the document. Entities can be of different types, such as – person, location, organization, dates, numerals, etc. These entities can be accessed through “.ents” property.
Dependency Parsing One of the most powerful feature of spacy is the extremely fast and accurate syntactic dependency parser which can be accessed via lightweight API. The parser can also be used for sentence boundary detection and phrase chunking. The relations can be accessed by the properties “.children” , “.root”, “.ancestor” etc.
Noun Phrases Dependency trees can also be used to generate noun phrases
Word to Vectors Integration Spacy also provides inbuilt integration of dense, real valued vectors representing distributional similarity information. It uses GloVe vectors to generate vectors. GloVe is an unsupervised learning algorithm for obtaining vector representations for words.
The data set contains about 1000 online reviews each for various items on Amazon, Yelp and IMDB and of these reviews about 500 were labelled positive and 500 were labelled negative reviews. For each company, the data was given the text format which are needed to be added to a dataframe
import spacy
import pandas as pd
import numpy as np
import sklearn as sk
import matplotlib.pyplot as plt
import seaborn as sns
Since the data for each of the company is seperately stored into txt files. Each of these files were seperately imported and joined with key fields.
data_yelp = pd.read_table('yelp_labelled.txt')
data_amazon = pd.read_table('amazon_cells_labelled.txt')
data_imdb = pd.read_table('imdb_labelled.txt')
# Joining the tables
combined_col= [data_amazon,data_imdb,data_yelp]
# To observe how the data in each individual dataset is structured
print(data_amazon.columns)
NOTE From the above output, it is evident that the tables do not have headers and the column content comprises of the review and then the label indicating 0 when negative review and 1 when positive review.
# In order to add headers for columns in each dataset
for colname in combined_col:
colname.columns = ["Review","Label"]
for colname in combined_col:
print(colname.columns)
# In order to recognize which dataset belonged to which company, a 'Company' column is added as a key
company = [ "Amazon", "imdb", "yelp"]
comb_data = pd.concat(combined_col,keys = company)
# Exploring the structure of the new data frame
print(comb_data.shape)
comb_data.head()
comb_data.to_csv("Sentiment_Analysis_Dataset")
print(comb_data.columns)
print(comb_data.isnull().sum())
In this stage, Spacy package of python is used to lemmatize and remove stop words from the obtained dataset.
import spacy
import en_core_web_sm
from spacy.lang.en.stop_words import STOP_WORDS
nlp = en_core_web_sm.load()
#nlp = spacy.load('en')
# To build a list of stop words for filtering
stopwords = list(STOP_WORDS)
print(stopwords)
Thus, the stop words have been enlisted
The data is initially split into test and training datasets prior to feeding into the machine learning pipeline. Then, a class object was defined as 'sent_predict' is used as the first step of the pipeline which would inherit from the TransformerMixin package and perform the cleaning of data. The second method of the pipeline is to vectorize the cleaned data.
Tokenized words needs to be lemmatized and filtered for pronouns, stopwords and punctuations using the defined method 'my_tokenizer' For that purpose count vectorizeor and tfidfVectorizer both have been tried subsequently to decide which is better.
Then the third step of the pipeline is the defining of the classifier. In this case, Linear Support Vector Machine classifier was chosen. Other methods could be explored in the furture
import string
punctuations = string.punctuation
# Creating a Spacy Parser
from spacy.lang.en import English
parser = English()
def my_tokenizer(sentence):
mytokens = parser(sentence)
mytokens = [ word.lemma_.lower().strip() if word.lemma_ != "-PRON-" else word.lower_ for word in mytokens ]
mytokens = [ word for word in mytokens if word not in stopwords and word not in punctuations ]
return mytokens
# ML Packages
from sklearn.feature_extraction.text import CountVectorizer,TfidfVectorizer
from sklearn.metrics import accuracy_score
from sklearn.base import TransformerMixin
from sklearn.pipeline import Pipeline
from sklearn.svm import LinearSVC
#Custom transformer using spaCy
class predictors(TransformerMixin):
def transform(self, X, **transform_params):
return [clean_text(text) for text in X]
def fit(self, X, y, **fit_params):
return self
def get_params(self, deep=True):
return {}
# Basic function to clean the text
def clean_text(text):
return text.strip().lower()
# Vectorization
vectorizer = CountVectorizer(tokenizer = spacy_tokenizer, ngram_range=(1,1))
classifier = LinearSVC()
# Using Tfidf
tfvectorizer = TfidfVectorizer(tokenizer = spacy_tokenizer)
# Splitting Data Set
from sklearn.model_selection import train_test_split
# Features and Labels
X = comb_data['Review']
ylabels = comb_data['Label']
X_train, X_test, y_train, y_test = train_test_split(X, ylabels, test_size=0.2, random_state=42)
# Create the pipeline to clean, tokenize, vectorize, and classify using"Count Vectorizor"
pipe_countvect = Pipeline([("cleaner", predictors()),
('vectorizer', vectorizer),
('classifier', classifier)])
# Fit our data
pipe_countvect.fit(X_train,y_train)
# Predicting with a test dataset
sample_prediction = pipe_countvect.predict(X_test)
# Prediction Results
# 1 = Positive review
# 0 = Negative review
for (sample,pred) in zip(X_test,sample_prediction):
print(sample,"Prediction=>",pred)
# Accuracy
print("Accuracy: ",pipe_countvect.score(X_test,y_test))
print("Accuracy: ",pipe_countvect.score(X_test,sample_prediction))
# Accuracy
print("Accuracy: ",pipe_countvect.score(X_train,y_train))
# Another random review
pipe.predict(["This was a great movie"])
example = ["I do enjoy my job",
"What a poor product!,I will have to get a new one",
"I feel amazing!"]
pipe.predict(example)
As observed thee model was about 79.4 % accurate.