Preprocessed 50,000 movie reviews, achieving 95% text consistency through text normalization and cleaning techniques.
Applied Bag of Words and TF-IDF vectorization, resulting in 10,000-dimensional feature vectors.
Built and evaluated Naive Bayes, Logistic Regression, and SVM models with an average accuracy of 85%.
Split dataset into 70% training and 30% testing sets, ensuring balanced classes.
Conducted EDA with word clouds for top 50 words in positive and negative reviews.
Laid the groundwork for advanced sentiment classification with a goal to improve accuracy by 5%.