Project Workflow
1. Data Extraction (Web Scraping + CSV + APIs)
IMDB Web Scraping → BeautifulSoup to collect Title, Revenue, Rating, Director, Year
Kaggle Datasets → CSV datasets for extended insights
Data Integration → Merged everything into one consolidated dataset with Pandas
2. Data Cleaning & Transformation ️
Handled missing values, removed duplicates, standardized formats
Normalized revenue columns for accurate comparisons
Built an automated ETL pipeline for consistency
3. Storage & Processing
Lightweight ETL architecture for structured & semi-structured data
Python + Pandas for processing raw data
Optimized pipeline for scalability and performance
4. Visualization & Insights (Power BI Dashboard)
Top-Grossing Movies
Highest-Rated Movies
Top Directors
Revenue Trends
5. Tech Stack Used
- Python (Pandas, BeautifulSoup, NumPy)
- ETL Pipeline (Extraction, Cleaning, Transformation)
- Power BI (Dashboards & Insights)
- Jupyter Notebook (Development & Analysis)