تفاصيل العمل

Data Cleaning & Standardization Project | SQL Objective: Cleaned and preprocessed a raw, unformatted global layoffs dataset containing thousands of records to make it ready for reliable exploratory data analysis (EDA).

Duplicate Removal: Developed a Common Table Expression (CTE) utilizing windows functions (ROW_NUMBER() with PARTITION BY) to safely identify and purge duplicate records, ensuring data integrity.

Data Standardization: Standardized inconsistent text entries across key categorical columns (e.g., merging variations of 'Crypto' and 'Fin-Tech' into unified categories) using advanced text pattern matching (LIKE and TRIM).

Handling Missing Values: Resolved critical missing data in the industry column by writing a self-join query to populate nulls based on matching company and location profiles.

Database Optimization: Filtered and removed useless/incomplete rows where both key metrics (total_laid_off and percentage_laid_off) were null to optimize dataset quality.

ملفات مرفقة

بطاقة العمل

عدد الإعجابات
0
تاريخ الإضافة
تاريخ الإنجاز
المهارات