Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

Β 

History

3 Commits
Β 
Β 
Β 
Β 

Repository files navigation

πŸš€ Customer Analytics Case Study

Segmentation + Churn Prediction + Next-Month Spend Forecast (End-to-End ML System)

An industry-grade data science project that simulates a real-world subscription-based e-commerce business and builds three production-relevant machine learning systems:

  • 🎯 Customer Segmentation (Unsupervised Learning)
  • ⚠️ Churn Prediction (Classification)
  • πŸ’° Next-Month Spend Forecasting (Regression)

🏒 Business Context

Modern e-commerce platforms generate large volumes of customer behavioral and transactional data. However, without structured analytics, businesses fail to:

  • Identify high-value customers
  • Prevent customer churn
  • Forecast revenue accurately

This project builds a data-driven decision system that enables businesses to:

  • Segment customers intelligently
  • Predict churn risk proactively
  • Forecast revenue at a granular level

🎯 Business Objectives

1️⃣ Customer Segmentation

  • Identify distinct customer groups
  • Enable targeted marketing strategies
  • Improve campaign efficiency

2️⃣ Churn Prediction

  • Predict probability of customer churn
  • Enable early retention intervention

3️⃣ Spend Forecasting

  • Predict next-month customer spending
  • Support revenue planning and personalization

🧠 Data Strategy (Simulation with Realistic Assumptions)

Dataset Scale

  • πŸ‘₯ Customers: 5,000
  • πŸ›’ Transactions: ~40,000
  • πŸ“… Simulated over multiple months

πŸ“Š Features Generated

πŸ‘€ Customer Attributes

  • Age (18–70)
  • Income ($5K – $150K)
  • Country (US 60%, UK 40%)
  • Tenure (1–60 months)

πŸ“± Behavioral Features

  • Visits per month (Poisson distributed)
  • Engagement intensity

πŸ›’ Transaction Features

  • Total orders
  • Total spend
  • Average order value (AOV)
  • Discount usage patterns

πŸ”Ž Exploratory Data Analysis (EDA)

Key Statistics

  • Average monthly visits: 5.8 visits/user
  • Average order value: $72.4
  • Average spend per customer: $1,180
  • High-income segment contributes ~48% of total revenue

Observations

  • Income distribution is right-skewed
  • Engagement strongly correlates with spending
  • Low-activity users show early signs of churn

🧾 Feature Engineering

Aggregated Features

  • Total Spend
  • Total Orders
  • Avg Order Value
  • Visit Frequency

Derived Features

  • Spend per visit
  • Engagement ratio
  • Tenure-adjusted activity

πŸ”· Part A β€” Customer Segmentation (KMeans + PCA)

βš™οΈ Approach

  • Feature Scaling using StandardScaler
  • PCA β†’ Reduced to 2 components (~81% variance explained)
  • KMeans clustering

πŸ” Optimal Cluster Selection

  • Elbow Method β†’ K = 5

πŸ“Š Cluster Summary

Cluster Size Avg Spend Avg Visits Segment Type
0 1,120 $1,950 8.2 High-value engaged
1 980 $420 2.1 Low-value inactive
2 1,050 $1,200 5.4 Mid-tier steady
3 890 $300 3.0 Price-sensitive
4 960 $2,400 9.5 Premium users

πŸ’‘ Insights

  • Top 20% customers generate ~55% revenue
  • Low-engagement users (<3 visits/month) have 2.5x higher churn risk
  • Premium segment shows highest retention

πŸ”· Part B β€” Churn Prediction (Classification)

βš™οΈ Models Evaluated

  • Logistic Regression
  • Random Forest βœ… (Best)

πŸ“Š Model Performance

Metric Logistic Regression Random Forest
Accuracy 0.78 0.86
Precision 0.74 0.83
Recall 0.69 0.81
F1 Score 0.71 0.82

πŸ” Key Drivers

  • Low visits/month
  • Low spend
  • Short tenure
  • High discount usage

πŸ’‘ Insights

  • Users with <3 visits/month β†’ 68% churn probability
  • Users with tenure >24 months β†’ <15% churn probability

πŸ”· Part C β€” Spend Forecasting (Regression)

βš™οΈ Model Used

  • Linear Regression

πŸ“Š Model Performance

  • RMSE: $142.6
  • RΒ² Score: 0.72

πŸ” Key Drivers of Spend

  1. Historical spend
  2. Visit frequency
  3. Avg order value
  4. Tenure

πŸ’‘ Insights

  • Users with β‰₯8 visits/month spend 2.3x more
  • High AOV users dominate revenue growth

πŸ—οΈ End-to-End Pipeline

Data Simulation
   ↓
EDA & Validation
   ↓
Feature Engineering
   ↓
Segmentation (K-Means + PCA)
   ↓
Churn Prediction (Random Forest)
   ↓
Spend Forecasting (Linear Regression)
   ↓
Business Insights


πŸ“Š Key Business Insights

  • Top 20% customers generate >50% of revenue
  • Engagement is the strongest predictor of churn
  • Segmentation enables targeted marketing
  • Spend prediction improves revenue planning

🧠 Key Learnings

  • Built multi-model ML pipeline (unsupervised + classification + regression)
  • Applied feature engineering for behavioral analytics
  • Connected ML outputs to business decisions
  • Learned trade-offs between model types

🧰 Tech Stack

  • Python
  • Pandas / NumPy
  • Matplotlib / Seaborn
  • Scikit-learn
  • KMeans
  • PCA
  • Random Forest
  • Linear Regression

πŸ“ Project Structure

Customer-Analytics-Case-Study/
β”‚
β”œβ”€β”€ CASE_STUDY_DS.ipynb
β”œβ”€β”€ README.md


πŸ”₯ Future Improvements

  • XGBoost / LightGBM models
  • Time-series forecasting
  • FastAPI deployment
  • Model monitoring
  • Cloud deployment (AWS/GCP)

πŸ‘¨β€πŸ’» Author

Birjung Thapa
Master’s in Data Science
University of Colorado Boulder


⭐ Support

If you found this project useful:

⭐ Star the repo
πŸ“’ Share with others
🀝 Connect for collaboration

About

End-to-end customer analytics ML system leveraging KMeans segmentation, Random Forest churn prediction, and regression-based spend forecasting to drive data-driven marketing, retention, and revenue strategies.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages