Customer churn is one of the most critical problems in subscription-based businesses.
This project focuses on predicting customer churn using Machine Learning techniques, with a strong emphasis on interpretability, statistical rigor, and real-world applicability.
The goal is not only to build predictive models but also to understand the drivers of churn, explain model decisions, and extract actionable insights for business decision-making.
This repository contains:
- Exploratory Data Analysis (EDA)
- Feature engineering
- Model training and evaluation
- Interpretation using coefficients and SHAP values
- A final production-ready notebook
A evasão de clientes (churn) é um dos principais desafios de negócios baseados em assinatura.
Este projeto tem como objetivo prever o churn de clientes utilizando técnicas de Machine Learning, com forte foco em interpretabilidade, fundamentação estatística e aplicação prática no mundo real.
Mais do que apenas prever, o projeto busca entender os fatores que levam ao cancelamento, explicar as decisões do modelo e gerar insights acionáveis para o negócio.
Este repositório contém:
- Análise exploratória dos dados
- Engenharia de atributos
- Treinamento e avaliação de modelos
- Interpretação via coeficientes e SHAP
- Um notebook final pronto para portfólio
Customer churn directly impacts revenue, growth, and long-term sustainability.
Being able to anticipate churn allows companies to act proactively, offering retention strategies before customers leave.
O churn impacta diretamente receita, crescimento e sustentabilidade do negócio.
Antecipar o cancelamento permite ações preventivas, como campanhas de retenção e melhorias de serviço.
- data_science/
- data/
- raw/
WA_Fn-UseC_-Telco-Customer-Churn.csv
- raw/
- notebooks/
01_lab_exploration.ipynb02_final_churn_model.ipynb
- data/
requirements.txtREADME.md
Exploratory Data Analysis & Feature Understanding
- Data loading and cleaning
- Handling missing values
- Distribution analysis
- Correlation analysis
- Initial business insights
Final Model & Interpretation
- Feature selection
- Train/test split
- Model training:
- Logistic Regression (primary and baseline model)
- Random Forest (comparative model for non-linear effects)
- Evaluation metrics:
- Accuracy
- Precision
- Recall
- Confusion Matrix
- Model interpretation:
- Coefficients (Logistic Regression)
- SHAP values (global and local explanations)
- Baseline and main model
- High interpretability
- Coefficients analyzed to understand feature impact
- Used as a comparative model
- Captures non-linear relationships
- Higher predictive power in some scenarios
- Feature importance and SHAP explanations
The dataset was split into training and testing sets to evaluate generalization performance.
Care was taken to avoid data leakage, ensuring that all preprocessing and feature engineering steps were applied consistently across splits.
O conjunto de dados foi dividido em treino e teste para avaliar a capacidade de generalização do modelo.
Foram adotados cuidados para evitar vazamento de dados, garantindo que as etapas de pré-processamento e engenharia de atributos fossem aplicadas de forma consistente.
Used to understand how each feature affects the probability of churn in linear models.
- Global feature importance
- Local explanations for individual predictions
- Clear visualization of model behavior
This combination ensures trustworthy and explainable AI.
- Contract type strongly influences churn
- Monthly charges have a significant impact
- Tenure is one of the most protective factors against churn
- Some features show non-linear effects, captured better by tree-based models
-
📺 YouTube Video (PT-BR)
https://www.youtube.com/watch?v=cbw1K6l7Bxg -
📝 LinkedIn Article
https://www.linkedin.com/posts/jessepmelo_apresento-aqui-um-projeto-completo-de-machine-activity-7424052162999058433-Phip
The video presentation is in Portuguese.
YouTube automatically provides AI-based audio translation for English viewers.
git clone https://github.com/JessePMelo/Telco-Customer-Churn.git
cd Telco-Customer-Churn
python -m venv venv
source venv/bin/activate # Linux/Mac
venv\Scripts\activate # Windows
pip install -r requirements.txt
- Open Google Colab
- Upload the notebooks
- Install dependencies: !pip install pandas numpy scikit-learn matplotlib shap statsmodels
pandas
numpy
scikit-learn
matplotlib
shap
statsmodels
- Real-world business problem
- Strong focus on explainability
- Combines statistics and machine learning
- Portfolio-ready and recruiter-friendly
- Clear communication of results
This project demonstrates the complete lifecycle of a data science solution, from problem formulation to interpretable results and business insights.
Jessé Pereira de Melo
Data Science | Machine Learning | Business Intelligence
GitHub: https://github.com/JessePMelo
LinkedIn: https://www.linkedin.com/in/jessepmelo/
This project is intended for educational and portfolio purposes.