This repository contains the material for a 30-hour course in Natural Language Processing (NLP), organised as 8 sessions of ~3 hours each, combining theoretical foundations (~1h) and hands-on practical sessions (travaux pratiques, ~2h).
The course follows a progressive path: symbolic methods → probabilistic models → classical ML → neural networks → transformers, with a strong emphasis on building things from scratch and controlled empirical comparison.
This course is part of a broader series of lecture modules:
- Introduction to Data Science 🧮
- Statistical Learning 📈
- Time Series ⌛
- Computer Vision Hands-On 🕶️
- Recommender Systems 🚀
- Deep Learning 💬
Each session is organised as:
- ~1h theory — concepts, intuitions, live demos
- ~2h practical session (TP) — guided implementation with a concrete deliverable
The practical sessions are cumulative: students progressively build a reusable NLP pipeline rather than starting from scratch at every step. The same benchmark dataset is used from Session 1 through Session 5, enabling direct comparison of every technique.
| # | Title | Key concepts | Deliverable |
|---|---|---|---|
| 01 | Tokenisation, Normalisation & Baselines | Text normalisation, tokenisation strategies, dataset hygiene, majority baseline | Tokeniser comparison table + baseline metrics |
| 02 | Language Models & N-grams | Chain rule, n-gram LMs, smoothing, perplexity, BPE | Perplexity results + BPE walkthrough |
| 03 | Classical Text Classification | TF-IDF, Naive Bayes, logistic regression, SVM, interpretability | Best config + top features + confusion matrix |
| 04 | Word Embeddings | Distributional hypothesis, Word2Vec, GloVe, polysemy limits | TF-IDF vs embeddings comparison |
| 05 | Recurrent Neural Networks | Sequences, LSTM/GRU, vanishing gradients, batching & masking | Training curves + error analysis |
| 06 | Transformer Architecture | Self-attention, multi-head attention, positional encoding | Attention visualisation notebook |
| 07 | Pre-training & Fine-tuning | BERT, GPT, masked LM, transfer learning, HuggingFace Trainer | Fine-tuned classifier + ablation report |
| 08 | Token Classification & Decoding | NER, BIO tagging, subword alignment, decoding strategies | NER F1 + decoding comparison |
nlp-course/
│
├── data/ # datasets (raw, processed, datasheets)
├── img/ # images used in README and notebooks
│
├── src/
│ ├── tp01_tokenization/
│ │ ├── tp01_tokenization.ipynb # student notebook
│ │ └── tp01_sol_tokenization.ipynb # solution (release after session)
│ │
│ ├── tp02_lm_ngrams/
│ │ ├── tp02_lm_ngrams.ipynb
│ │ ├── tp02_sol_lm_ngrams.ipynb
│ │ └── ngram_lm.py # n-gram LM implementation
│ │
│ ├── tp03_classical_models/
│ │ ├── tp03_classical_models.ipynb
│ │ ├── tp03_sol_classical_models.ipynb
│ │ ├── naive_bayes.py # Naive Bayes from scratch
│ │ └── logistic_regression.py # logistic regression helpers
│ │
│ ├── tp04_embeddings/
│ │ ├── tp04_embeddings.ipynb
│ │ └── tp04_sol_embeddings.ipynb
│ │
│ ├── tp05_rnn/
│ │ ├── tp05_rnn.ipynb
│ │ └── tp05_sol_rnn.ipynb
│ │
│ ├── tp06_transformer/
│ │ ├── tp06_transformer.ipynb
│ │ └── tp06_sol_transformer.ipynb
│ │
│ ├── tp07_pretraining/
│ │ ├── tp07_pretraining.ipynb
│ │ └── tp07_sol_pretraining.ipynb
│ │
│ ├── tp08_ner_decoding/
│ │ ├── tp08_ner_decoding.ipynb
│ │ └── tp08_sol_ner_decoding.ipynb
│ │
│ └── utils/ # shared helpers (metrics, data loading, training loops)
│ └── __init__.py
│
├── .github/ # CI / issue templates
├── .gitignore
├── Dockerfile
├── Makefile
├── environment.yml
├── requirements.txt
├── requirements-macm1.txt
└── README.md
- Student notebooks (
tpXX_*.ipynb) — theory cells are complete; code cells contain# YOUR CODE HEREstubs with type hints and docstrings. Students fill in the logic, not boilerplate. - Solution notebooks (
tpXX_sol_*.ipynb) — identical structure, all stubs implemented. Released after each session. Include an Instructor notes cell with expected results, common mistakes, and a timing guide. .pycompanion modules — present only when meaningful reusable code exists (e.g.ngram_lm.py,naive_bayes.py). Not every session requires one.
The following notebook is used as a lecture companion in Session 02. It is a conceptual walkthrough (no code) and does not have a student/solution split.
src/tp02_lm_ngrams/LanguageModels.ipynb— Language models: definition, chain rule, n-gram approximation, smoothing, perplexity.
git clone https://github.com/oscar-defelice/NLP-Course.git
cd NLP-Course
python -m venv .venv && source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -r requirements.txt
jupyter lab src/conda env create -f environment.yml
conda activate nlp-course
jupyter lab src/make
# Jupyter will be running at http://localhost:8888No local setup required — open any notebook directly in your browser:
| Platform | Link |
|---|---|
| Google Colab | Open a notebook → File → Open in Colab |
| Binder | (badge coming — add after first public push) |
I am a theoretical physicist working at the intersection of machine learning, natural language processing, and computational biology. I write (occasionally) on Medium and keep personal open-source projects on GitHub.
Found a bug or have a question? Please open an issue or send an email.
If you find these lectures useful, consider buying me a coffee!

