Skip to content

About

Repository containing slides, exercises, and exams for my course "Natural Language Processing"@epita

Topics

Resources

Stars

4 stars

Watchers

0 watching

Forks

Repository files navigation

Natural Language Processing — Course Material

This repository contains the material for a 30-hour course in Natural Language Processing (NLP), organised as 8 sessions of ~3 hours each, combining theoretical foundations (~1h) and hands-on practical sessions (travaux pratiques, ~2h).

The course follows a progressive path: symbolic methods → probabilistic models → classical ML → neural networks → transformers, with a strong emphasis on building things from scratch and controlled empirical comparison.

This course is part of a broader series of lecture modules:

  1. Introduction to Data Science 🧮
  2. Statistical Learning 📈
  3. Time Series ⌛
  4. Computer Vision Hands-On 🕶️
  5. Recommender Systems 🚀
  6. Deep Learning 💬

Course structure

Each session is organised as:

  • ~1h theory — concepts, intuitions, live demos
  • ~2h practical session (TP) — guided implementation with a concrete deliverable

The practical sessions are cumulative: students progressively build a reusable NLP pipeline rather than starting from scratch at every step. The same benchmark dataset is used from Session 1 through Session 5, enabling direct comparison of every technique.

Session outline

# Title Key concepts Deliverable
01 Tokenisation, Normalisation & Baselines Text normalisation, tokenisation strategies, dataset hygiene, majority baseline Tokeniser comparison table + baseline metrics
02 Language Models & N-grams Chain rule, n-gram LMs, smoothing, perplexity, BPE Perplexity results + BPE walkthrough
03 Classical Text Classification TF-IDF, Naive Bayes, logistic regression, SVM, interpretability Best config + top features + confusion matrix
04 Word Embeddings Distributional hypothesis, Word2Vec, GloVe, polysemy limits TF-IDF vs embeddings comparison
05 Recurrent Neural Networks Sequences, LSTM/GRU, vanishing gradients, batching & masking Training curves + error analysis
06 Transformer Architecture Self-attention, multi-head attention, positional encoding Attention visualisation notebook
07 Pre-training & Fine-tuning BERT, GPT, masked LM, transfer learning, HuggingFace Trainer Fine-tuned classifier + ablation report
08 Token Classification & Decoding NER, BIO tagging, subword alignment, decoding strategies NER F1 + decoding comparison

Repository structure

nlp-course/
│
├── data/                          # datasets (raw, processed, datasheets)
├── img/                           # images used in README and notebooks
│
├── src/
│   ├── tp01_tokenization/
│   │   ├── tp01_tokenization.ipynb        # student notebook
│   │   └── tp01_sol_tokenization.ipynb    # solution  (release after session)
│   │
│   ├── tp02_lm_ngrams/
│   │   ├── tp02_lm_ngrams.ipynb
│   │   ├── tp02_sol_lm_ngrams.ipynb
│   │   └── ngram_lm.py                    # n-gram LM implementation
│   │
│   ├── tp03_classical_models/
│   │   ├── tp03_classical_models.ipynb
│   │   ├── tp03_sol_classical_models.ipynb
│   │   ├── naive_bayes.py                 # Naive Bayes from scratch
│   │   └── logistic_regression.py         # logistic regression helpers
│   │
│   ├── tp04_embeddings/
│   │   ├── tp04_embeddings.ipynb
│   │   └── tp04_sol_embeddings.ipynb
│   │
│   ├── tp05_rnn/
│   │   ├── tp05_rnn.ipynb
│   │   └── tp05_sol_rnn.ipynb
│   │
│   ├── tp06_transformer/
│   │   ├── tp06_transformer.ipynb
│   │   └── tp06_sol_transformer.ipynb
│   │
│   ├── tp07_pretraining/
│   │   ├── tp07_pretraining.ipynb
│   │   └── tp07_sol_pretraining.ipynb
│   │
│   ├── tp08_ner_decoding/
│   │   ├── tp08_ner_decoding.ipynb
│   │   └── tp08_sol_ner_decoding.ipynb
│   │
│   └── utils/                     # shared helpers (metrics, data loading, training loops)
│       └── __init__.py
│
├── .github/                       # CI / issue templates
├── .gitignore
├── Dockerfile
├── Makefile
├── environment.yml
├── requirements.txt
├── requirements-macm1.txt
└── README.md

Notebook conventions

  • Student notebooks (tpXX_*.ipynb) — theory cells are complete; code cells contain # YOUR CODE HERE stubs with type hints and docstrings. Students fill in the logic, not boilerplate.
  • Solution notebooks (tpXX_sol_*.ipynb) — identical structure, all stubs implemented. Released after each session. Include an Instructor notes cell with expected results, common mistakes, and a timing guide.
  • .py companion modules — present only when meaningful reusable code exists (e.g. ngram_lm.py, naive_bayes.py). Not every session requires one.

Companion material

The following notebook is used as a lecture companion in Session 02. It is a conceptual walkthrough (no code) and does not have a student/solution split.


Getting started

Option 1 — pip

git clone https://github.com/oscar-defelice/NLP-Course.git
cd NLP-Course
python -m venv .venv && source .venv/bin/activate   # Windows: .venv\Scripts\activate
pip install -r requirements.txt
jupyter lab src/

Option 2 — conda

conda env create -f environment.yml
conda activate nlp-course
jupyter lab src/

Option 3 — Docker (zero dependency conflicts)

make
# Jupyter will be running at http://localhost:8888

Running in the cloud

No local setup required — open any notebook directly in your browser:

Platform Link
Google Colab Open a notebook → File → Open in Colab
Binder (badge coming — add after first public push)

Your lecturer 👨‍🏫

I am a theoretical physicist working at the intersection of machine learning, natural language processing, and computational biology. I write (occasionally) on Medium and keep personal open-source projects on GitHub.

📫 oscar.defelice@gmail.com

GitHub Website Twitter LinkedIn


Questions & contributions

questions

Found a bug or have a question? Please open an issue or send an email.

Support ☕️

If you find these lectures useful, consider buying me a coffee!


 

About

Repository containing slides, exercises, and exams for my course "Natural Language Processing"@epita

Topics

Resources

Stars

4 stars

Watchers

0 watching

Forks

Releases

Sponsor this project

Contributors

Languages