Official ICLR 2026 implementation: evaluate class-incremental learners through performance distributions and extreme class orders.
Paper · PDF · Get Started · Citation · Issues
Guannan Lai · Da-Wei Zhou · Xin Yang · Han-Jia Ye
This repository contains the official implementation of EDGE (Extreme case–based Distribution & Generalization Evaluation), an evaluation protocol for Class-Incremental Learning (CIL).
Mainstream CIL evaluation typically reports the mean (and sometimes variance) over only a small number of randomly sampled class sequences. However, CIL performance can vary substantially across sequences, and limited sampling may lead to biased mean estimates and a severe underestimation of the true variance in the performance distribution.
Our paper argues that robust CIL evaluation should characterize the full performance distribution, and introduces extreme sequences as a principled tool to approximate distributional boundaries efficiently.
EDGE evaluates a CIL method by estimating not only the central tendency but also the distributional boundaries (e.g., near-best / near-worst sequences).
At a high level, EDGE:
- Computes or estimates inter-task similarity between incremental tasks/classes.
- Uses the similarity signal to search for extreme sequences.
- Samples sequences adaptively to better approximate the true performance distribution.
- [2026-01] 🌟 Accepted by ICLR 2026
- [2025-10] 🌟 Initial release of EDGE evaluation code
- [2025-09] 🌟 arXiv preprint released
We integrate EDGE into two widely-used CIL toolboxes:
This repo contains two subfolders, PILOT/ and PyCIL/. Our main modifications include:
- Updating
main.pyandtrainer.py - Adding
utils/edge.py(EDGE core logic)
EDGE is largely decoupled from the core training pipelines of PILOT and PyCIL. Therefore, if PILOT/PyCIL updates in the future, this integration can be adapted with minimal changes.
Concretely, we introduce an --eval argument:
--eval random: the baseline using randomly sampled class orders--eval edge: the released implementation evaluates similarity-guided hard/easy class orders and a random order
Both entry points default to --eval edge. The trainers report maximum, minimum, mean, and standard deviation of accuracy over the evaluated sequences.
git clone https://github.com/AIGNLAI/EDGE
cd EDGE- Python >= 3.8
- PyTorch >= 2.0
- torchvision
- timm
- numpy / scipy
- tqdm
We recommend using a clean conda environment.
conda create -n edge python=3.10 -y
conda activate edge
# Example environment
pip install torch==2.0.0+cu118 torchvision==0.15.1+cu118 --extra-index-url https://download.pytorch.org/whl/cu118
pip install git+https://github.com/openai/CLIP.git
pip install scipy
pip install timm==0.6.12
pip install tqdmRun commands from the selected toolbox directory. For example, from the repository root, use PILOT's ImageNet-R configuration:
cd PILOT
python main.py --config ./exps/simplecil_inr.json --eval edgeTo use the random-sequence baseline:
python main.py --config ./exps/simplecil_inr.json --eval randomFor PyCIL, run cd PyCIL from the repository root and choose a configuration in its exps/ directory, using the same --config and --eval flags. Dataset preparation, model dependencies, and dataset paths are described in the PILOT guide and PyCIL guide.
Dataset setup: the released extreme-sequence search uses a fixed 200-class assumption. PILOT includes class-name lists for CIFAR-100, CUB-200, ImageNet-R, and ImageNet-A; PyCIL currently includes CIFAR-100 only. For a new dataset or a different class count, configure the class-name registry and class-count/search settings in the selected toolbox's utils/edge.py before using EDGE. The supplied 100-class configurations therefore need the search settings adjusted. Choose init_cls and increment in the experiment JSON to match the intended task partition.
Training logs are written under logs/<model>/<dataset>/<init_cls>/<increment>/ within the selected toolbox. The summary describes the evaluated orders; consult the paper for the full evaluation protocol and experimental settings.
If you find this repo useful, please consider citing:
@inproceedings{lai2026lie,
title = {The Lie of the Average: How Class Incremental Learning Evaluation Deceives You?},
author = {Lai, Guannan and Zhou, Da-Wei and Yang, Xin and Ye, Han-Jia},
booktitle = {International Conference on Learning Representations (ICLR)},
year = {2026},
url = {https://openreview.net/forum?id=19LHXi9uLw}
}We thank the following repos/projects for helpful components:
- PILOT
- PyCIL
For questions and feedback, please open an issue or contact:
- Guannan Lai (laign@lamda.nju.edu.cn)

