diff --git a/labs/AgentStream/README.md b/labs/AgentStream/README.md new file mode 100644 index 00000000..58fb6d24 --- /dev/null +++ b/labs/AgentStream/README.md @@ -0,0 +1,40 @@ +
+ +

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?

+ +

+ Dong Yan1,2,3, + Jian Liang1,3†, + Dapeng Hu2†, + Ran He1,3, + Nicholas Jing Yuan2, + Qi Zhang2, + Tieniu Tan1,3,4 +

+ +

+ 1School of Artificial Intelligence, University of Chinese Academy of Sciences
+ 2Microsoft
+ 3Institute of Automation, Chinese Academy of Sciences
+ 4Nanjing University +

+ +

+ 📧 liangjian92@gmail.com   + dapenghu@microsoft.com +

+ +
+ +## 🚀 News +* **[2026/07]** Code is under preparation. Stay tuned! + +## 📖 Overview +Large language model (LLM) agents can self-evolve by continually improving from their own accumulated experience. +However, existing studies predominantly adopt independent evaluation. +Consequently, the behavior of self-evolving agents in realistic streaming settings, where agents adapt to diverse and complex task streams, remains poorly understood. +To address this gap, we introduce AgentStream, a unified framework that evaluates self-evolving agents spanning diverse evolution components by organizing agentic benchmarks into a configurable task stream and instantiating the `Isolated`, `Sequential`, and `Interleaved` streaming scenarios at test time, which progressively vary the scope and domain composition of the stream. +Over these scenarios, we combinatorially evaluate five representative self-evolving methods across three frontier foundation models, disentangling how model capability, method architecture, and streaming scenario jointly shape self-evolution. +Our results show that self-evolution reliability varies across streaming scenarios, the benefit of self-evolution is gated by model capability and non-monotonic in model strength, and no single method dominates across models and scenarios. +These findings offer concrete guidance for selecting self-evolving methods across models and streaming scenarios. +Overall, we advocate that self-evolving agents should be evaluated under realistic task streams rather than isolated single-task settings.