From 0ace8449b16e0d55189d8aec94ca64de1b033ccd Mon Sep 17 00:00:00 2001
From: ttang911 <8541896+ttang911@users.noreply.github.com>
Date: Tue, 4 Aug 2026 14:24:00 +0800
Subject: [PATCH] Add the initial file and folder for Sico labs
---
labs/AgentStream/README.md | 40 ++++++++++++++++++++++++++++++++++++++
1 file changed, 40 insertions(+)
create mode 100644 labs/AgentStream/README.md
diff --git a/labs/AgentStream/README.md b/labs/AgentStream/README.md
new file mode 100644
index 00000000..58fb6d24
--- /dev/null
+++ b/labs/AgentStream/README.md
@@ -0,0 +1,40 @@
+
+
+
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?
+
+
+ Dong Yan1,2,3,
+ Jian Liang1,3†,
+ Dapeng Hu2†,
+ Ran He1,3,
+ Nicholas Jing Yuan2,
+ Qi Zhang2,
+ Tieniu Tan1,3,4
+
+
+
+ 1School of Artificial Intelligence, University of Chinese Academy of Sciences
+ 2Microsoft
+ 3Institute of Automation, Chinese Academy of Sciences
+ 4Nanjing University
+
+
+
+ 📧 liangjian92@gmail.com
+ dapenghu@microsoft.com
+
+
+
+
+## 🚀 News
+* **[2026/07]** Code is under preparation. Stay tuned!
+
+## 📖 Overview
+Large language model (LLM) agents can self-evolve by continually improving from their own accumulated experience.
+However, existing studies predominantly adopt independent evaluation.
+Consequently, the behavior of self-evolving agents in realistic streaming settings, where agents adapt to diverse and complex task streams, remains poorly understood.
+To address this gap, we introduce AgentStream, a unified framework that evaluates self-evolving agents spanning diverse evolution components by organizing agentic benchmarks into a configurable task stream and instantiating the `Isolated`, `Sequential`, and `Interleaved` streaming scenarios at test time, which progressively vary the scope and domain composition of the stream.
+Over these scenarios, we combinatorially evaluate five representative self-evolving methods across three frontier foundation models, disentangling how model capability, method architecture, and streaming scenario jointly shape self-evolution.
+Our results show that self-evolution reliability varies across streaming scenarios, the benefit of self-evolution is gated by model capability and non-monotonic in model strength, and no single method dominates across models and scenarios.
+These findings offer concrete guidance for selecting self-evolving methods across models and streaming scenarios.
+Overall, we advocate that self-evolving agents should be evaluated under realistic task streams rather than isolated single-task settings.