Skip to content

Latest commit

 

History

History

Folders and files

NameName
Last commit message
Last commit date

parent directory

..
 
 
 
 
 
 
 
 

README.md

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?

Dong Yan1,2,3, Jian Liang1,3†, Dapeng Hu2†, Ran He1,3, Nicholas Jing Yuan2, Qi Zhang2, Tieniu Tan1,3,4

1School of Artificial Intelligence, University of Chinese Academy of Sciences
2Microsoft
3Institute of Automation, Chinese Academy of Sciences
4Nanjing University

📧 liangjian92@gmail.com   dapenghu@microsoft.com

arXiv

🚀 News

  • [2026/08] Code is released!
  • [2026/07] Code is under preparation. Stay tuned!

📖 Overview

Large language model (LLM) agents can self-evolve by continually improving from their own accumulated experience. However, existing studies predominantly adopt independent evaluation. Consequently, the behavior of self-evolving agents in realistic streaming settings, where agents adapt to diverse and complex task streams, remains poorly understood. To address this gap, we introduce AgentStream, a unified framework that evaluates self-evolving agents spanning diverse evolution components by organizing agentic benchmarks into a configurable task stream and instantiating the Isolated, Sequential, and Interleaved streaming scenarios at test time, which progressively vary the scope and domain composition of the stream. Over these scenarios, we combinatorially evaluate five representative self-evolving methods across three frontier foundation models, disentangling how model capability, method architecture, and streaming scenario jointly shape self-evolution. Our results show that self-evolution reliability varies across streaming scenarios, the benefit of self-evolution is gated by model capability and non-monotonic in model strength, and no single method dominates across models and scenarios. These findings offer concrete guidance for selecting self-evolving methods across models and streaming scenarios. Overall, we advocate that self-evolving agents should be evaluated under realistic task streams rather than isolated single-task settings.

Framework of AgentStream

⚡️ Getting Started

AgentStream is built on a locally adapted snapshot of the Exgentic framework, bundled under exgentic. This copy adds the self-evolving agents and streaming experiment runners used by AgentStream. AgentStream-specific changes to the snapshot are maintained in Sico, and the bundled package is not published independently from this repository. The five self-evolving agents live under exgentic/src/exgentic/agents, and the benchmarks are orchestrated through exgentic's installation and runner infrastructure.

1. Requirements

  • Python >= 3.11
  • uv
  • Docker (optional)

2. Install the local exgentic (agent side)

Clone the repo and create an editable environment from the bundled exgentic:

git clone https://github.com/microsoft/Sico.git
cd Sico/labs/AgentStream/exgentic

# Install the local ./src/exgentic in editable mode into .venv/
uv sync

# Activate the environment
source .venv/bin/activate

Verify that the self-evolving agents are visible from the local install:

uv run exgentic list agents

3. Install benchmarks (benchmark side)

Each benchmark is installed into isolated venv environment:

cd Sico/labs/AgentStream/exgentic


uv run exgentic install --benchmark tau2
uv run exgentic install --benchmark bfcl
uv run exgentic install --benchmark hle
uv run exgentic install --benchmark appworld
uv run exgentic install --benchmark swebench
uv run exgentic install --benchmark browsecompplus

4. API credentials

The runners call LLMs through LiteLLM. Set the credentials for your provider in the exgentic/scripts/<method>/run_experiment.sh:

export OPENAI_API_KEY="..."
export OPENAI_API_BASE="..."

5. Run the streaming experiments

Each method has its own runner under exgentic/scripts/<method>. The shell script selects the streaming scenario via MODE (isolated | sequential | interleaved), the model, the seed, and the benchmark stream:

cd Sico/labs/AgentStream/exgentic/scripts/ace

bash run_experiment.sh

🙏 Acknowledgement

This work is based on Exgentic. We sincerely thank the authors and contributors of these excellent open-source projects.

📚 Citation

If you find our work helpful, please consider citing:

@article{yan2026agentstream,
  title={AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?},
  author={Yan, Dong and Liang, Jian and Hu, Dapeng and He, Ran and Yuan, Nicholas Jing and Zhang, Qi and Tan, Tieniu},
  journal={arXiv preprint arXiv:2608.00155},
  year={2026}
}

📄 License

The contents of this AgentStream directory are licensed separately under the Apache License 2.0.