Skip to content

[BUG] evaluation: BM25/Embedding index filenames mismatch when running with --from-conv/--to-conv, causing empty retrieval #127

Description

@Donghua-Cai

I encountered a problem when reproducing the evaluation experiments on the longmemeval dataset.

When using sliced runs (e.g. --from-conv 234 --to-conv 264) can make search fail with:

BM25 index not found: bm25_index_conv_234.pkl

Root cause: index files may be built with local sequential ids (0..N-1) while search looks up global conversation ids (234..263).

Result: search_results are empty, and evaluation scores drop for the wrong reason.

Repro:

uv run python -m evaluation.cli
--dataset longmemeval
--system evermemos
--from-conv 234 --to-conv 264
--run-name test
Expected: sliced runs should generate/load indexes with consistent conversation ids so search can find bm25_index_conv_<conv_id>.pkl.

Activity

  1. added a commit that references this issue on Mar 18, 2026
    3e63e5a
  2. cyfyifanchen commented on Jun 6, 2026

    @cyfyifanchen
    Collaborator

    Closing as part of the EverOS 1.0 issue triage. This issue targets the old benchmark/evaluation code path. The current repo uses the 1.0 server API for LoCoMo reproduction; see docs/locomo_benchmark.md. PR #258 adds migration notes for legacy benchmark reports: #258. If this still reproduces with current main and the 1.0 benchmark commands, please open a fresh issue with the exact command and output.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions