Skip to content

Evaluation的环境配置有问题 #22

Description

@Jarod-Leo

某些文件的配置似乎采用了硬编码,yaml和.env配置没有被全部覆盖到
在evermemos.yaml配置了
llm:
provider: "openai"
model: "Qwen/Qwen2.5-7B-Instruct-AWQ"
api_key: "${LLM_API_KEY}"
base_url: "${LLM_BASE_URL:http://192.168.111.4:8899/v1}"
max_tokens: 8192
但是max_tokens在generation result环节似乎没有广播到某些文件依然显示32764,另外LLM Judge model似乎也没有被设置的model name覆盖。

Activity

  1. Jarod-Leo commented on Dec 12, 2025

    @Jarod-Leo
    Author

    是否支持vllm接入的本地模型评估准确率

  2. freshpomelo commented on Dec 15, 2025

    @freshpomelo

    是否支持vllm接入的本地模型评估准确率

    Yes — we do support using a self-hosted vLLM for evaluation. You may change the "model" and "base_url" in the evermemos.yaml to try with your local model if it follows the openai convention.

  3. freshpomelo commented on Dec 15, 2025

    @freshpomelo

    某些文件的配置似乎采用了硬编码,yaml和.env配置没有被全部覆盖到 在evermemos.yaml配置了 llm: provider: "openai" model: "Qwen/Qwen2.5-7B-Instruct-AWQ" api_key: "${LLM_API_KEY}" base_url: "${LLM_BASE_URL:http://192.168.111.4:8899/v1}" max_tokens: 8192 但是max_tokens在generation result环节似乎没有广播到某些文件依然显示32764,另外LLM Judge model似乎也没有被设置的model name覆盖。

    Regarding the issues you mentioned:

    • Hard-coded max_tokens
      You're right — the answer-generation step currently uses a hard-coded max_tokens, so the value you added in the evermemos.yaml isn’t fully applied. This hard-coded max_tokens will be removed in an upcoming release.
      As a temporary workaround, you can remove or modify the hard-coded parameter in that module.

    • Judge LLM configuration
      The Judge LLM settings are placed in the evaluation/config/datasets YAMLs to ensure consistent evaluation across systems. If you need to change the Judge model, you can edit the dataset YAML directly. After changing it, please check the model’s instruction-following behavior.
      We will also improve the robustness of Judge output parsing for other LLMs in future updates.

  4. cyfyifanchen commented on Jun 6, 2026

    @cyfyifanchen
    Collaborator

    Closing as part of the EverOS 1.0 issue triage. This issue targets the old benchmark/evaluation code path. The current repo uses the 1.0 server API for LoCoMo reproduction; see docs/locomo_benchmark.md. PR #258 adds migration notes for legacy benchmark reports: #258. If this still reproduces with current main and the 1.0 benchmark commands, please open a fresh issue with the exact command and output.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions