Skip to content

Agent-pipeline e2e is the only coverage of the skill chain and CI never runs it #397

Description

@Kendrick-Song

PR #393 fixed tests/e2e/test_add_flush_agent_pipeline_e2e.py so it actually exercises the agent-skill chain — before that fix it structurally could not, and that is why a defect making the chain fail 4/4 reached a release.

The original failure had two layers:

  1. tests/conftest.py:68-88 is an autouse fixture forcing EmbeddingCapability(provider=None) for every test (hermeticity — correct in itself). trigger_skill_clustering and extract_agent_skill both body-guard on get_embedding_capability().available, so both returned early. Measured log counts from an actual run: agent_case_extracted=5, skill_cluster_updated=0, agent_skills_extracted=0, strategy_gated_off_embedding_unavailable=10.
  2. Its three skill assertions were assert len(...) >= 0 — always true. The inline comment rationalized this as LLM-dependent flakiness; the count was in fact necessarily zero.

#393 opts the test in to a real embedding capability and gives it a meaningful floor plus a dead-letter assertion. But it keeps the slow + live_llm markers, and CI injects no provider credentials — so this remains the only end-to-end coverage of the chain, and it still never runs automatically.

Worth deciding how to close that gap. Options: a scheduled (nightly / weekly) job with credentials in secrets, running just the live_llm set; a recorded-fixture replay so the chain can be exercised without credentials (the examples/langfuse trace-replay work may be reusable); or accept it and add a release-checklist step that runs it manually.

Related: the same class of gap likely applies to other live_llm tests — worth auditing which of them, if any, are the sole coverage of their path.

Activity

  1. BrierAinz commented on Sep 19, 2026

    @BrierAinz

    I would close this gap with a two-tier gate rather than choosing only live credentials or only fixtures.

    Tier 1, required on every PR: a recorded or deterministic replay that proves the chain does not structurally short-circuit. It should assert the events that matter, for example agent_case_extracted > 0, skill_cluster_updated > 0, agent_skills_extracted > 0, and no unexpected dead-letter/gated-off reason. This catches the exact class of "assert >= 0" regression without needing provider secrets.

    Tier 2, scheduled or release-gated: the current live LLM path with credentials, allowed to be slower and quarantined from ordinary PR noise.

    The key acceptance criterion I would add is: every live_llm test that is sole coverage for a production path must have either a deterministic replay twin or an explicit release checklist owner. Otherwise the suite can look green while the only real coverage is sitting outside automation.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions