Skip to content

Epic: Fleet traffic & consumption analytics — recurring markdown report #60

Description

@jsboige

Goal

Instrument the fleet's traffic and subscription consumption so we can master our plans (not discover their walls reactively). One recurring markdown report the operator reads, with metrics at several scales and points of attention.

Context (user, 2026-08-26): "il est temps qu'on s'instrumente pour mieux maîtriser notre forfait". No new spend — Qwen and Mistral onboarding were disappointments; the question is visibility into what we already pay for.

Deliverable: recurring report

Markdown report produced off-peak, covering:

  1. Fleet scale — day × model: requests, input/output tokens (resp captures), activation counts
  2. Post scale — machine × workspace × model (who consumed what; scripts/traffic-mxws.py already pairs req/resp by (day, reqN))
  3. Plan scale — per subscription lane: tokens burned since last reset vs known wall dates (Qwen Token Plan, GLM coding ~5h window, MiniMax weekly, Mistral one-shot credit, Codex/OpenAI)
  4. Points d'attention — near-saturated contexts, lanes trending toward a wall, anomalies vs previous period

Hard constraints

  • Runs ONLY in the 05-07Z off-peak window (incident 2026-08-26 14:05-14:46Z: in-day analysis jobs degraded the prod hub 40 min — event-loop starvation + disk-queue saturation)
  • 7z archive extractions go to D: (captures disk), never C: (Docker/WSL2 vhdx lives there)
  • Never blocks the 3h surveillance cron's light checks

Sub-items

  • Evolution report 24→26 Aug (machine × workspace × model) — extractions already staged
  • Recurring scheduled generation of the report (05-07Z window)
  • Qwen accounting discriminator: console token count vs capture-measured tokens since 22:28Z reset → determines if the Token Plan bills input or output
  • Mistral vibe credit consumption via scheduled task (credit expires 2026-09-01)
  • Codex leftover-credit check behind the OpenAI $200/mo ChatGPT subscription; consume via schtask or IDE lane if present

Activity

  1. jsboige commented on Sep 11, 2026

    @jsboige
    OwnerAuthor

    Report slices to add — from the 2026-09-11 user mandate

    The mandate sharpened what the report must answer. Beyond the four scales already listed, the operator needs these three slices (they are what turn the numbers into decisions — plug leaks / adjust crons / make harnesses condense):

    1. Top consumers by Machine:Workspace:Harness — the "who" axis. Same three-level key as Epic: fleet consumption observability — live per-lane table, provider dashboards, cron-to-cost mapping #41; harness fingerprint = cc version (from ua=) + hash of the position-0 harness block + resolved model. Ranked by read tokens (the dominant term), not request count.
    2. Fixed harness share vs. conversation vs. memories — the "what" axis. Per session: tokens of the position-0 harness block, of memories (topic files), and of the actual conversation; plus the ratio to fresh-session cost. harness-injection-measure.py already computes the chars-per-token conversion; [EPIC] Harness slimming — mesurer 'chargé vs mobilisé' via les traces, et aider les workspaces à rester slim #23 carries the measurement method.
    3. Compaction slice — compactions/day per lane and tokens re-paid by compaction (the harness block is re-read at every compaction, so this is the multiplier that makes a heavy harness expensive rather than merely large). Detector is a separate issue; the report should reserve the columns now.

    Plus one operational line the operator reads first: fleet cascade depth — which role is currently served by a fallback step, and which provider is absorbing it (a wall on the nominal is invisible in token totals alone, and it is exactly when the expensive PAYG takes over).

    Constraint already established, restated

    Off-peak 05-07Z only (the 2026-08-26 incident: in-day analysis degraded the hub 40 min). Since the 2026-09-05 migration the hub is po-2025, and a NOMINAL relay writes no local capture — the report must be generated hub-side, from the hub's own captures, never from a relay's.

  2. jsboige commented on Sep 15, 2026

    @jsboige
    OwnerAuthor

    Rattachée comme sub-issue de #116 (EPIC parapluie datastudy consommation)

    Cette issue est absorbée comme couche P2 (restitution) de #116 — non fermée, non remplacée.

    Ce qui est repris tel quel :

    • Les quatre échelles du rapport (flotte jour×modèle · machine×workspace×modèle · plan/forfait · points d'attention).
    • Les trois tranches ajoutées le 11/09 (top consommateurs par Machine:Workspace:Harnais · part harnais fixe / mémoires / conversation · tranche compaction) et la ligne opérationnelle profondeur de cascade.
    • La contrainte dure 05-07Z uniquement (incident 2026-08-26 : un job d'analyse en journée a dégradé le hub 40 min) et l'extraction 7z vers D:, jamais C:.
    • Le fait établi depuis la migration du 05/09 : le rapport se génère côté hub, jamais depuis un relais (un sidecar NOMINAL n'écrit aucune capture).

    Ce que le parapluie ajoute : le rapport devient un dashboard avec graphiques et interprétations écrites (P2), alimenté par une série multi-lanes (P1) et une table de prix (P0), et complété par les métriques projet (P3) — la forme « rapport markdown » reste possible comme sortie, mais n'est plus le livrable unique.

  3. jsboige commented on Oct 6, 2026

    @jsboige
    OwnerAuthor

    [po-203] État de l'epic — file 06/10 12:00Z (R1)

    #60 est absorbée comme couche P2 (restitution) de #116 depuis le 15/09 (c.2) — pas fermée, pas remplacée. État vérifié sous-item par sousitem, ce jour :

    Sous-item du body État (vérifié 06/10)
    Rapport d'évolution 24→26/08 (machine × workspace × modèle) Supersédé — les extractions ont nourri le seed P2 ; traffic-mxws.py existe et paire req/resp par (day, reqN)
    Génération planifiée récurrente (fenêtre 05-07Z) Non fait — scripts/fleet-dashboard.py (seed P2, mergé via #120) régénère en une commande, mais aucun appelant planifié n'existe dans le dépôt (grep : zéro référence hors le script lui-même)
    Discriminateur comptable Qwen (input vs output) Tranché par l'analyse du 14/08 (body fondateur) : 200 M brûlés, 76 % cache hit, 1 M output → le Token Plan facture l'entrée, cache compté brut — la question du sous-item est répondue
    Crédit Mistral « vibe » (expirait 01/09) Obsolète tel qu'écrit — l'arrangement a changé (lane active sous murs 402 jusqu'au 11/10, pas un crédit one-shot) ; le sous-item ne décrit plus la réalité
    Check crédit résiduel Codex (OpenAI $200/mois) Inconnu — jamais mesuré, aucun artefact

    Recommandation : les deux sous-items encore actionnables (la planification 05-07Z de fleet-dashboard.py, le check Codex) sont des grains bornés qui appartiennent désormais au plan de couches de #116 (P2 restitution / P5 modèle de coût). Proposer : (a) re-pointer ces deux grains sur #116 et fermer #60 comme absorbée-résolue, ou (b) garder #60 ouverte comme conteneur P2 et cocher ce qui est tranché. Arbitrage coordinateur/user — aucune des deux options ne perd d'information (la table ci-dessus fait foi).

    🤖 Generated with Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions