Skip to content

fix(pipeline): tell the customer when the AI backend fails (CRM-236) - #9

Merged
gomessguii merged 6 commits into
developfrom
fix/CRM-236-degraded-provider-feedback
Aug 22, 2026
Merged

fix(pipeline): tell the customer when the AI backend fails (CRM-236)#9
gomessguii merged 6 commits into
developfrom
fix/CRM-236-degraded-provider-feedback

Conversation

@pastoriniMatheus

@pastoriniMatheus pastoriniMatheus commented Aug 22, 2026

Copy link
Copy Markdown

Problema

Quando o provedor de LLM degrada, o turno estoura o teto e o pipeline apenas logava e limpava o estado — nada chegava ao chat. Para o cliente é indistinguível de um bot que o está ignorando.

E é pior do que parece: o efeito colateral da ferramenta já foi aplicado. Na execução ao vivo o card do funil moveu aos ~20s e o timeout disparou aos 30s — o funil andou e a conversa ficou muda.

Medição (a premissa original do card estava errada)

O card dizia "30s não cobre turno com tool-calling". Falso. Turno completo, em condições normais:

5s · 16s · 1s · ~12s

O que oscila é o provedor. Mesma chamada trivial ("responda apenas: ok") ao gemini-2.5-flash, cinco vezes seguidas:

0.74s · 0.79s · 1.36s · 10.34s · 20.40s
min 0.74s · max 20.40s · média 6.73s

Variação de 27× na mesma chamada. Um turno com ferramenta faz ao menos duas dessas idas (decidir a tool, depois redigir a resposta), então duas caudas ruins sozinhas passam de 30s sem nada de errado no código.

E sob cota estourada o litellm ainda retenta internamente:

litellm.RateLimitError: 429 RESOURCE_EXHAUSTED
"Quota exceeded ... limit: 20, model: gemini-2.5-flash ... Please retry in 10.35s"
litellm.ServiceUnavailableError: 503 "This model is currently experiencing high demand."

Mudanças

1. AI_CALL_TIMEOUT_SECONDS 30 → 90. Cobre duas chamadas na cauda mais o round trip do CRM, e ainda limita um provedor genuinamente travado.

Isto não é a correção sozinha: elevar um teto só o desloca. É a mensagem abaixo que protege o cliente quando o teto é atingido.

2. Aviso no chat em vez de silêncio. No timeout ou erro, o pipeline despacha uma frase para a conversa. O erro cru do provedor nunca chega ao cliente — ele carrega nome de modelo, id de cota e URLs — e vai para o log do operador como cause:

WARN pipeline.ai.failure_notice.sending contact_id=… conversation_id=… cause="…"

AI_FAILURE_NOTICE sobrescreve o texto; defini-la vazia mantém o silêncio de hoje, para quem preferir.

Testes

pkg/pipeline/service/ai_failure_notice_test.go6 testes: o cliente é avisado; o erro do provedor nunca vaza (verifica que gemini, Quota, RateLimitError, googleapis não aparecem na mensagem); o texto é sobrescrevível; vazio desativa; postback ausente não quebra; o default vale quando a env não existe.

go build ./...                              OK
go vet ./pkg/... ./internal/... ./cmd/...   OK
go test ./pkg/... ./internal/...            todos ok (Redis real)

Deliberadamente fora deste PR

Dois defeitos reais que pertencem ao lado do processor, não ao bot-runtime:

  1. O processor continua trabalhando depois que o bot-runtime desistiu — gastando cota que já estava esgotada.
  2. O 429/503 do provedor vira 500 INTERNAL_ERROR genérico na rota A2A, então nem o operador vê "cota excedida" sem ler o log do container.

Ambos estão registrados no card. Fiz aqui o que resolve o sintoma para o cliente; aquilo exige mexer no caminho de erro do processor e merece PR próprio.

Trade-off assumido

Um teto de 90s significa que um provedor travado segura o turno por mais tempo antes de o cliente receber o aviso. É o preço de não cortar turnos legítimos com cauda alta — e o aviso garante que, mesmo no pior caso, a conversa não termina em silêncio.

Summary by Sourcery

Ensure customers receive a safe, actionable response when AI processing times out or fails instead of being left in silence.

New Features:

  • Notify customers in chat when an AI timeout or backend error prevents a response, using a safe default message or an operator-configurable notice.

Bug Fixes:

  • Prevent AI backend failures from leaving conversations silently unresponsive while downstream tool side effects may already have occurred.
  • Keep provider-specific error details out of customer-facing messages while retaining them in operator logs.

Enhancements:

  • Increase the default AI call timeout to 90 seconds to accommodate normal tool-calling turns with provider latency spikes.
  • Allow operators to disable the failure notice or customize its text through environment configuration.

Tests:

  • Add coverage for customer notification, provider-error redaction, custom and disabled notices, missing postbacks, and default configuration behavior.

⚠️ Este PR sozinho NÃO tem efeito — precisa do par

Descoberto ao verificar o CRM-236 na stack rodando: depois de rebuildar com este fix, o container ainda reportava AI_CALL_TIMEOUT_SECONDS=30.

Todo compose entregue setava a env explicitamente em 30, e env explícita ganha do default do código. Elevar o default para 90 aqui não muda nada onde importa.

Corrigido em PR separado no repo raiz — evo-crm-community#175 (docker-compose.swarm.yaml e internal/review/docker-compose.yml). Mergear os dois juntos.

.env.example:224 carrega o mesmo 30 e é arquivo protegido no ambiente onde trabalhei, então precisa de um mantenedor — sem isso, toda instalação nova nasce com o teto antigo.

Ordem de merge — PRs irmãos do CRM-236

#9 (este) e #175 devem subir juntos; o #52 é independente e pode ir em qualquer ordem.

Repo PR Papel
evo-bot-runtime este teto 90s + aviso ao cliente
evo-ai-processor-community #52 cancelar execução órfã + classificar erro do provedor (429/503/502 em vez de 500)
evo-crm-community #175 composes deixando de anular o default

Verificação ao vivo do aviso (item C do card)

Feita após este PR: bot-runtime rebuildado, teto forçado a 1s, evento real injetado numa conversa.

pipeline.ai.timeout contact_id=3 conversation_id=3
pipeline.ai.failure_notice.sending cause="bot_runtime: ai call timeout"
pipeline.dispatch.completed parts_total=1

A conversa foi de 7 para 8 mensagens e a nova é o aviso, message_type=1 (outgoing), persistida no CRM. O cliente vê texto em vez de silêncio.

Summary by Sourcery

Ensure customers receive a clear response when AI processing fails or times out instead of being left in silence.

New Features:

  • Notify customers in chat when AI processing times out or fails, using a safe default message that operators can customize or disable.

Bug Fixes:

  • Prevent AI backend failures from leaving conversations silent while preserving the state and bookkeeping of subsequent turns.
  • Keep provider-specific failure details out of customer-facing messages while retaining them in operator logs.

Enhancements:

  • Increase the AI call timeout to 90 seconds and provide a bounded dispatch window for failure notices.
  • Strengthen pipeline isolation coverage to verify that successful responses and failure notices are delivered to the correct conversations.

Deployment:

  • Update the Kubernetes ConfigMap to use the 90-second AI call timeout.

Tests:

  • Add coverage for failure notification behavior, error redaction, configuration overrides, disabled notices, missing postbacks, turn-state preservation, segmented dispatch timing, and end-to-end pipeline isolation.

Review — críticos 2, 3 e item 9

🔴 2 — o aviso podia apagar o turno seguinte

Confirmado, e o cenário é justamente o que a feature mira.

sendAIFailureNotice despachava via runDispatchStage, que é dono do bookkeeping do turno: o caminho de sucesso roda SetState(StageDone)ClearStateentries.Delete(pairKey). Esse Delete é o que os comentários da própria função proíbem: "A Delete here would race with the new event's Store and could delete the replacement entry."

O cliente espera, desiste, manda "oi?" durante o despacho do aviso → startDebounce guarda o turno novo → o aviso termina e deleta essa entry, órfã o turno novo, e sobem dois pipelines no mesmo par: resposta duplicada.

Agora o aviso despacha direto. Não havia o que book-keepar de todo jeito: os dois call sites já rodam clearStateWithLog antes.

Prova negativa: restaurando a chamada, falha DoesNotTouchTheEntryOfTheNextTurn com a mensagem do turno órfão.

🟡 6 — fechou junto

cleanupCtx são 5s documentados para cleanup (ClearState, SetState). Um Dispatch segmenta por TextSegmentationLimit e dorme DelayPerCharacter por runa entre as partes: com segmentação ligada, o aviso de 84 runas virava 2-3 partes e só os delays estouravam os 5s → ErrDispatchInterrupted → o log dizia "New message arrived" sem que mensagem nenhuma tivesse chegado. Agora noticeCtx() (30s).

🔴 3 — o teto seguia anulado dentro deste repo

.env.example:8 e k8s/configmap.yaml:8 corrigidos para 90. O critério é do próprio PR: se env explícita ganha do default, deixar 30 no ConfigMap que serve staging/prod anula o fix onde ele mais importa.

Não tocado: evolution-ecosystem/k8s/base/bot-runtime.yaml:44-45 também fixa 30, mas é o SaaS, fora destes três PRs.

🟡 9 — aviso default em pt-BR hardcoded

O bot-runtime roda em instalações community/self-hosted no mundo todo, e clientes de instalações que nunca escolheram português eram respondidos nele. Agora inglês, com AI_FAILURE_NOTICE documentada em .env.example para localização — e string vazia ainda restaura o silêncio. Há teste que falha se o pt-BR voltar.

🟢 11, 12, 13

Doc comment devolvido ao clearStateWithLog (meu bloco novo tinha caído entre ele e a função, então o godoc de aiFailureNoticeEnv começava com "clearStateWithLog calls ClearState…"), contains trocado por strings.Contains, e os.Unsetenv agora restaura via t.Cleanup.

Testes

10 (eram 6): entry do turno seguinte intacta, estado não reescrito, orçamento do dispatch, default sem pt-BR. go build/go vet limpos, suíte completa verde com Redis real.

gofmt: pipeline_service.go já está fora de formato em develop (CRLF); deixado como está para não jogar o arquivo inteiro neste diff.

Em aberto, dito por mim

10 (nada exercita sendAIFailureNotice via runAIStage) e 14 (AI_FAILURE_NOTICE lido por os.LookupEnv em vez de passar pelo internal/config).

When the LLM provider degrades the turn exceeds the ceiling and the pipeline
only logged and cleared state — nothing reached the chat. For the customer that
is indistinguishable from a bot ignoring them, and it is worse than it looks:
the tool's side effect may ALREADY be applied. In the live run the pipeline card
moved at ~20s and the timeout fired at 30s, so the funnel advanced while the
conversation stayed silent.

Two changes.

1. AI_CALL_TIMEOUT_SECONDS default 30 -> 90. Measured against gemini-2.5-flash,
   the SAME trivial prompt answered in 0.74s / 0.79s / 1.36s / 10.34s / 20.40s —
   a 27x spread on the provider's tail. A tool-calling turn makes at least two of
   those round trips (decide the tool, then write the reply), so two bad tails
   alone exceed 30s with nothing wrong in the code. 90s covers that and still
   bounds a genuinely hung provider.

   Note this is NOT the fix on its own: raising a ceiling only moves it. The
   notice below is what protects the customer when the ceiling IS reached.

2. On timeout or error, dispatch a plain sentence to the conversation instead of
   silence. The provider's raw error never reaches the customer — it carries
   model names, quota ids and URLs (litellm.RateLimitError: ... limit: 20,
   model: gemini-2.5-flash) — it goes to the operator's log as `cause`.
   AI_FAILURE_NOTICE overrides the text; setting it empty keeps today's silence
   for operators who prefer it.

Tests: 6 in pkg/pipeline/service/ai_failure_notice_test.go — the customer is
told, the provider error never leaks, the text is overridable, empty disables
it, a missing postback url is not a crash, and the default holds when the env is
unset. go build + go vet clean; full suite green
(pkg/... and internal/..., Redis-backed).

Not addressed here, deliberately: the processor keeps working after the
bot-runtime gives up, and the provider's 429/503 still surfaces as a generic
500 on the A2A route. Both are real and belong to the processor side.
@sourcery-ai

sourcery-ai Bot commented Aug 22, 2026

Copy link
Copy Markdown

Reviewer's Guide

Extends AI pipeline robustness by increasing the AI call timeout and proactively notifying customers in chat when the AI backend times out or errors, with configurable and non‑leaky messaging, plus tests to validate the behavior.

Sequence diagram for AI failure customer notification

sequenceDiagram
    participant Pipeline
    participant AIBackend
    participant Chat
    participant OperatorLog

    Pipeline->>AIBackend: runAIStage
    AIBackend-->>Pipeline: timeout or error
    Pipeline->>Pipeline: clearStateWithLog
    Pipeline->>OperatorLog: pipeline.ai.failure_notice.sending(cause)
    Pipeline->>Chat: runDispatchStage(notice)
    Chat-->>Pipeline: postback delivered
Loading

Flow diagram for configurable AI failure notice

flowchart TD
    A[AI timeout or error] --> B[clearStateWithLog]
    B --> C{AI_FAILURE_NOTICE configured?}
    C -->|empty| D[Return silently]
    C -->|unset| E[Use default notice]
    C -->|non-empty| F[Use configured notice]
    E --> G{postbackURL available?}
    F --> G
    G -->|no| H[Log no_postback and return]
    G -->|yes| I[Log provider error as cause]
    I --> J[runDispatchStage with customer-safe notice]
Loading

File-Level Changes

Change Details Files
Increase the AI call timeout ceiling to better accommodate high tail latency from the LLM provider.
  • Raise AI_CALL_TIMEOUT_SECONDS default from 30 to 90 seconds in configuration loading.
  • Document the latency rationale and trade-offs in inline comments near the timeout configuration.
internal/config/config.go
Introduce a failure-notice mechanism so customers receive a safe, configurable message when the AI backend times out or errors instead of silence.
  • Add sendAIFailureNotice helper on pipelineService that logs the provider error only to operators and dispatches a canned message to the chat.
  • Make the notice text configurable via AI_FAILURE_NOTICE env var, with empty string disabling the notice and default text used when unset.
  • Ensure the notice is only sent when a postbackURL is available, with structured logging for disabled/no-postback cases.
  • Wire sendAIFailureNotice into AI-stage error handling paths, both for timeouts and other AI errors, after state is cleared.
pkg/pipeline/service/pipeline_service.go
Add tests to guarantee the AI failure notice behavior and prevent provider details from leaking to end users.
  • Create unit tests covering timeout notification, non-leakage of provider error details, env-based customization, disablement via empty env, handling of missing postback URL, and default behavior when env is unset.
  • Introduce a small helper to capture dispatched messages and a local contains function for substring checks in tests.
pkg/pipeline/service/ai_failure_notice_test.go

Tips and commands

Interacting with Sourcery

  • Trigger a new review: Comment @sourcery-ai review on the pull request.
  • Continue discussions: Reply directly to Sourcery's review comments.
  • Generate a GitHub issue from a review comment: Ask Sourcery to create an
    issue from a review comment by replying to it. You can also reply to a
    review comment with @sourcery-ai issue to create an issue from it.
  • Generate a pull request title: Write @sourcery-ai anywhere in the pull
    request title to generate a title at any time. You can also comment
    @sourcery-ai title on the pull request to (re-)generate the title at any time.
  • Generate a pull request summary: Write @sourcery-ai summary anywhere in
    the pull request body to generate a PR summary at any time exactly where you
    want it. You can also comment @sourcery-ai summary on the pull request to
    (re-)generate the summary at any time.
  • Generate reviewer's guide: Comment @sourcery-ai guide on the pull
    request to (re-)generate the reviewer's guide at any time.
  • Resolve all Sourcery comments: Comment @sourcery-ai resolve on the
    pull request to resolve all Sourcery comments. Useful if you've already
    addressed all the comments and don't want to see them anymore.
  • Dismiss all Sourcery reviews: Comment @sourcery-ai dismiss on the pull
    request to dismiss all existing Sourcery reviews. Especially useful if you
    want to start fresh with a new review - don't forget to comment
    @sourcery-ai review to trigger a new review!

Customizing Your Experience

Access your dashboard to:

  • Enable or disable review features such as the Sourcery-generated pull request
    summary, the reviewer's guide, and others.
  • Change the review language.
  • Add, remove or edit custom review instructions.
  • Adjust other review settings.

Getting Help

@sourcery-ai sourcery-ai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hey - I've found 1 issue

Prompt for AI Agents
Please address the comments from this code review:

## Individual Comments

### Comment 1
<location path="pkg/pipeline/service/pipeline_service.go" line_range="422" />
<code_context>
+			// CRM-236: silence is indistinguishable from "the bot is ignoring you",
+			// and the tool's side effect may already be applied (the card moved at
+			// ~20s, the timeout fired at 30s). Tell the customer something.
+			s.sendAIFailureNotice(contactID, conversationID, cfg, postbackURL, err)
 		default:
 			slog.Error("pipeline.ai.error",
</code_context>
<issue_to_address>
**issue (bug_risk):** The failure notice is dispatched with a fresh background context after the AI stage has cleared its state, so a newer message can start a replacement pipeline before this dispatch completes. When the notice dispatch succeeds, `runDispatchStage` clears the state and deletes the entry for the same contact/conversation, thereby deleting or clearing the newer pipeline and causing its response to be lost.

**Triggers:** When a new customer message arrives after the AI failure has been recorded but before the fallback notice finishes dispatching.

**Suggested fix:** Revalidate ownership of the pipeline entry before dispatching and before cleanup, or route the notice through an ownership-aware dispatch path that cannot clear state belonging to a newer pipeline.
</issue_to_address>

Sourcery assessment

Needs a human reviewer. 1 finding to address first, and if the new timeout or default failure message is the wrong product decision, every affected failed AI turn will wait up to 90 seconds and may send a customer-facing message; those messages cannot be undone by reverting the change. Reverting prevents future notices and restores the old timeout, but it cannot retract notifications already delivered.

Blocking findings: pkg/pipeline/service/pipeline_service.go:422


Sourcery is free for open source - if you like our reviews please consider sharing them ✨
Help me be more useful! Please click 👍 or 👎 on each comment and I'll use the feedback to improve your reviews.

Comment thread pkg/pipeline/service/pipeline_service.go
Matheus Pastorini and others added 5 commits August 22, 2026 14:19
…-236 review)

Addresses the CRITICAL findings on the bot-runtime side.

1. The notice could delete the follow-up turn's entry.

   sendAIFailureNotice dispatched through runDispatchStage, which owns the
   turn's bookkeeping: its success path runs SetState(StageDone) -> ClearState
   -> entries.Delete(pairKey). That Delete is what runDispatchStage's own
   comments forbid ("A Delete here would race with the new event's Store and
   could delete the replacement entry").

   The race lands on the very scenario this feature targets — the customer who
   waited and follows up. They send "oi?" while the notice is being dispatched,
   startDebounce stores the new turn, then the notice finishes and deletes THAT
   entry and clears its state. The new turn is orphaned: the next message cannot
   cancel it, so two pipelines run concurrently on the same pair and the
   customer gets a duplicated reply.

   The notice now dispatches directly. There was nothing to book-keep anyway:
   both call sites already run clearStateWithLog before asking for it.

2. The 5s budget truncated the notice and then lied about why.

   cleanupCtx's 5s are documented for cleanup calls (ClearState, SetState). A
   Dispatch is a different animal: it segments by TextSegmentationLimit and
   sleeps DelayPerCharacter per rune between parts. With segmentation on, the
   84-rune default became 2-3 parts and the delays alone exceeded 5s ->
   ErrDispatchInterrupted -> the log said "New message arrived" when nothing had
   arrived. Now bounded by noticeCtx (30s).

3. The 90s ceiling was still pinned at 30 inside this repo.

   .env.example:8 and k8s/configmap.yaml:8 both set it explicitly, and by this
   PR's own reasoning ("an explicit env beats the code default") the fix had no
   effect where it runs — the ConfigMap is what actually serves staging/prod.

   NOT fixed here: evolution-ecosystem/k8s/base/bot-runtime.yaml:44-45 pins 30
   too, but that is the SaaS repo, outside these three PRs.

4. The default notice was hardcoded pt-BR (finding 9).

   bot-runtime ships in community/self-hosted installs worldwide, so customers
   of installations that never chose Portuguese were answered in it. Now English,
   with AI_FAILURE_NOTICE documented in .env.example for localisation and the
   empty value still restoring silence.

Tests: 10 (was 6). The new ones pin that the notice leaves the follow-up entry
and its state untouched, that its budget fits a segmented dispatch, and that the
default carries no pt-BR.

Negative proof: restoring the runDispatchStage call fails
"DoesNotTouchTheEntryOfTheNextTurn" with the orphaned-turn message.

gofmt: pipeline_service.go is already unformatted on develop (CRLF); left as-is
rather than reformatting the whole file into this diff.
11 — the doc comment of clearStateWithLog had been orphaned: the new constants
     landed between it and its function, so godoc showed aiFailureNoticeEnv
     documented as "clearStateWithLog calls ClearState...". Moved back.

12 — the test helper reimplemented strings.Contains in 9 lines with a closure.

13 — os.Unsetenv mutated the process env without restoring it, so every test
     running after that one inherited the change. Restores via t.Cleanup now.

No behaviour change. build/vet clean, full suite green with a real Redis.
CI was red on TestE2E_PipelineIsolation, and the test was right to fail:

    postback content = "We are having a temporary issue…", want "pair-b response"

It asserted callCount == 1 with the comment "only pair B should deliver". That
described the OLD behaviour, where pair A's AI error was swallowed — log, clear
state, nothing reaching the chat. Making that pair speak is the entire point of
this card, so the assertion had to move.

Updated rather than relaxed. The test is about isolation, and it now checks that
more strictly than before: both pairs deliver, and each must receive ITS OWN
message. A crossed delivery fails here even though the count would be right —
the previous version could not have caught that.

The notice is pinned through AI_FAILURE_NOTICE instead of reaching for the
package constant: it keeps test/e2e out of service's internals and exercises the
operator-facing env on the way.

On how this reached CI: I ran ./pkg/... and ./internal/... and called it "the
full suite". test/e2e was never in it. Same mistake as the migration parity spec
on CRM-210 — a hand-picked list reported as a green suite. Ran ./... this time.
…eview)

The round-1 defect was runDispatchStage's success bookkeeping running for the
failure notice: SetState(StageDone) -> ClearState -> entries.Delete(pairKey),
against a pair that may already belong to a follow-up turn. The fix is in, but
nothing pinned it: the existing test read Redis after the fact, and SetState
followed by ClearState leaves nothing to read, so it passed with or without the
regression.

DoesNotTouchTheEntryOfTheNextTurn seeds the pair the way startDebounce leaves it
(entry in the map, StageDebounce in Redis) and asserts both survive. Restoring
the runDispatchStage call fails all three assertions.

Also make DoesNotWriteTurnState fail on a read error instead of passing on it,
and drop its unused rdb binding.
… 14)

The measurement table, the incident timeline and the race walkthrough belong in
the PR body, where they already are. In the source they were 11 lines above one
call to getEnvIntOrDefault, 18 above one call to Dispatch, and 10 above a
context.WithTimeout.

What is kept is what the code cannot say: why 90 and not 30, why the notice must
not go through runDispatchStage, why noticeCtx is not cleanupCtx. 102 added
comment lines to 58, no behaviour touched.
@gomessguii
gomessguii merged commit 7ec7db4 into develop Aug 22, 2026
5 checks passed
@gomessguii
gomessguii deleted the fix/CRM-236-degraded-provider-feedback branch August 22, 2026 23:13
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants