Skip to content

fix(sdk-broker): stop waiting on trampoline close - #6225

Closed
probepark wants to merge 122 commits into
mainfrom
gc-g4-broker-reparent-r2
Closed

probepark wants to merge 122 commits into
mainfrom
gc-g4-broker-reparent-r2

Conversation

@probepark

Copy link
Copy Markdown
Collaborator

Agent

omo/cliproxy/gpt-6.1-sol via paseo on studio (worktree gc-g4-broker-reparent-r2-r5)

Task

REVIEW_FIX for PR #6221 (branch gc-g4-broker-reparent-r2, head 182b1ac, base dev).
Work on top of branch gc-g4-broker-reparent-r2. Do not rebase onto dev unless needed; keep the PR's intent.

Goal

Fix the 2 PR-caused CI failures (run 36942588664) so "Affected path validation" goes green.

Where to change

  1. packages/coding-agent/src/sdk/broker/ensure.ts — biome format violations (job 110641756818):
    L419 child.on("error", () => { }); -> () => {}; L501-504 Boolean(discovery && ...) indentation.
    Run bunx biome format --write packages/coding-agent/src/sdk/broker/ensure.ts (this file only).
  2. packages/coding-agent/test/sdk-broker-reparent-process.test.ts — on ubuntu-22.04 CI the test
    reparents the shared broker outside the spawner process tree times out at 30000ms
    (job 110641756844, 0 pass 1 fail, "killed 1 dangling process"). On macOS it passes in 7.7s.
    Find the root cause FIRST: waitForFile (200x25ms = 5s) should throw if the ready file never appears,
    so reaching 30s means ensureBroker never resolves or await parent.exited never resolves on Linux.
    Check the Linux setsid/detach path in ensure.ts / runtime.ts (e.g. detached child keeping stdio pipes
    open so the parent never exits, missing unref(), stdio inherited instead of "ignore", setsid binary
    absent/behaving differently on Linux). If the cause is in src (Linux reparent path), fix it there.
    Only raise the test timeout if you can

Local CI gate (prepush)

PREPUSH_CMD rc=0 bun run --workspaces --if-present check:types
PREPUSH_CMD rc=0 bun scripts/telegram-daemon-generation-guard.ts "$(git merge-base origin/dev HEAD)" "$(git rev-parse HEAD)"

Agent log (tail)

Failed to load pi_natives native addon for linux-arm64.

  • No timeout increase was retained.

Verification

  • bunx biome check packages/coding-agent/src/sdk/broker/ensure.ts packages/coding-agent/test/sdk-broker-reparent-process.test.ts
    • Exit 0
    • Checked 2 files ... No fixes applied.
  • bun test packages/coding-agent/test/sdk-broker-reparent-process.test.ts
    • Exit 0
    • 1 pass, 0 fail
  • bun --cwd=packages/coding-agent run check
    • Exit 0
    • Type check passed.
    • One existing warning remains for src/session/agent-session.ts exceeding the configured 1 MiB file-size threshold.
  • LSP diagnostics:
    • No errors on either changed file.
    • One existing TypeScript hint in ensure.ts about converting a function to async.
  • git show --check
    • Passed.
  • Final commit changed exactly:
    • packages/coding-agent/src/sdk/broker/ensure.ts
    • packages/coding-agent/test/sdk-broker-reparent-process.test.ts

Commits

  • 82b3ebe9b — fix(sdk-broker): stop waiting on trampoline close
  • Existing PR-scope commits retained:
    • 233f3c6bb — formatter corrections
    • b2706c7a2 — earlier bounded startup attempt, superseded in behavior by the final test restoration

Assumptions

  • Ubuntu CI has the expected Linux native addon available; the local Docker reproduction could not validate runtime behavior because that addon was not built for linux-arm64.
  • The existing large-file Biome warning and async hint are pre-existing and unrelated to these changes.
    [Thought] Confirming root cause with code evidence

Documenting local verification with assumptions
[Compacted]
[Compacted]

gjc-stack and others added 30 commits September 30, 2026 20:27
The embedded codex-pro profile still routed core roles through GPT-6 Sol and Terra. Point the builtin profile at GPT-6.1 Sol while preserving Astra for architecture, and lock the role mapping in the catalog regression test.
…-opus-codex intent

Address review on #6163:
- codex-pro is now a like-for-like Sol generation swap: executor stays on
  gpt-5.6-terra:medium, critic keeps :max, architect stays Sol (gpt-6.1-sol:xhigh)
  instead of moving to gpt-6-astra. The all-6.1 + Astra-architect mapping
  remains available as codex-sol61.
- packages/ai fragment now only documents the af1dc0f catalog entry; the preset
  change gets its own coding-agent fragment, including the registry precedence note.
- fable-opus-codex test asserts it still projects codex-pro executor and architect
  tier/effort while intentionally staying on the gpt-6-sol generation.
Every built-in profile role bound to openai-codex/gpt-6-sol now binds
openai-codex/gpt-6.1-sol at the same effort. Opus bindings are already on
claude-opus-5-5. The gpt-6-sol catalog entry stays for explicit selection.
…dex-sol61

codex-pro now carries GPT-6.1 Sol, so the additive codex-sol61 profile from
#6152 is redundant and is removed. All built-in profiles that bound
gpt-6-sol (codex-medium, codex-pro, astra-*, opus-codex, codex-opencodego,
fable-opus-codex) move to gpt-6.1-sol at the same effort. The two changelog
fragments are merged into one.
Keep persisted 0.18.2 codex-sol61 defaults working while preserving user-defined profile precedence.
Describe legacy resolution precedence and the catalog context limit for migrated Sol profiles.
…-only

Address re-review on #6163:
- Every built-in Codex role on gpt-5.6-terra now uses gpt-6.1-sol at the
  same effort; codex-medium and codex-pro are GPT-6.1 Sol in every role.
- codex-eco is codex-medium with Sol lowered to gpt-6-luna at the same
  effort. The catalog test asserts that derivation instead of pinning
  literals.
- Provider-agnostic open-weights-* profiles keep gpt-5.6-luna: gpt-6-luna
  is only on openai-codex, while gpt-5.6-luna is on 12 providers.
- Changelog documents the Terra tier move, the 372K to 272K compaction
  change for those roles, and how codex-pro differs from the removed
  codex-sol61 mapping.
feat(models): move built-in Codex profiles to GPT-6.1 Sol, Luna-only codex-eco, drop codex-sol61
…ed message (#6159)

* fix(kiro): non-eventstream 200 body is masked as eventstream: truncated message

When CodeWhisperer returns HTTP 200 with a non-eventstream body (JSON/HTML error),
the first 4 bytes are parsed as a frame length and the user only sees
'eventstream: truncated message at end of stream', hiding the real server error.

Before decoding, check if the response content-type is application/vnd.amazon.eventstream.
If not, read the text body and throw an error with status, content-type and the first
1000 chars of the body, following the existing !response.ok pattern with withHttpStatus.

Also improve the truncated-stream error in aws-eventstream.ts to report how many
trailing bytes were left.

Tested: new kiro-codewhisperer-endpoint test mocking 200 application/json response
with error message, asserting the message is surfaced instead of truncated error.

* fix(kiro): match the eventstream media type case-insensitively and add #6158 changelog fragment

Test the error event the stream actually emits (the stream never throws), cover a
mixed-case media type with parameters, and assert the bearer token never leaks.

* fix(kiro): bound the non-eventstream diagnostic read and redact the bearer token from it

* fix(kiro): keep a non-eventstream 2xx out of retryable transport failures

* fix(kiro): name the 2xx status in the non-eventstream diagnostic without making it parseable

* fix(kiro): never parse a status out of the non-eventstream diagnostic body

* fix(kiro): mark the non-eventstream diagnostic terminal for session retry

* fix(kiro): classify a bodyless 2xx as a protocol mismatch

* fix(kiro): classify any bodyless 2xx as a protocol mismatch regardless of content type

* fix(kiro): sanitize non-eventstream body separately so the prefix keeps its 1000-char budget (#6158)

---------

Co-authored-by: Gajae Bot <bot@gajae.dev>
…6166) (#6169)

- Set userInputMessage.origin=AI_EDITOR; without it CodeWhisperer ignores
  modelId and routes every request to auto.
- Send every trailing result of a parallel tool batch in currentMessage;
  splitting them across history and currentMessage is rejected with
  TOOL_USE_RESULT_MISMATCH (HTTP 400).
* fix(auth): recover unavailable pins during profile startup

A resumed session can pass saved-model restoration and still crash when its configured default profile probes the same unavailable pinned provider. OAuth selection can also invalidate a pin after the saved-model preflight check. Classify scoped pin failures as recoverable profile credential errors and skip only explicitly configured fallback candidates without selecting a different account.

Confidence: high

Scope-risk: auth-startup

Reversibility: simple-revert

Tested: 299 focused tests, coding-agent check, state-writer gate

Not-tested: live upstream resume against the reporter's private credential store

Supersedes: e94acf6

* docs(changelog): note resume profile credential recovery

The prior durable-pin fix was already released, so this user-visible follow-up needs its own fragment for the next release.

Confidence: high

Scope-risk: low

Reversibility: simple-revert

Tested: bun scripts/changelog-fragments.ts check

* fix(auth): honor runtime and models.yml keys over unavailable pins in profile startup

Profile activation, default-chain recovery, and startup availability checks now
treat an unavailable session pin as blocking only when no runtime --api-key or
owner-scoped models.yml provider key exists, matching AuthStorage.getApiKey
precedence. Addresses snowykr P2 on #6025.

* fix(auth): honor explicit keys over unavailable pins when restoring resumed models

The saved-model restore, settings-default fallback, and post-extension
retry still consulted the raw unavailable-pin marker, so a runtime
--api-key or this registry's models.yml key could not restore the
resumed model even though AuthStorage.getApiKey would resolve it.
Route all five startup pin checks through isSessionCredentialPinBlocking
so they share getApiKey precedence; stored accounts stay blocked.

Addresses snowykr review on session.ts:2250-2252, :2303-2307, :4243-4246.

* test(auth): guard late-registered pinned providers without explicit keys

Cover the post-extension retry's blocking branch: an extension-registered
saved model whose pin is unavailable and has no runtime key must stay
unselected and expose no stored-account key. Pin the pre-extension
restore assertion and narrow the changelog to the explicit-key cases
the precedence helper actually covers.

* fix(auth): keep apiKeyEnv keys from retargeting unavailable pins

isSessionCredentialPinBlocking counted any owner-scoped models.yml key
as explicit. For an apiKeyEnv key, AuthStorage.getApiKey prefers a
ranked stored api_key account before it checks the unavailable pin, so
a resumed session with a stale pin could restore its model on a
different stored account instead of prompting to re-pin. Only a runtime
key or a literal models.yml apiKey now unblocks the pin, matching
peekApiKey; add AuthStorage.hasLiteralConfigApiKey to expose that.

* test(ai): cover hasLiteralConfigApiKey owner and env-sourced cases

Pin the accessor contract the startup pin helper relies on: an apiKeyEnv registration is a config key but not a literal one, a literal re-registration flips it, and another owner never sees it.

* docs(auth): scope literal config key wording to startup pin policy

---------

Co-authored-by: Bellman <54757707+Yeachan-Heo@users.noreply.github.com>
Co-authored-by: Gajae Bot <bot@gajae.dev>
A failed or cancelled prompt can be settled at the ACP boundary while the SDK loop is still winding down. Retry a busy acknowledgement only after the session publishes idle, with a bounded watchdog deadline.
Exercise both terminal failure and immediate cancellation paths while the SDK remains busy, and verify the next prompt is accepted after idle.
The router may surface a busy response as SdkClientError. Treat it like the adapter error and retry once after the bounded winding-down wait.
Make the fixture return a real busy control error on the successor attempt and cover both failed and cancelled predecessor turns.
…pt cache survives (#6167) (#6168)

* fix(session): maintain volatile context cache prefix by keeping ephemeral messages in agent.state

Volatile project context and untrusted MCP server instructions were being removed
from agent.state.messages immediately after being sent, which destroyed the cache
prefix and caused every subsequent request to diverge at that early index.

This fix keeps ephemeral messages in agent.state.messages after they are sent,
allowing them to serve as the prefix for subsequent requests. Each new request
appends a fresh copy of volatile context before the new prompt, extending the
cache prefix naturally.

Ephemeral messages are still excluded from persistent storage (sessionManager
branch and compaction input) as intended.

Regression test added to verify prefix-extension behavior across multiple turns.

Fixes #6167

* fix(session): drop retained volatile context on rescope and exclude it from handoff input (#6167)

* fix(session): keep retained volatile context out of fork seeds, contribution prep, and session export (#6167)

* test(session): use named ContributionPrepResult type instead of ReturnType (#6167)

* fix(session): keep retained volatile context out of side-channel ephemeral turns (#6167)

* fix(session): skip appending an ephemeral copy identical to the latest retained one (#6167)

---------

Co-authored-by: Gajae Bot <bot@gajae.dev>
A busy acknowledgement from a cancelled preflight belongs to the retired prompt and must follow cancellation settlement. Gate the settle wait and retry on the waiter remaining active and uncancelled, with a regression covering the delayed busy response.

Confidence: high
Scope-risk: narrow
Reversibility: easy
Tested: sdk-acp-production-path, sdk-acp-prompt-terminal, biome, tsc
Not-tested: none
The catch path replaces every thrown run failure with 'Agent run failed.' so
provider errors never leak into transcripts. The raw cause was not recorded
anywhere, which left frequent failures (e.g. pre-stream credential refresh
errors) undiagnosable. Log it at warn with redactCrashSecrets applied and the
text bounded; transcript and SDK surfaces are unchanged.
… Linux (#6179)

The resident-cache GC reaped the instance directory of a live session on
Linux hosts whose wall clock is slewed (WSL2 in particular), so later turns
replaced tool results with "[Session resident text blob missing: ...]" and
logged "Resident cache trust rejection ... blob_create_failed ENOENT".

The owner lease recorded the process start from `ps -o lstart`. procps
renders that as btime + starttime ticks, and btime is re-derived from the
current wall clock minus uptime on every read. When the clock is slewed the
same live process reports a start time that drifts (observed: a lease
written at 03:08:21 UTC read back as 03:23:43 UTC eleven hours later), and
residentCacheOwnerIsStale read any mismatch as PID reuse.

On Linux the lease now stores /proc/<pid>/stat field 22 (start ticks since
boot), which never moves for a live process, under the basis
`linux-proc-ticks:<boot_id>` so tick counts are only compared within one
kernel boot. Other platforms keep the UTC-pinned `ps` render. Staleness is
only proven when the recorded basis equals the reader's basis, so leases
from older builds (basis "utc" or unlabelled) or another boot are never
reaped while their PID is alive; dead PIDs are still collected.
Wait for a cancelled prompt to settle before admitting one successor, while preserving immediate conflicts for active turns.
Concurrent writers in one managed scope could promote or retire another
writer's live receipt, turning a valid append into identity_mismatch.
Bind receipts to their exact publishing attempt and defer peer recovery
until owner exit or PID reuse is proven. Keep exact orphan and transcript
identity checks intact.

Lore-id: 56bbf432
Constraint: preserve live publishers and external replacement fences
Rejected: refresh cached inode after failure | can overwrite another writer
Tested: 126 focused tests passed; 8 platform skips; two-process same-scope writes
Tested: package typecheck and touched-file Biome checks
Not-tested: full suite, native rebuild, production rollout
Confidence: high
Scope-risk: managed-storage
Reversibility: coordinated-writer-rollout
Wait for owner-fenced Telegram control requests before force escalation on hard-termination platforms, and report post-update recovery failures without uncaught exceptions.
Owner-fenced control requests must interrupt in-flight Windows long polls before controller escalation.
Yeachan-Heo and others added 20 commits October 1, 2026 17:43
fix(acp): admit a fresh prompt after the previous turn settles while the SDK winds down
fix(exec): stop bash shell supervisor busy-spinning while idle (#5972)
… stamps row (#6146)

* fix: remove message and artifactDigest from unconditional startupFailure construction

* fix: only include message and artifactDigest in startupFailure for recovery-stamped terminal_uncertain responses

* fix(lifecycle): properly type-guard response.error access and always include artifactDigest in startupFailure

When evidence exists and response is a failure, properly type-guard access to response.error
using !response.ok before accessing response.error.code. Always include artifactDigest and
message from the evidence artifact to maintain trustworthy durable evidence binding.

---------

Co-authored-by: Gajae Bot <bot@gajae.dev>
…unting (#6199)

AgentSession.messages now hides the ephemerals that #6167 retains in
agent.state.messages, but four internal accounting paths still read that
public view: the usage anchor lookup, the usage cache key, the compaction
token correction ratio, and the context token estimate. They undercounted
the real provider request, so auto-compaction could be suppressed until the
window overflowed. A 15,000 character retained ephemeral left the estimate
unchanged at 46 tokens.

Read agent.state.messages in those paths so the anchor, indices and estimate
share one authoritative list. Transcript-facing readers keep the public view.

Found by the boundary architect and QA lanes on the 5951386 change set.

Constraint: #6167 cache-prefix behavior must stay intact
Rejected: unfilter the public getter | reintroduces the dev CI failure
Confidence: high
Scope-risk: narrow
Tested: volatile-context-prefix (new test fails without the fix), fallback-transaction, 10 compaction and context-usage files; coding-agent check
Not-tested: full suite (EBADF spawn failures locally)

Co-authored-by: Yeachan-Heo <yeachan-heo@gajae.dev>
… word-diff A/B adapter (#6194)

* feat(natives): add word-diff A/B bench adapter

* fix(natives): drive the word-diff adapter through renderDiff so base and head share an entrypoint

* fix(natives): init theme and settings in the word-diff adapter

* docs(rust-porting): adopt diff, edit, js-testing and highlight rows; rename non-fallback identifiers

Rows A-PI-DIFF, B-DIFF, A-PI-EDIT, B-EDIT, B-JS-TESTING and C-HIGHLIGHT have merged PRs and A/B evidence. Identifiers named *Fallback in edit paths tripped the no-fallback gate although none is a native-to-TS fallback.

* docs(rust-porting): record 2.7 adopted and 2.8-2.10 rejected under D9

* docs(rust-porting): record the tty-write A/B result for B-TTY-WRITER

* docs(rust-porting): separate scenario-wide and per-token-window E-M01 attribution and cite the right run

* feat(natives): add word-diff A/B bench adapter

* fix(natives): drive the word-diff adapter through renderDiff so base and head share an entrypoint

* fix(natives): init theme and settings in the word-diff adapter

* docs(rust-porting): adopt diff, edit, js-testing and highlight rows; rename non-fallback identifiers

Rows A-PI-DIFF, B-DIFF, A-PI-EDIT, B-EDIT, B-JS-TESTING and C-HIGHLIGHT have merged PRs and A/B evidence. Identifiers named *Fallback in edit paths tripped the no-fallback gate although none is a native-to-TS fallback.

* docs(rust-porting): record 2.7 adopted and 2.8-2.10 rejected under D9

* perf(natives): transcode TtyWriter output straight from N-API UTF-8 into the pump buffer

Shaves the JS-thread enqueue cost of a single frame (4.5us -> 2.0us on an 8KB frame) and wakes only the pump on enqueue.

Tested: cargo test -p pi-natives --lib tty_writer (7 pass); PTY byte parity with Buffer.from(utf8) incl. NUL and lone surrogates
Not-tested: Windows (unix-only module)

* docs(rust-porting): record tty-write A/B after the direct-UTF-8 enqueue change

* fix(natives): write N-API UTF-8 into spare capacity so clippy::uninit_vec passes

---------

Co-authored-by: Yeachan-Heo <yeachan-heo@gajae.dev>
…chema (#6204)

`web_search` declared `limit` and `num_search_results` as plain
`z.number()`, so a model could send `0`, `-3` or `2.5`. Providers
forward the value as `num_results`, and openai-compatible, xai and kimi
slice sources with it, so a negative count silently dropped results
from the end. Use `z.number().int().min(1)`, the same constraint
`search_tool_bm25` already applies to its `limit`, and regenerate the
tool catalog.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…6203)

`Flags.integer` accepts `0` and negative integers, and `runSearchCommand`
only rejected NaN. The value then reached providers as `num_results`,
and openai-compatible, xai and kimi slice sources with it, so
`--limit=-3` silently dropped the last three results. Require a
positive safe integer. Also stop `parseSearchArgs` from keeping the
leading digits of values like `5x` or `2.5`.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* fix(ssh): reject non-integer --port in `gjc ssh add`

`gjc ssh add <name> --host <addr> --port 22oops` passed validation and
saved port 22, because `Number.parseInt` stops at the first non-digit.
`2222.5`, `+22`, ` 22` and `1e3` were coerced the same way. Require the
whole token to be digits before parsing, and store the validated value
instead of parsing the flag a second time. Interactive `/ssh add` got
the same fix in #6133.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(ssh): restore the stdout.write spy after each ssh-cli port test

The spy was installed in beforeEach and never restored, so each test
stacked another mock on the global. Keep the handle and call
`mockRestore()` in afterEach, as `stats-cli.test.ts` does.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* fix(sdk): flush pending agent_end event after auto compaction

After auto-compaction resets the message history, a pending agent_end event
may have been parked before or during compaction. Without flushing it, the
next prompt fails with 'Timed out waiting for prior agent run to finish before
prompting' because the agent remains in a busy state.

Fixes #6004

Tested: regression test added to ensure prompt succeeds after auto-compaction
without timing out.

* chore: revert unrelated read.md edit

* test(sdk): add deterministic regression test for issue #6004

Add a new test that explicitly verifies the fix for auto-compaction
pending agent_end flushing. The test:

1. Drives auto-compaction
2. Manually parks an agent_end event using test seams
3. Calls applyCompactionPostAppendForTests (which contains the fix)
4. Verifies that the parked agent_end is published/flushed

The test runs 5 times without timing races to ensure deterministic
behavior. Without the fix in #applyCompactionPostAppend, the pending
agent_end would remain parked and cause the next prompt to timeout
waiting for the agent to become idle.

Regression test for issue #6004.

* fix(sdk): publish pending agent_end after compaction when no continuation is scheduled

After auto-compaction completes, if no continuation prompt is scheduled, any pending
agent_end event must be published to allow the run to settle properly. This prevents
the session from becoming wedged with 'Timed out waiting for prior agent run' errors
on subsequent prompts (#6004).

The fix:
- Adds #flushPendingAgentEndAfterCompaction() that bypasses the handler-in-flight guard
- Publishes pending agent_end in overflow recovery path when no continuation is scheduled
- Preserves continuation logic by not publishing if a continuation will run
- Normal handler unwinding (line 7038) handles non-overflow compaction paths

Threshold compaction paths will publish the agent_end through normal handler
unwinding, but overflow recovery needs explicit flushing here to avoid deadlock
when the continuation isn't reachable.

* chore: drop unrelated read.md edit

* fix(sdk): flush pending agent_end after auto-compaction to prevent timeout on next prompt

After threshold auto-compaction resets the message history, a pending agent_end
event may have been parked before or during the compaction. If this event is not
flushed before the compaction completes, the session's persistent state will not
properly mark the end of the current agent run.

When the session is later loaded in a new process, the agent_end event has been
lost (it's an in-memory variable), and the turn appears unfinished. This causes
the next prompt() call to timeout waiting for the agent to finish its prior run.

The fix is to call #flushPendingAgentEnd() at the end of #applyCompactionPostAppend()
for both threshold and overflow auto-compaction paths. This ensures the agent_end
event is properly published and persisted before the compaction is complete.

Fixes #6004

Test: Added deterministic regression test that verifies prompt() succeeds after
auto-compaction without timing out. Test passes 5/5 runs.

* chore: add changelog fragment for issue #6004 fix

* fix(sdk): properly flush pending agent_end after overflow compaction

- Remove ineffective flush call at threshold path (line 16434) - it's a no-op because handlers are in flight
- Remove problematic flush call at overflow path (line 22767) - it bypasses handler barrier and publishes with undefined lease
- Move flush to after auto_compaction_end event is emitted, with proper handler barrier in place
- Remove unused #flushPendingAgentEndAfterCompaction() function that bypassed safety checks
- Update test to exercise real overflow path instead of mocked threshold path

The fix ensures pending agent_end is published only after:
1. auto_compaction_end event is emitted to all handlers
2. Handler barrier properly protects the publication
3. eventLease is properly transferred before publication

Fixes issue #6004 where pending agent_end could timeout on next prompt after compaction.

* fix(test): remove incorrect assertion from auto-compaction regression test

The test was checking if promptResult is defined, but session.prompt()
returns Promise<void> which resolves to undefined. This was causing a
false assertion failure that masked the actual test intent: verifying
that the prompt does not time out after overflow compaction.

* fix(sdk): remove dead code and revert unrelated blank-line change

- Remove unreachable flush code after overflow compaction (lines 22751-22758)
  that was guarded by handler barriers, making it unreachable
- Revert unrelated blank-line removal in packages/natives/native/index.d.ts

* fix: correct changelog and test documentation for auto-compaction changes

- Change changelog from 'Fixes #6004' to accurately describe as test-only
- Update test name from 'regression test for issue #6004' to 'allows prompting after overflow compaction completes'
- Remove misleading claims that test reproduces the issue described in #6004
- The test passes at base without code changes, so cannot be a regression test for #6004
- Test exercises the overflow path to verify prompts work after compaction transitions

---------

Co-authored-by: Bellman <bellman@example.com>
Co-authored-by: Gajae Bot <bot@gajae.dev>
Co-authored-by: gaebal-gajae <clawdbot@users.noreply.github.com>
Memory maintenance must work with models that do not expose reasoning controls.

Co-authored-by: probe <probe@gjc.dev>
The manual paste prompt asked for "the authorization code (or full
redirect URL)" without saying the redirect is a zcode:// custom-protocol
URL that browsers never surface in the address bar. Users pasted the
chat.z.ai address-bar URL, which carries no code param, so parsing found
no code and the login silently re-prompted with the same wording. State
the exact zcode://oauth/callback?code=…&state=… shape in the auth
instructions, point at DevTools → Network as the only place that URL is
visible, and give glm-zcode its own manual-input prompt so a re-prompt
after a rejected paste stays self-contained.

Lore-id: glm-zcode-paste-guidance
Constraint: instructions must keep the unofficial/ToS and desktop-app (broker error 2007) warnings — pinned by existing tests
Tested: bun test packages/ai/test/oauth-glm-zcode.test.ts (32 pass; the 2 new guards fail on the pre-change source)
Not-tested: live Z.AI broker exchange (code path unchanged)
Confidence: high
Scope-risk: narrow
Reversibility: trivial
Restore Biome formatting after the memory thinking-effort fix.
…guidance

fix(ai): name the zcode:// callback shape in glm-zcode login guidance
* test(acp): cover newSession phase timing

Capture success and attach-failure logging contracts for ACP newSession.

* fix(acp): log newSession phase timings

Emit one debug event for each ACP newSession outcome without changing the response contract.
A detached trampoline lets the owned broker reparent before the spawner can be process-tree reaped, while startup evidence remains bound to the real broker identity.
@probepark

Copy link
Copy Markdown
Collaborator Author

Opened by mistake by an automated continuation run. The real PR for this work is #6221 (head 871b7d75d, fork branch with the same name), and Affected path validation is green there. This PR targets main, carries 136 files of unrelated dev history, and should not be merged. Leaving close and branch cleanup to a maintainer.

probepark pushed a commit to probepark/gajae-code that referenced this pull request Oct 2, 2026
Port of 82b3ebe (pushed by mistake to Yeachan-Heo:gc-g4-broker-reparent-r2 / Yeachan-Heo#6225) onto the Yeachan-Heo#6221 head.
@probepark

probepark commented Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator Author

Opened by mistake (coder) — superseded by PR #6221

This PR was opened automatically by our mac-agent continuation step, and that was a bug. The REVIEW_FIX run for PR #6221 (head probepark:gc-g4-broker-reparent-r2, base dev) checked for an open PR on the upstream repo's branch of the same name. It found none, so it pushed the fix commit to Yeachan-Heo:gc-g4-broker-reparent-r2 and opened this PR against main. That's why the diff is 136 files and 100 commits: it's every change dev has that main doesn't, plus the broker fix.

A maintainer can close it and delete the stray upstream branch gc-g4-broker-reparent-r2. We don't close PRs or delete branches in this repo.

@probepark

Copy link
Copy Markdown
Collaborator Author

CI triage (coder) — no fix pushed to this PR

Run 36948368582 @ 82b3ebe failed in test (coding-agent shards 1 and 12):

  • AC10c: usable autorouting tiers and disabled autorouting produce no thought warning
  • SDK broker identity and discovery > terminates and reaps the spawned broker when discovery times out (29s timeout)

No fix is going here on purpose. As noted above, this PR was opened by mistake. It targets main and carries 136 files of unrelated dev history, so its CI result says nothing about the broker change. That work lives in #6221 (base dev). Any broker-test failure will be triaged and fixed there, not on this branch.

Still recommended: close this PR and delete the upstream branch gc-g4-broker-reparent-r2. That is left to a maintainer.

@probepark probepark closed this Oct 2, 2026
@probepark

Copy link
Copy Markdown
Collaborator Author

Closing: this PR was opened by mistake by automation. A follow-up run for #6221 pushed to an upstream branch that has the same name as #6221's fork branch and opened it against main. The work belongs to #6221 (base dev). Nothing here needs review.

Yeachan-Heo pushed a commit that referenced this pull request Oct 2, 2026
…ess tree (#6221)

* fix(sdk-broker): reparent shared broker outside spawner tree

A detached trampoline lets the owned broker reparent before the spawner can be process-tree reaped, while startup evidence remains bound to the real broker identity.

* fix(sdk-broker): force-reap trampoline-reported broker on discovery timeout

Once the trampoline has reported the real broker pid and exited, the
retained child handle no longer owns the broker, so a graceful SIGTERM
cannot be observed within the discovery deadline. Reap the reported
identity with SIGKILL (identity-fenced against pid reuse). No timeout or
retry budgets change.

* test(sdk-broker): bound reparent startup failures

* style(sdk-broker): biome format ensure.ts

* fix(sdk-broker): stop waiting on trampoline close

Port of 82b3ebe (pushed by mistake to Yeachan-Heo:gc-g4-broker-reparent-r2 / #6225) onto the #6221 head.

* Revert "fix(sdk-broker): stop waiting on trampoline close"

This reverts commit e8128d2.

---------

Co-authored-by: probe <probe@gjc.dev>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

9 participants