Repository navigation
ci: wait out the uninstall before the crash harness reinstalls - #54
Merged
Merged
Conversation
The Android crash harness sometimes failed a native case with "expected exactly one native record, received []" although the OS had filed the crash. The fault is in the harness, not in the library. _Device.reinstall uninstalled the app and installed the next build at once, but adb uninstall returns when the package manager is done, not the activity manager. The activity manager drops the uninstalled package's exit records later, when the removal broadcast reaches it, and it drops them by package name, which takes the new install's records with them. In the two failures that left a log, the broadcast landed 16 and 17 seconds after the uninstall, after the new install had crashed and the OS had filed its record. In 27 passing runs it landed 0.2 to 1.8 seconds after the uninstall, long before the crash. reinstall now dumps the exit history, uninstalls, and polls until the dump lists no record and the history has been written since, every 500 ms for at most 120 s and then a StateError, before it installs. It prints how long that took. The wait is on the OS's own state: no assertion is relaxed, no run is retried and the library is unchanged. The readings of the dump move to tool/exit-info.dart, tested against dumps captured from the CI emulators, and a launch 2 that received no record now says whether the OS still holds it or removed it. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XLy7TiPAxtVjPKEwQ9iwBa
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this fixes
The Android crash harness fails a native case now and then with
launch 2: expected exactly one "native" record, received [], although the OS had filed the crash. It is the harness, not the library: the library path is unchanged and nothing it asserts is touched.Cause
_Device.reinstallranadb uninstalland thenadb installstraight away.adb uninstallreturns when the package manager is done, not the activity manager. The package manager then broadcastsACTION_PACKAGE_REMOVED; when the activity manager's receiver gets it,AppExitInfoTracker.onPackageRemoveddestroys every exit-info container filed under the package name (the new install's uid included, replacements excepted viaEXTRA_REPLACING) and persists the history. When that broadcast lands after the new install has already crashed and had its record filed, the record is destroyed with the old install's, the launch that follows reads nothing, and the crash is never reported. AOSP source read: android14-releaseonPackageRemoved, its receiver, andremovePackageLocked; android11-release has the same code.Evidence
adb uninstallreceived []received []In the failing job the exit-info at the end of the run holds only the force stops of the new uid, and its
Last Timestamp of Persistenceis 22:04:11.860, the moment the removal was handled; the OS had reported the crash filed at 22:04:09. In both failures the broadcast landed after the crash, in every passing run well before it. The two dumps are the fixtures of the new tests.Ruled out: the double launch (a symptom), force-stop eviction (it destroys a whole container and was filed later), a tombstone not written (it is written before
am_crash), a relaunch before the exit record exists (waitForExitRecordsaw it), and the watermark's granularity.The change
tool/exit-info.dart(new, pure, every public member doc-commented):holdsExitRecords,exitInfoPersistedAt(throwsFormatExceptionwhen the line is missing),holdsExitRecord(moved out ofwaitForExitRecord) andremovalSettled, which holds when no record is listed and the history has been persisted since the dump taken before the uninstall. The timestamp also covers an old install that had no records at all._Device.reinstalldumps the exit history, uninstalls, and, if something was uninstalled, pollsremovalSettledevery 500 ms for at most 120 s (aStateErrorafter that, as_rundoes whenadbfails), printsprevious install left the exit history after <ms> ms, and only then installs. Its doc comment gives the mechanism and whyupgrade()needs no wait (a replacing install is ignored by the receiver)._ReleaseRefusalalready goes throughreinstall.waitForExitRecordusesholdsExitRecord.<ts>). The assertion above it has already failed, so no verdict changes.test/exit-info_test.dart(new, 19 tests) with fixtures taken from the captured dumps; README section "Android: host-driven, on an emulator" updated.Not done, on purpose: no library change, no assertion relaxed, no retry. The only wait is on the tracker's own state.
.github/workflows/crash-harness.ymlis not touched (the emulator pin is on another branch).Verification so far
dart format --output=none --set-exit-if-changed .: clean, 49 files, 0 changed.flutter analyze: No issues found (public_member_api_docsis an error here).flutter test: 268 tests pass, 19 of them new; the new file failed to compile beforetool/exit-info.dartexisted.reinstallexercised against a stateful fakeadb(scratch copy, not committed): wait of 3.1 s when the removal landed after 3 s, no wait when the uninstall fails (not installed), a timestamp-only wait of 2.1 s for an old install without records, and aStateErrorwith no install when the removal never lands (cap shortened to 3 s for the test).Verification results
Every run below is on this branch's head, 436e5ee, and green on both API levels. The two jobs of each run passed every case (API 30: 12
PASSlines, API 34: 32), and none of them printed aFAILor a problem line. The concurrency group iscrash-harness-${{ github.ref }}; the pull_request run isrefs/pull/54/mergeand the dispatches are the branch ref, so no run cancelled another. The first dispatch (37871241036) was started together with the pull_request run; the other three were dispatched one at a time, each after the previous run had finished.That is 210
previous install left the exit history after <ms> mslines. API 30: 38 to 533 ms. API 34: 27 ms to 13125 ms, and 6232, 5963 and 13125 are the only ones over 5 s.One wait of 10 s or more: 13125 ms, in run 37875030890, API 34 job 113641543061, the
nativecase that followsjvm. It is the first real uninstall on that fresh boot, which is where the largest wait fell in four of the five API 34 jobs. The two failing runs in Evidence had their removal handled 16 to 17 s after the uninstall and their new install crash about 13 to 15 s after it; a removal 13 s after the uninstall sits in that same window. The harness waited it out, installed afterwards, and the case passed withreceivedcarrying the one native record. It was not seen to fail without the wait; this proves the wait happens and the case passes, not that the old harness would have failed that exact run.No API 34 job hung on emulator 37.1.11, none reached the 120 s cap, and no assertion, retry or case was touched.
Other checks on the head sha, all green: analyze and test, kotlin unit tests, swift unit tests and ios build, Super-linter, Trivy, title and squash-message conventional-commit checks.
Main has moved since the branch point (39e6ef6 to 20b8ac6, #55). It changed
.github/workflows/crash-harness.ymlonly, which this branch does not touch, so the overlap is empty and the branch was not rebased or re-merged. #55 frees memory and keeps the logs on every run; it does not alter how the harness is invoked.Not seen on a device by hand: the OS's removal timing is read from the CI emulators' logs only.
🤖 Generated with Claude Code
https://claude.ai/code/session_01XLy7TiPAxtVjPKEwQ9iwBa
Generated by Claude Code
<source media="(prefers-color-scheme: dark)" srcset="https://pr-comments-assets.blacksmith.sh/codesmith/view-with-codesmith-dark-v3.svg"><source media="(prefers-color-scheme: light)" srcset="https://pr-comments-assets.blacksmith.sh/codesmith/view-with-codesmith-light-v3.svg"><img alt="View with [code]smith" src="https://pr-comments-assets.blacksmith.sh/codesmith/view-with-codesmith-dark-v3.svg">https://backend.blacksmith.sh/track/enable-autofix?expires=1794102246&installation_model_id=440339&pr_number=54&ref=codesmith_pr_footer&repository=vaam-apps%2Fflutter-otel-zone&return_to=https%3A%2F%2Fgithub.com%2Fvaam-apps%2Fflutter-otel-zone%2Fpull%2F54&signature=4279f46facbd3db0d0172ba4a04afb62b472036d8e21e19a9322d5cef7bf909b"><picture><source media="(prefers-color-scheme: dark)" srcset="https://pr-comments-assets.blacksmith.sh/codesmith/autofix-with-codesmith-dark.svg"><source media="(prefers-color-scheme: light)" srcset="https://pr-comments-assets.blacksmith.sh/codesmith/autofix-with-codesmith-light.svg"><img alt="Autofix with [code]smith" src="https://pr-comments-assets.blacksmith.sh/codesmith/autofix-with-codesmith-dark.svg">Need help on this PR? Tag
@codesmith-botwith what you need. Autofix is disabled.Generated by Claude Code