Skip to content

Commit 618fd23

Browse files
solderzzcclaude
andauthored
test: add a coverage report that names where the suites are blind (#146)
Tier 4 of #128, which asked for `--enable-code-coverage` to find blind spots objectively rather than by eye. This runs it and aggregates the result into something readable: coverage per area, plus the largest files no test executes. Both suites are measured, because reading either alone misleads. From the umbrella package the model architectures look like 0%, which is not because they are untested but because their tests live in the submodule. Measured separately: SwiftLM package server 24.4%, inference core 57.7%, SwiftBuddy 5.3% mlx-swift-lm submodule LLM models 24.1%, VLM models 0.3%, MLXLMCommon 41.3% Two findings worth acting on. #128 named `MLXVLM/Models/Gemma4.swift` as having no tests despite being where SharpAI/mlx-swift-lm#45's changes landed. Measurement confirms it and adds that it is the single largest untested file in the model layer, at 1771 lines. More striking is the pattern around it: the VLM half of the model layer is at **0.3% over 13,304 lines** — Qwen3VL, Qwen35, LFM2VL, GlmOcr, Qwen25VL, Gemma3, Pixtral, FastVLM are all at zero. CI does run vision and omni end-to-end tests, so these paths are exercised; nothing touches them at unit level, which is why a shape or dtype mistake inside one surfaces only as a wrong image answer. Deliberately not wired into CI. A coverage gate rejects work for touching an already-thin file and invites tests written to raise a number, and the run costs a full debug rebuild of both packages. This is a tool to answer "where are we blind" when someone asks, not a check to pass. The script says so itself: coverage measures execution, not correctness — a covered line may be asserted on by nothing at all. Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
1 parent a08ffd1 commit 618fd23

1 file changed

Lines changed: 126 additions & 0 deletions

File tree

‎scripts/coverage-report.sh‎

Lines changed: 126 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,126 @@
1+
#!/usr/bin/env bash
2+
# coverage-report.sh — measure where the test suites are blind.
3+
#
4+
# Tier 4 of #128. The suites in this repository have never surfaced a defect: every
5+
# bug in the #108/#110/#112 cycle was caught by loading a real checkpoint or by code
6+
# review. Knowing *which* code they never execute turns that from an impression into
7+
# a list, and tells you whether adding tests to what exists is worth more than adding
8+
# a fixture for a shape that is missing.
9+
#
10+
# There are two suites and they cover different things, so both are measured:
11+
# - the SwiftLM package: server, inference core, the app
12+
# - the mlx-swift-lm submodule: the model architectures
13+
# Reading either alone is misleading. MLXLLM/Models looks like 0% from the umbrella
14+
# package because its tests live in the submodule, not because it is untested.
15+
#
16+
# Coverage is a map of what was executed, not evidence that behaviour is correct — a
17+
# line can be run by a test that asserts nothing. Treat a low number as a question and
18+
# a high number as no answer at all.
19+
#
20+
# Usage:
21+
# ./scripts/coverage-report.sh # both suites
22+
# ./scripts/coverage-report.sh package # SwiftLM package only
23+
# ./scripts/coverage-report.sh submodule # mlx-swift-lm only
24+
25+
set -uo pipefail
26+
27+
WHICH="${1:-both}"
28+
REPO_ROOT="$(cd "$(dirname "$0")/.." && pwd)"
29+
YELLOW='\033[1;33m'; NC='\033[0m'
30+
log() { echo -e "${YELLOW}[coverage]${NC} $*"; }
31+
32+
# Aggregates an llvm-cov JSON export by area and lists the largest untouched files.
33+
# Kept inline so the script is one file to copy and run.
34+
summarise() {
35+
python3 - "$1" "$2" <<'PY'
36+
import json, sys, collections
37+
38+
export_path, label = sys.argv[1], sys.argv[2]
39+
with open(export_path) as fh:
40+
data = json.load(fh)["data"][0]
41+
42+
def area(path):
43+
if "/MLXLLM/Models/" in path: return "model architectures (LLM)"
44+
if "/MLXVLM/Models/" in path: return "model architectures (VLM)"
45+
if "/MLXLMCommon/" in path: return "MLXLMCommon"
46+
if "/Sources/SwiftLM/" in path: return "server"
47+
if "/Sources/MLXInferenceCore/" in path: return "inference core"
48+
if "/SwiftBuddy/" in path: return "SwiftBuddy app"
49+
if "/Tests/" in path or "/tests/" in path: return "test code"
50+
if "/mlx-swift-lm/" in path: return "submodule (other)"
51+
return None # third-party dependencies; not ours to cover
52+
53+
agg = collections.defaultdict(lambda: [0, 0, 0])
54+
untouched = []
55+
for f in data["files"]:
56+
a = area(f["filename"])
57+
if a is None:
58+
continue
59+
lines = f["summary"]["lines"]
60+
agg[a][0] += lines["covered"]
61+
agg[a][1] += lines["count"]
62+
agg[a][2] += 1
63+
# Only flag files big enough that never executing them means something.
64+
if lines["percent"] == 0 and lines["count"] >= 200:
65+
untouched.append((lines["count"], f["filename"].split("/")[-1], a))
66+
67+
print(f"\n ── {label} " + "─" * max(0, 56 - len(label)))
68+
print(f" {'AREA':<28} {'COVERED':>8} {'LINES':>8} {'PCT':>7} FILES")
69+
for a, (cov, tot, n) in sorted(agg.items(), key=lambda kv: -kv[1][1]):
70+
pct = 100 * cov / tot if tot else 0.0
71+
print(f" {a:<28} {cov:>8} {tot:>8} {pct:>6.1f}% {n}")
72+
73+
if untouched:
74+
print(f"\n largest files never executed by this suite:")
75+
for count, name, a in sorted(untouched, reverse=True)[:10]:
76+
print(f" {count:>6} lines {name:<28} ({a})")
77+
PY
78+
}
79+
80+
run_suite() {
81+
local dir="$1" label="$2"
82+
log "Running $label suite with coverage (this rebuilds in debug; give it a minute)"
83+
(
84+
cd "$dir" || exit 1
85+
swift test --enable-code-coverage > /tmp/coverage-$$.log 2>&1
86+
status=$?
87+
# Test failures still leave usable coverage data, so report and continue.
88+
if [ "$status" -ne 0 ]; then
89+
echo " note: tests exited $status — coverage below reflects what ran"
90+
grep -aE "error:|failed" /tmp/coverage-$$.log | head -3
91+
fi
92+
grep -aE "Executed [0-9]+ tests" /tmp/coverage-$$.log | tail -1 | sed 's/^/ /'
93+
94+
local profdata bundle binary
95+
profdata="$(dirname "$(swift test --show-codecov-path 2>/dev/null)")/default.profdata"
96+
bundle="$(ls -d .build/*/debug/*.xctest 2>/dev/null | head -1)"
97+
if [ ! -f "$profdata" ] || [ -z "$bundle" ]; then
98+
echo " could not locate coverage artifacts for $label — skipping"
99+
exit 1
100+
fi
101+
binary="$bundle/Contents/MacOS/$(basename "$bundle" .xctest)"
102+
xcrun llvm-cov export -format=text -instr-profile "$profdata" "$binary" \
103+
> /tmp/coverage-export-$$.json 2>/dev/null
104+
summarise /tmp/coverage-export-$$.json "$label"
105+
rm -f /tmp/coverage-$$.log /tmp/coverage-export-$$.json
106+
)
107+
}
108+
109+
# The package suite aborts on its second test without this; see #128 Tier 3.
110+
if [ -x "$REPO_ROOT/scripts/install-test-metallib.sh" ]; then
111+
bash "$REPO_ROOT/scripts/install-test-metallib.sh" >/dev/null 2>&1
112+
fi
113+
114+
case "$WHICH" in
115+
package) run_suite "$REPO_ROOT" "SwiftLM package" ;;
116+
submodule) run_suite "$REPO_ROOT/mlx-swift-lm" "mlx-swift-lm submodule" ;;
117+
both)
118+
run_suite "$REPO_ROOT" "SwiftLM package"
119+
run_suite "$REPO_ROOT/mlx-swift-lm" "mlx-swift-lm submodule"
120+
;;
121+
*) echo "usage: $0 [package|submodule|both]"; exit 2 ;;
122+
esac
123+
124+
echo
125+
log "Coverage measures execution, not correctness. A covered line may still be"
126+
log "asserted on by nothing at all."

0 commit comments

Comments
 (0)