Repository navigation
Commit 539066a
bench(tabular): soft-label distillation results + lock the blob format
Full 51-dataset re-run adding gbt<-tabfm soft (TabFM's predict_proba distilled
via proba/classes). Result is modest but exactly what the mechanism predicts:
soft vs hard is 16W/11T/8L (mean +0.003) -- most datasets tie because the
teacher was already confident -- but soft gbt<-tabfm beats the knn5-taught
student on 74% (up from hard's 65%), and on 6 datasets where hard-label
distillation had lost to the cheap knn5 teacher, soft caught back up or passed
it. Soft labels stop discarding TabFM's calibrated edge; what they can't fix is
the representational half of the gap (axis-aligned trees vs a warped boundary),
which is the next student, not the next target. tabarena-full.md updated with
the soft column, head-to-head, and analysis.
ARCHITECTURE.md: point at the now-normative, versioned blob format (RFC 4.1.6,
PSTREE01/PSGBT01) -- a stored student stays servable across upgrades,
snapshots, and forks, and a serving-only module can execute a blob it never
trained.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Signed-off-by: mstrathman <matthew.strathman@gmail.com>1 parent 8b365b8 commit 539066a
3 files changed
Lines changed: 140 additions & 111 deletions
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
103 | 103 | | |
104 | 104 | | |
105 | 105 | | |
106 | | - | |
107 | | - | |
| 106 | + | |
| 107 | + | |
| 108 | + | |
| 109 | + | |
| 110 | + | |
108 | 111 | | |
109 | 112 | | |
110 | 113 | | |
| |||
0 commit comments