Skip to content

Draft: concretize AI-detector-flagged sections (ai-first + one-on-one) - #1954

Merged
benbalter merged 3 commits into
mainfrom
post/detector-concreteness-draft
Aug 18, 2026
Merged

Draft: concretize AI-detector-flagged sections (ai-first + one-on-one)#1954
benbalter merged 3 commits into
mainfrom
post/detector-concreteness-draft

Conversation

@benbalter

Copy link
Copy Markdown
Owner

Draft for review — do not merge as-is. Ran Fast-DetectGPT (Llama-3-8B pair) over the 2026 posts; this branch reworks the passages it flagged as most machine-like in two posts.

The core finding

The detector separates concrete/specific prose (reads human) from abstract/summarizing prose (reads machine) — not by structure. Bold-label lists score anywhere from 20% to 97% depending on whether their items are named scenarios or generic categories. So the lever is the same as Ben's own voice principle: replace abstractions with a number, a named tool, a lived moment. List structure is kept intact.

Section scores (Llama-3, before → after)

post / section before after
one-on-one · "What belongs in a 1:1" 93% 30%
ai-first · "What changes for PMs" 97% 20%
ai-first · "What doesn't change" 87% 46%
ai-first · orchestra metaphor → "The judgment is still yours" 93% 60%
ai-first · "From async" 77% 68% (resists — abstract throughline)

Important caveat

Whole-post scores barely move (and can rise when an anecdote adds length) because the Fast-DetectGPT criterion grows ~√N. The whole-post % mostly measures length, not AI-ness — it is not a meaningful optimization target. The section scores are the actionable signal, and every edited section improved. These edits stand on their own as voice improvements regardless of the detector.

Notes / open decisions

  • The launch anecdote in "The judgment is still yours" uses illustrative specifics (SnippetGPT, "lgtm, mostly", auth edge case) — please swap in the real story if the details differ.
  • agentic-workflows was excluded (its 100% is the intentional "agentic" gimmick).

🤖 Generated with Claude Code

benbalter and others added 3 commits August 17, 2026 23:18
Re-analysis (same instrument, run on the Open & Async book with finer per-era
buckets) found the kept-ON top signal, AnthropomorphicJustification, conflates
human/collective subjects ("the people doing the work", "deserve credit") with
the real inanimate tell. On the book, splitting by subject cut it from 6.7x
whole-rule to ~2.4x inanimate-only. Adds a note to re-measure the 16x on the
inanimate subset before trusting it, plus per-era corroboration that
FigurativeFalls/Idioms are revived early-2010s habits (0.11/1k in 2010-13).

Docs-only: comments in the advisory config; no rule behavior changes.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Fast-DetectGPT (Llama-3-8B) flagged abstract/summarizing passages as
machine-like. Ground them in specific, named scenarios (keeping the list
structure, per the concreteness-not-structure finding):

- one-on-one "What belongs in a 1:1": 93% -> 30%
- ai-first "What changes for PMs": 97% -> 20%
- ai-first orchestra metaphor -> concrete judgment examples: 93% -> 71%

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Second pass on the ai-first post:
- 'The judgment is still yours': replace abstract triplet with a concrete
  launch anecdote (green board, 'lgtm, mostly', held a week): 71% -> 60%
- 'What doesn't change': concretize + drop 'The opposite is true'/'not less': 87% -> 46%
- 'From async': drop two antithesis reversals: 77% -> 68%

Note: section scores all down, but whole-post rose 83->88% because the
anecdote adds length and the criterion grows ~sqrt(N). Whole-post % is a
length artifact, not a meaningful target; section scores are the signal.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@benbalter
benbalter merged commit 6166f54 into main Aug 18, 2026
12 checks passed
@benbalter
benbalter deleted the post/detector-concreteness-draft branch August 18, 2026 17:11
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant