The signature was valid. The bytes were intact. A byte-identical copy of the file, sitting in the same directory, started instantly while the original hung. That left exactly one thing it could be, and it was not the file.
Two days of looking fine#
The reason this hid for two days is the part worth keeping, and it has nothing to do with the bug itself.
When an upgrade replaces a binary, the old file is unlinked, but a process that is already running holds its inode open. The bytes stay on disk, with no name pointing at them, until the last process using them exits. So every session started before the upgrade kept running the old copy and never touched the broken one.
The window between breaking this and noticing it was therefore set by how long I keep sessions open, which on this machine is days.
The one comparison that ruled out everything else#
My first instinct was that the file was damaged. It was not:
$ codesign --verify --strict /opt/homebrew/bin/claude
$ echo $?
0
That returned clean in 0.15 seconds, signed by Developer ID Anthropic PBC. Content and signature were both intact. A damaged download would have failed right there.
Then I tried to take the environment out of the picture:
$ env -i HOME=/tmp/probe PATH=/usr/bin:/bin /opt/homebrew/bin/claude --version
# still hangs
Empty environment, fresh HOME, no config to read. Same hang. I also cleared
com.apple.quarantine, which changed nothing, and found that
com.apple.provenance cannot be removed at all.
The experiment that actually decided it was much simpler. I copied the file next to itself and ran the copy:
$ cp /opt/homebrew/bin/claude /opt/homebrew/bin/claude-copy
$ /opt/homebrew/bin/claude-copy --version
# instant
One command, and most of my remaining theories died together. The copy sat in the same directory as the original, held the same bytes, ran under the same environment, and lived on the same filesystem. Path, content, config and disk were all ruled out in a single comparison. What is left, when a file and a copy of that file in the same folder behave differently, is the inode.
The fix follows from that. Copy over the original, which gives it a fresh inode holding the same bytes:
$ cp /opt/homebrew/bin/claude-copy /opt/homebrew/bin/claude
$ codesign --verify --strict /opt/homebrew/bin/claude && echo signature ok
signature ok
Why the old inode hung, I still cannot tell you. The upgrade looks like it was interrupted partway through, and this was a 302 MB cask binary, so there was plenty of time to interrupt. I can describe the shape of the fault and the test that identifies it. The mechanism underneath stays unexplained, and I would rather write that down than invent something that sounds complete.
Two traps that gave me confident wrong answers#
Before the copy experiment, I wasted real time on two things that looked like evidence.
sample showed only _dyld_start, and I believed it. A truncated stack
with nothing but the loader entry point looks exactly like “stuck in the
dynamic linker”. It is not evidence of that. claude is a single-file
bundled binary, and bundlers embed a runtime whose frames do not
symbolicate, so the stack is cut short rather than genuinely short. I had a
theory about dyld before I had a result that supported one.
My first three probes never ran claude at all. This one is funnier and
more useful. I wrapped the call in a timeout, like this:
# reports rc=0 and never executes claude once
timeout 25 command claude --versioncommand is a shell builtin. timeout is a program, and a program that
execs cannot see builtins, shell functions, or aliases. So timeout looked
for a binary named command, and the whole thing reported success without
the probed program ever starting. Three runs, three clean exit codes, zero
actual measurements.
That one is not even a new lesson here. My own rules already say a builtin
can intercept something you took for a binary, which is how log show once
reported zero errors from a log holding tens of thousands. This is the same
confusion pointing the other way: anything that execs, so timeout,
xargs, nohup, sudo, cannot see what the shell resolves for you. Wrap a
shell-resolved name in one of those and you quietly probe something else.
The check could not live where the other checks live#
The obvious place for a new guard was the block that already asks each tool to parse its config and report back. I could not put it there, and the reason is the whole design.
Those checks assume the tool exits. This failure is a tool that does not exit. A check written in that style would have inherited the hang and taken the entire maintenance run down with it. A check that can hang the thing it is checking is worse than no check, so this one refuses to run unbounded.
dm_launch_probe() { # $1 = seconds, rest = command and args
local secs="$1" rc tmo
shift
tmo="${TIMEOUT_CMD:-}"
[ -n "$tmo" ] || tmo=$(command -v timeout 2>/dev/null || command -v gtimeout 2>/dev/null)
command -v "$tmo" >/dev/null 2>&1 || { echo skipped; return 0; }
"$tmo" "$secs" "$@" >/dev/null 2>&1
rc=$?
case "$rc" in
0) echo ok ;;
124) echo hang ;;
*) echo "fail $rc" ;;
esac
}Four verdicts, because two would mislead#
| Verdict | What it means | Where the fix is |
|---|---|---|
ok | the binary started and reached its own code | nothing to do |
fail <rc> | the tool is rejecting something it can see | its config or arguments |
hang | it never reached its own code | the filesystem, so give it a new inode |
skipped | there is no usable timeout command | the probe itself, not the probed tool |
The split between hang and fail is the point. A non-zero exit means the
program started, read something, and said no. A timeout means it never got
as far as its own code, so nothing it can see is relevant and its config is
the wrong place to look. My repo has already paid once for printing one
remedy for two different causes, so these two print different remedies, and
the hang branch tells me to re-verify the signature after replacing the
inode.
The fourth verdict earns its place too. GNU timeout reserves 124 for
a real timeout and 127 for “command cannot be found”, so a TIMEOUT_CMD
pointing at something missing would have reported fail 127. That reads as
a verdict about claude when it is really a verdict about my own probe.
Misattribution like that is the exact thing this function exists to prevent,
so a broken probe says skipped and stays honest.
All four verdicts were produced on purpose before I trusted them, and two
were proven by breaking them: folding 124 into the fail branch turns the
hang test red, and dropping the command -v validation turns the skipped
test red.
run_test "launch probe: a clean exit is ok" \
"[ \"\$(dm_launch_probe 5 /usr/bin/true)\" = ok ]"
run_test "launch probe: a non-zero exit is fail, with the code" \
"[ \"\$(dm_launch_probe 5 /usr/bin/false)\" = 'fail 1' ]"
run_test "launch probe: a hang is hang, not fail" \
"[ \"\$(dm_launch_probe 2 /bin/sleep 30)\" = hang ]"
run_test "launch probe: a broken timeout command skips, not fails" \
"[ \"\$(TIMEOUT_CMD=/nonexistent-timeout dm_launch_probe 2 /usr/bin/true)\" = skipped ]"CI caught what my machine could not#
My Mac has coreutils installed, so timeout is always there and all four
tests pass locally. A GitHub runner does not: macOS ships no timeout at
all. On CI the probe correctly returned skipped for every case, which
turned three of my four assertions red.
The fourth one passed.
That asymmetry is the tell, and it is the same shape as the inverted test in
the previous post of this series. Three tests that assert a specific verdict
go red the moment every verdict becomes skipped. The survivor is the one
comparing skipped against skipped, so it had never been able to tell any
two outcomes apart in the first place.
The guard went on the machine rather than on the repo, because having a timeout command is a property of the machine. It also announces itself rather than disappearing:
skipped (no timeout command)Four tests vanished behind a rename in this repo earlier the same month and the suite stayed green, so a block that silently stops running is a shape I now actively look for. A suite has three outcomes, and the third one has to be visible.
The number I got wrong, and the claim I narrowed#
Two corrections came out of review, and both made the thing smaller.
I set the timeout at 20 seconds and justified it badly. I wrote that a real hang would cost 20 seconds of every maintenance run forever, which sounds expensive until you notice runs are daily. It costs 20 seconds once a day, which is nothing. The failure that actually costs something is the opposite one: a false hang reported on a working tool teaches me to ignore the check, and then the check is decoration. That argument points at a generous cap rather than a tight one. The measured warm start is 0.04 to 0.05 seconds over three runs, so 20 seconds is roughly 400 times the warm path. It stays at 20, and if it ever moves it moves longer.
The honest part: that reasoning covers the warm path, which I measured. The cold path, with Gatekeeper assessment and a cold loader cache right after an upgrade, I have measured exactly zero times. The cap is a guess about a number I do not have. What replaces the guess is already scheduled, because the next cask upgrade is the cold path, so I time it then.
The second correction was to what the check claims. My first version printed “claude launches OK”. It does not know that. It knows one narrow thing:
claude launch probe (--version) started--version proves the binary started and reached its own code. It proves
nothing about whether a real session works. A narrower claim is an accurate
one, and this particular narrow claim happens to sit directly on top of the
failure it guards: the break was at the loader and inode layer, and
--version exercises that same layer. The check is not standing in for the
fault. It sits on the fault’s own ground.
Lessons#
- A file and a copy of that file in the same directory differ in exactly one way. When a binary misbehaves and its copy does not, you have already found the layer.
- An upgrade that only breaks new processes has a delay built in, as long as your longest-lived session. Probe on a schedule, not on symptoms.
- A hang and a non-zero exit need separate remedies. One means the tool read something and refused; the other means it never reached its own code.
- A verdict about your own instrument must never be printed as a verdict about the thing it measures.
- Programs that exec cannot see builtins, shell functions, or aliases. Wrap
a shell-resolved name in
timeoutand you measure something else entirely. - When one assertion survives a change that reddens its neighbours, suspect the survivor. It is usually the one that cannot distinguish anything.
- Write down which half of your reasoning rests on a measurement and which half is a guess, along with the event that will replace the guess.
References#
- GNU coreutils:
timeoutinvocation (exit status 124 for a timeout, 127 when the command cannot be found) - Bun: single-file executables
(one example of the bundled-runtime shape that produces a truncated
samplestack) - Homebrew cask: claude-code
codesign(1)andxattr(1)on macOS, plus Apple’s notarization documentation- The probe, its four tests, and the triage playbook: nickboy/dotfiles
