↓ Skip to main content
  1. Posts/

A cask upgrade left the binary hanging. A copy of it started instantly.

Nick Liu
Author
Nick Liu
Building infrastructure for Facebook Feed Ranking at Meta. Previously at Walmart, Twitter, AWS, and eBay. MS in Computer Science at Georgia Tech.
Table of Contents
Hardening My Dotfiles - This article is part of a series.
Part 7: This Article
A Homebrew cask upgrade on 2026-08-28 left my `claude` binary hanging forever on `--version`. No error, no exit, just a prompt that never came back. For two days nothing looked wrong, because every session already open kept working perfectly. The first person to notice was the first person to open a new one.

The signature was valid. The bytes were intact. A byte-identical copy of the file, sitting in the same directory, started instantly while the original hung. That left exactly one thing it could be, and it was not the file.

Two days of looking fine
#

The reason this hid for two days is the part worth keeping, and it has nothing to do with the bug itself.

When an upgrade replaces a binary, the old file is unlinked, but a process that is already running holds its inode open. The bytes stay on disk, with no name pointing at them, until the last process using them exits. So every session started before the upgrade kept running the old copy and never touched the broken one.

08-28: cask upgrade
replaces the binary

old inode is unlinked,
but still held open by
every running session

those sessions keep
working perfectly

the NEW file on disk
hangs forever on --version

nothing looks wrong
for two days

08-30: someone opens
a new session

the hang finally
becomes visible

08-28: cask upgrade
replaces the binary

old inode is unlinked,
but still held open by
every running session

those sessions keep
working perfectly

the NEW file on disk
hangs forever on --version

nothing looks wrong
for two days

08-30: someone opens
a new session

the hang finally
becomes visible

The window between breaking this and noticing it was therefore set by how long I keep sessions open, which on this machine is days.

The one comparison that ruled out everything else
#

My first instinct was that the file was damaged. It was not:

$ codesign --verify --strict /opt/homebrew/bin/claude
$ echo $?
0

That returned clean in 0.15 seconds, signed by Developer ID Anthropic PBC. Content and signature were both intact. A damaged download would have failed right there.

Then I tried to take the environment out of the picture:

$ env -i HOME=/tmp/probe PATH=/usr/bin:/bin /opt/homebrew/bin/claude --version
# still hangs

Empty environment, fresh HOME, no config to read. Same hang. I also cleared com.apple.quarantine, which changed nothing, and found that com.apple.provenance cannot be removed at all.

The experiment that actually decided it was much simpler. I copied the file next to itself and ran the copy:

$ cp /opt/homebrew/bin/claude /opt/homebrew/bin/claude-copy
$ /opt/homebrew/bin/claude-copy --version
# instant

One command, and most of my remaining theories died together. The copy sat in the same directory as the original, held the same bytes, ran under the same environment, and lived on the same filesystem. Path, content, config and disk were all ruled out in a single comparison. What is left, when a file and a copy of that file in the same folder behave differently, is the inode.

The fix follows from that. Copy over the original, which gives it a fresh inode holding the same bytes:

$ cp /opt/homebrew/bin/claude-copy /opt/homebrew/bin/claude
$ codesign --verify --strict /opt/homebrew/bin/claude && echo signature ok
signature ok

Why the old inode hung, I still cannot tell you. The upgrade looks like it was interrupted partway through, and this was a 302 MB cask binary, so there was plenty of time to interrupt. I can describe the shape of the fault and the test that identifies it. The mechanism underneath stays unexplained, and I would rather write that down than invent something that sounds complete.

Two traps that gave me confident wrong answers
#

Before the copy experiment, I wasted real time on two things that looked like evidence.

sample showed only _dyld_start, and I believed it. A truncated stack with nothing but the loader entry point looks exactly like “stuck in the dynamic linker”. It is not evidence of that. claude is a single-file bundled binary, and bundlers embed a runtime whose frames do not symbolicate, so the stack is cut short rather than genuinely short. I had a theory about dyld before I had a result that supported one.

My first three probes never ran claude at all. This one is funnier and more useful. I wrapped the call in a timeout, like this:

# reports rc=0 and never executes claude once
timeout 25 command claude --version

command is a shell builtin. timeout is a program, and a program that execs cannot see builtins, shell functions, or aliases. So timeout looked for a binary named command, and the whole thing reported success without the probed program ever starting. Three runs, three clean exit codes, zero actual measurements.

That one is not even a new lesson here. My own rules already say a builtin can intercept something you took for a binary, which is how log show once reported zero errors from a log holding tens of thousands. This is the same confusion pointing the other way: anything that execs, so timeout, xargs, nohup, sudo, cannot see what the shell resolves for you. Wrap a shell-resolved name in one of those and you quietly probe something else.

The check could not live where the other checks live
#

The obvious place for a new guard was the block that already asks each tool to parse its config and report back. I could not put it there, and the reason is the whole design.

Those checks assume the tool exits. This failure is a tool that does not exit. A check written in that style would have inherited the hang and taken the entire maintenance run down with it. A check that can hang the thing it is checking is worse than no check, so this one refuses to run unbounded.

dm_launch_probe() {   # $1 = seconds, rest = command and args
    local secs="$1" rc tmo
    shift
    tmo="${TIMEOUT_CMD:-}"
    [ -n "$tmo" ] || tmo=$(command -v timeout 2>/dev/null || command -v gtimeout 2>/dev/null)
    command -v "$tmo" >/dev/null 2>&1 || { echo skipped; return 0; }
    "$tmo" "$secs" "$@" >/dev/null 2>&1
    rc=$?
    case "$rc" in
        0)   echo ok ;;
        124) echo hang ;;
        *)   echo "fail $rc" ;;
    esac
}

Four verdicts, because two would mislead
#

VerdictWhat it meansWhere the fix is
okthe binary started and reached its own codenothing to do
fail <rc>the tool is rejecting something it can seeits config or arguments
hangit never reached its own codethe filesystem, so give it a new inode
skippedthere is no usable timeout commandthe probe itself, not the probed tool

The split between hang and fail is the point. A non-zero exit means the program started, read something, and said no. A timeout means it never got as far as its own code, so nothing it can see is relevant and its config is the wrong place to look. My repo has already paid once for printing one remedy for two different causes, so these two print different remedies, and the hang branch tells me to re-verify the signature after replacing the inode.

The fourth verdict earns its place too. GNU timeout reserves 124 for a real timeout and 127 for “command cannot be found”, so a TIMEOUT_CMD pointing at something missing would have reported fail 127. That reads as a verdict about claude when it is really a verdict about my own probe. Misattribution like that is the exact thing this function exists to prevent, so a broken probe says skipped and stays honest.

All four verdicts were produced on purpose before I trusted them, and two were proven by breaking them: folding 124 into the fail branch turns the hang test red, and dropping the command -v validation turns the skipped test red.

run_test "launch probe: a clean exit is ok" \
    "[ \"\$(dm_launch_probe 5 /usr/bin/true)\" = ok ]"
run_test "launch probe: a non-zero exit is fail, with the code" \
    "[ \"\$(dm_launch_probe 5 /usr/bin/false)\" = 'fail 1' ]"
run_test "launch probe: a hang is hang, not fail" \
    "[ \"\$(dm_launch_probe 2 /bin/sleep 30)\" = hang ]"
run_test "launch probe: a broken timeout command skips, not fails" \
    "[ \"\$(TIMEOUT_CMD=/nonexistent-timeout dm_launch_probe 2 /usr/bin/true)\" = skipped ]"

CI caught what my machine could not
#

My Mac has coreutils installed, so timeout is always there and all four tests pass locally. A GitHub runner does not: macOS ships no timeout at all. On CI the probe correctly returned skipped for every case, which turned three of my four assertions red.

The fourth one passed.

That asymmetry is the tell, and it is the same shape as the inverted test in the previous post of this series. Three tests that assert a specific verdict go red the moment every verdict becomes skipped. The survivor is the one comparing skipped against skipped, so it had never been able to tell any two outcomes apart in the first place.

The guard went on the machine rather than on the repo, because having a timeout command is a property of the machine. It also announces itself rather than disappearing:

skipped (no timeout command)

Four tests vanished behind a rename in this repo earlier the same month and the suite stayed green, so a block that silently stops running is a shape I now actively look for. A suite has three outcomes, and the third one has to be visible.

The number I got wrong, and the claim I narrowed
#

Two corrections came out of review, and both made the thing smaller.

I set the timeout at 20 seconds and justified it badly. I wrote that a real hang would cost 20 seconds of every maintenance run forever, which sounds expensive until you notice runs are daily. It costs 20 seconds once a day, which is nothing. The failure that actually costs something is the opposite one: a false hang reported on a working tool teaches me to ignore the check, and then the check is decoration. That argument points at a generous cap rather than a tight one. The measured warm start is 0.04 to 0.05 seconds over three runs, so 20 seconds is roughly 400 times the warm path. It stays at 20, and if it ever moves it moves longer.

The honest part: that reasoning covers the warm path, which I measured. The cold path, with Gatekeeper assessment and a cold loader cache right after an upgrade, I have measured exactly zero times. The cap is a guess about a number I do not have. What replaces the guess is already scheduled, because the next cask upgrade is the cold path, so I time it then.

The second correction was to what the check claims. My first version printed “claude launches OK”. It does not know that. It knows one narrow thing:

claude launch probe (--version) started

--version proves the binary started and reached its own code. It proves nothing about whether a real session works. A narrower claim is an accurate one, and this particular narrow claim happens to sit directly on top of the failure it guards: the break was at the loader and inode layer, and --version exercises that same layer. The check is not standing in for the fault. It sits on the fault’s own ground.

Lessons
#

  • A file and a copy of that file in the same directory differ in exactly one way. When a binary misbehaves and its copy does not, you have already found the layer.
  • An upgrade that only breaks new processes has a delay built in, as long as your longest-lived session. Probe on a schedule, not on symptoms.
  • A hang and a non-zero exit need separate remedies. One means the tool read something and refused; the other means it never reached its own code.
  • A verdict about your own instrument must never be printed as a verdict about the thing it measures.
  • Programs that exec cannot see builtins, shell functions, or aliases. Wrap a shell-resolved name in timeout and you measure something else entirely.
  • When one assertion survives a change that reddens its neighbours, suspect the survivor. It is usually the one that cannot distinguish anything.
  • Write down which half of your reasoning rests on a measurement and which half is a guess, along with the event that will replace the guess.

References
#

Hardening My Dotfiles - This article is part of a series.
Part 7: This Article

Related

My 2 MB dotfiles repo grew to 31 GB. The cleaner found nothing to clean.

··1498 words·8 mins
My Mac mini froze hard enough that I held the power button to get it back. My first suspect was the macOS beta. My second was a terminal multiplexer I had been trialling. Both were wrong, and I had skipped the one check that would have pointed at the answer in about four seconds: the disk was full. My dotfiles repo tracks 2 MB of content, and it had grown to 31 GB. 26 GB of that was garbage, and two disk cleaners had already looked straight at it and reported nothing to clean.

Five wrong answers in one day. One nearly deleted 10 GB of coursework.

··1903 words·9 mins
`mdls kMDItemLastUsedDate` returned `(null)` for Microsoft Word. I read the null as "never opened" and put Office on a removal list: 10.1 GB, four apps. One last check saved me. My home directory held 150 Office documents, a conference presentation edited two weeks earlier, and a PowerPoint lock file, which only exists while the file is open. The proof that the null was misleading had been sitting in my own diagnostic report for an hour. That was one of five. In a single day of hardening this machine, five different tools told me things that were not true. None of the answers looked like an error. Each one arrived as a clean, confident finding, and under each one a check had quietly failed or asked the wrong question. All five had the same shape underneath. Once I could name the shape, I stopped falling for it.

My dotfiles had a no-exceptions test gate. It had never run once.

··1331 words·7 mins
My dotfiles repo has a CLAUDE.md, and the CLAUDE.md has a rule in bold: every commit must pass the test suite, no exceptions. Within the first hour of an audit this July, I learned that this rule had been enforced exactly zero times since the day it was written. The hook file existed, its contents were correct, it even had its executable bit. It was just sitting at a path that yadm stopped reading a major version ago. No error message. No warning. To yadm, a hook in the wrong place and no hook at all are the same thing.