26 GB of that was garbage, and two disk cleaners had already looked straight at it and reported nothing to clean.
What a cleaner cannot see#
The repo is managed with yadm, so the bare repo lives outside my home directory. One command describes the damage:
$ GIT_DIR=$(yadm introspect repo) git count-objects -vH
size-garbage: 17.2 GiB
size: 8.6 GiB
Git has a precise name for the first number. garbage counts files in the
object database that are neither valid loose objects nor valid packs, and
size-garbage is the space they occupy. Mine was 17.2 GB of
objects/pack/tmp_pack_*: half-written pack files left behind by a git gc
or repack that was killed partway through. Git labels them and then leaves
them alone forever. They only ever accumulate.
The other 8.6 GB was unreachable loose objects, the residue of aborted adds and dropped stashes.
Here is the part that cost me the most time. A general-purpose disk cleaner
cannot find any of this. I ran mole’s “Large files” scan, and it reported
“Nothing to clean” while 26 GB sat in that directory. That is not a bug in
the cleaner. Every such tool judges a repository by its files, and an
objects directory full of pack files is exactly what a healthy repo looks
like from outside. Reachability is a property only git can compute, and no
cleaner is going to walk your commit graph to decide which blobs still
matter.
So the disk filled up quietly, and a full disk stops every tool on the machine at once. That is how this was found: not by noticing a large repo, but by losing the machine.
I blamed the beta OS, then a TUI#
Before I looked at the disk, I built two theories.
The first was the macOS 27 beta, which both of my machines were running. The second was herdr, the terminal multiplexer I was a few weeks into trialling at the time, because the freeze landed on the same day upstream published a version, reverted it, and removed the tag. That is a suspicious coincidence, and a coincidence is all it was.
I went looking for corroboration and found none. A search turned up zero
reports of herdr freezing an entire machine. The closest candidate, upstream
issue #2592, I read and ruled out: it describes an allocator lock convoy on
a 316-core Linux host, it happens server-side, and the machine in that
report never froze at all. Its reporter kept running perf and starting new
clients throughout.
The detail that actually settled it was one I already had. SSH into the
frozen machine was unresponsive. sshd and WindowServer are independent,
so a UI that has locked up still leaves SSH answering. A machine that
refuses SSH has stopped scheduling work, which puts the fault at the kernel,
memory pressure, saturated I/O or hardware. A terminal program cannot
produce that. The frozen machine was also the client rather than the server,
and herdr’s heavy work happens on the server side.
I want to be accurate about how much that proves. I captured no diagnosis before restarting, so all of this is reasoning after the fact rather than evidence. The check that would have ended the question takes one line, and it is now the first thing I reach for, because userspace cannot produce a kernel panic:
ls -la /Library/Logs/DiagnosticReports/*.panicTwo suspects, both plausible, both wrong, and the disk I never checked was
sitting at 228 GB capacity with 26 GB of invisible garbage on it. My triage
playbook now opens with df -h and git count-objects -vH, before anything
touches a log file at all.
The lock guard was wrong in both directions#
The obvious way to clean up tmp_pack_* safely is to skip the cleanup while
a git gc is running, and git leaves a gc.pid file behind precisely so you
can tell. I wrote that version first. Review took it apart, and the
interesting thing is that it failed in both possible directions at once.
Too permissive, because gc is not the only writer of that filename.
index-pack is the receiving side of every fetch, clone and pull, it writes
the same tmp_pack_* template, and it takes no lock at all. A maintenance
run racing a yadm pull would have deleted a pack file that was still being
written.
Too restrictive, because of how the garbage is created in the first place.
These files exist because a gc was killed, and a killed gc leaves its
gc.pid behind. The guard would therefore be switched off in exactly the
situation it was written for.
There was a third problem under the implementation. Detecting the running
process by grepping ps output produced four false matches on a machine
with no gc running at all, which is the kind of result that makes a guard
worse than no guard.
Age, not a lock#
The fix drops the idea of knowing who is writing and asks a question that needs no coordination: has anyone touched this file in the last hour?
find "$objdir/pack" -name 'tmp_pack_*' -mmin +60 -deleteAn in-flight fetch writes continuously, so its pack file is never an hour
stale. A pack left by a process that died last Tuesday always is. The
threshold sits in DM_PACK_GARBAGE_MIN_AGE_MIN so it can be raised on a
machine with slower links, and the cleanup reports what it actually removed
rather than assuming the delete worked.
Loose objects take the slower path on purpose. They go through plain
git gc, whose default expiry is two weeks, and I left that default alone:
| What | How it is cleaned | Why |
|---|---|---|
tmp_pack_* garbage | find -mmin +60 -delete | age needs no lock, and nothing legitimate writes a pack for an hour |
| unreachable loose objects | git gc, expiry two weeks | keeps the recovery window that has already saved real work here |
git gc --prune=now would have reclaimed the loose objects immediately, and
my earlier version of this check printed exactly that as the suggested
remedy. It was also the one command the policy refused to run on a schedule,
for two documented reasons: git’s own manual says --prune=now raises the
risk of corruption when another process is writing concurrently, and running
it destroys the git fsck --unreachable window that let me recover real
work from this repo once before. So the check was recommending, to a human
reading the log, the exact command it would not run itself. That line is
gone.
Lessons#
- Reachability is not a property of files, so no disk cleaner can see git
garbage. Ask git with
git count-objects -vH, and ask it about bare repos you forgot you had. - Check free space before reading a single log line. The cheap check feels less like investigating, which is why it gets skipped.
- A guard that asks “who is writing right now” needs a complete list of
writers. Mine missed
index-pack, which is the busiest one. - When the condition that creates a mess also disables your guard against it, the guard is backwards. Garbage from a killed process cannot be gated on that process looking dead.
- Age is a weaker signal than a lock and it needs no cooperation from anyone. Prefer it when the writers are not all under your control.
- If a remedy is too dangerous to run on a schedule, stop printing it as the recommended fix.
- A frozen UI and an unreachable machine are different diagnoses. SSH still answering means userspace; SSH refusing means scheduling has stopped.
References#
git count-objects(thegarbageandsize-garbagedefinitions quoted above)git gc(the two-week default expiry, and the concurrency warning on--prune=now)git index-pack(the receiving side of fetch, clone and pull)git fsck(--unreachable, the recovery net worth keeping)- yadm common commands
(
yadm introspect repo, which locates the bare repo) - The cleanup, its tests, and the triage playbook: nickboy/dotfiles
