↓ Skip to main content
  1. Posts/

My 2 MB dotfiles repo grew to 31 GB. The cleaner found nothing to clean.

··1498 words·8 mins·
Nick Liu
Author
Nick Liu
Building infrastructure for Facebook Feed Ranking at Meta. Previously at Walmart, Twitter, AWS, and eBay. MS in Computer Science at Georgia Tech.
Table of Contents
Hardening My Dotfiles - This article is part of a series.
Part 8: This Article
My Mac mini froze hard enough that I held the power button to get it back. My first suspect was the macOS beta. My second was a terminal multiplexer I had been trialling. Both were wrong, and I had skipped the one check that would have pointed at the answer in about four seconds: the disk was full. My dotfiles repo tracks 2 MB of content, and it had grown to 31 GB.

26 GB of that was garbage, and two disk cleaners had already looked straight at it and reported nothing to clean.

What a cleaner cannot see
#

The repo is managed with yadm, so the bare repo lives outside my home directory. One command describes the damage:

$ GIT_DIR=$(yadm introspect repo) git count-objects -vH
size-garbage: 17.2 GiB
size:          8.6 GiB

Git has a precise name for the first number. garbage counts files in the object database that are neither valid loose objects nor valid packs, and size-garbage is the space they occupy. Mine was 17.2 GB of objects/pack/tmp_pack_*: half-written pack files left behind by a git gc or repack that was killed partway through. Git labels them and then leaves them alone forever. They only ever accumulate.

The other 8.6 GB was unreachable loose objects, the residue of aborted adds and dropped stashes.

Here is the part that cost me the most time. A general-purpose disk cleaner cannot find any of this. I ran mole’s “Large files” scan, and it reported “Nothing to clean” while 26 GB sat in that directory. That is not a bug in the cleaner. Every such tool judges a repository by its files, and an objects directory full of pack files is exactly what a healthy repo looks like from outside. Reachability is a property only git can compute, and no cleaner is going to walk your commit graph to decide which blobs still matter.

So the disk filled up quietly, and a full disk stops every tool on the machine at once. That is how this was found: not by noticing a large repo, but by losing the machine.

I blamed the beta OS, then a TUI
#

Before I looked at the disk, I built two theories.

The first was the macOS 27 beta, which both of my machines were running. The second was herdr, the terminal multiplexer I was a few weeks into trialling at the time, because the freeze landed on the same day upstream published a version, reverted it, and removed the tag. That is a suspicious coincidence, and a coincidence is all it was.

I went looking for corroboration and found none. A search turned up zero reports of herdr freezing an entire machine. The closest candidate, upstream issue #2592, I read and ruled out: it describes an allocator lock convoy on a 316-core Linux host, it happens server-side, and the machine in that report never froze at all. Its reporter kept running perf and starting new clients throughout.

The detail that actually settled it was one I already had. SSH into the frozen machine was unresponsive. sshd and WindowServer are independent, so a UI that has locked up still leaves SSH answering. A machine that refuses SSH has stopped scheduling work, which puts the fault at the kernel, memory pressure, saturated I/O or hardware. A terminal program cannot produce that. The frozen machine was also the client rather than the server, and herdr’s heavy work happens on the server side.

I want to be accurate about how much that proves. I captured no diagnosis before restarting, so all of this is reasoning after the fact rather than evidence. The check that would have ended the question takes one line, and it is now the first thing I reach for, because userspace cannot produce a kernel panic:

ls -la /Library/Logs/DiagnosticReports/*.panic

Two suspects, both plausible, both wrong, and the disk I never checked was sitting at 228 GB capacity with 26 GB of invisible garbage on it. My triage playbook now opens with df -h and git count-objects -vH, before anything touches a log file at all.

The lock guard was wrong in both directions
#

The obvious way to clean up tmp_pack_* safely is to skip the cleanup while a git gc is running, and git leaves a gc.pid file behind precisely so you can tell. I wrote that version first. Review took it apart, and the interesting thing is that it failed in both possible directions at once.

guard: skip cleanup
while a gc is running

direction 1:
too permissive

direction 2:
too restrictive

index-pack writes the same
tmp_pack_* on every fetch
and pull, and takes NO lock

cleanup deletes a pack
a live fetch is still writing

a KILLED gc leaves
gc.pid behind

guard stays off in exactly
the case that creates
the garbage

guard: skip cleanup
while a gc is running

direction 1:
too permissive

direction 2:
too restrictive

index-pack writes the same
tmp_pack_* on every fetch
and pull, and takes NO lock

cleanup deletes a pack
a live fetch is still writing

a KILLED gc leaves
gc.pid behind

guard stays off in exactly
the case that creates
the garbage

Too permissive, because gc is not the only writer of that filename. index-pack is the receiving side of every fetch, clone and pull, it writes the same tmp_pack_* template, and it takes no lock at all. A maintenance run racing a yadm pull would have deleted a pack file that was still being written.

Too restrictive, because of how the garbage is created in the first place. These files exist because a gc was killed, and a killed gc leaves its gc.pid behind. The guard would therefore be switched off in exactly the situation it was written for.

There was a third problem under the implementation. Detecting the running process by grepping ps output produced four false matches on a machine with no gc running at all, which is the kind of result that makes a guard worse than no guard.

Age, not a lock
#

The fix drops the idea of knowing who is writing and asks a question that needs no coordination: has anyone touched this file in the last hour?

find "$objdir/pack" -name 'tmp_pack_*' -mmin +60 -delete

An in-flight fetch writes continuously, so its pack file is never an hour stale. A pack left by a process that died last Tuesday always is. The threshold sits in DM_PACK_GARBAGE_MIN_AGE_MIN so it can be raised on a machine with slower links, and the cleanup reports what it actually removed rather than assuming the delete worked.

Loose objects take the slower path on purpose. They go through plain git gc, whose default expiry is two weeks, and I left that default alone:

WhatHow it is cleanedWhy
tmp_pack_* garbagefind -mmin +60 -deleteage needs no lock, and nothing legitimate writes a pack for an hour
unreachable loose objectsgit gc, expiry two weekskeeps the recovery window that has already saved real work here

git gc --prune=now would have reclaimed the loose objects immediately, and my earlier version of this check printed exactly that as the suggested remedy. It was also the one command the policy refused to run on a schedule, for two documented reasons: git’s own manual says --prune=now raises the risk of corruption when another process is writing concurrently, and running it destroys the git fsck --unreachable window that let me recover real work from this repo once before. So the check was recommending, to a human reading the log, the exact command it would not run itself. That line is gone.

Lessons
#

  • Reachability is not a property of files, so no disk cleaner can see git garbage. Ask git with git count-objects -vH, and ask it about bare repos you forgot you had.
  • Check free space before reading a single log line. The cheap check feels less like investigating, which is why it gets skipped.
  • A guard that asks “who is writing right now” needs a complete list of writers. Mine missed index-pack, which is the busiest one.
  • When the condition that creates a mess also disables your guard against it, the guard is backwards. Garbage from a killed process cannot be gated on that process looking dead.
  • Age is a weaker signal than a lock and it needs no cooperation from anyone. Prefer it when the writers are not all under your control.
  • If a remedy is too dangerous to run on a schedule, stop printing it as the recommended fix.
  • A frozen UI and an unreachable machine are different diagnoses. SSH still answering means userspace; SSH refusing means scheduling has stopped.

References
#

  • git count-objects (the garbage and size-garbage definitions quoted above)
  • git gc (the two-week default expiry, and the concurrency warning on --prune=now)
  • git index-pack (the receiving side of fetch, clone and pull)
  • git fsck (--unreachable, the recovery net worth keeping)
  • yadm common commands (yadm introspect repo, which locates the bare repo)
  • The cleanup, its tests, and the triage playbook: nickboy/dotfiles
Hardening My Dotfiles - This article is part of a series.
Part 8: This Article

Related

A cask upgrade left the binary hanging. A copy of it started instantly.

A Homebrew cask upgrade on 2026-08-28 left my `claude` binary hanging forever on `--version`. No error, no exit, just a prompt that never came back. For two days nothing looked wrong, because every session already open kept working perfectly. The first person to notice was the first person to open a new one. The signature was valid. The bytes were intact. A byte-identical copy of the file, sitting in the same directory, started instantly while the original hung. That left exactly one thing it could be, and it was not the file.

Five wrong answers in one day. One nearly deleted 10 GB of coursework.

··1903 words·9 mins
`mdls kMDItemLastUsedDate` returned `(null)` for Microsoft Word. I read the null as "never opened" and put Office on a removal list: 10.1 GB, four apps. One last check saved me. My home directory held 150 Office documents, a conference presentation edited two weeks earlier, and a PowerPoint lock file, which only exists while the file is open. The proof that the null was misleading had been sitting in my own diagnostic report for an hour. That was one of five. In a single day of hardening this machine, five different tools told me things that were not true. None of the answers looked like an error. Each one arrived as a clean, confident finding, and under each one a check had quietly failed or asked the wrong question. All five had the same shape underneath. Once I could name the shape, I stopped falling for it.

My commands vanished with exit 0. The culprit was a file named env.

··1441 words·7 mins
`env -u VAR command` did nothing. Exit code 0, no output, no error. A different command, same thing. A different variable, same thing. Any invocation that started with `env` just quietly evaporated. What finally made the problem visible was a `git init` that reported success while creating no `.git` directory at all. That was a year ago. The case closed last month, and the culprit was not uv, not some third-party installer, not anything exotic. It was this repo’s own bootstrap script. The thing that caught it was the regression test I had written for the original incident.