My machine started hanging a few days ago. Applications crashed, the desktop stuttered, and nothing I did made it better. The numbers I reached for first all looked healthy, which is exactly why it took me a while to find the real cause.
Here is what I saw:
- 476 GB partition, roughly 20% still free
- CPU around 2%
- RAM around 29%
- no runaway process, no memory pressure, no disk that was actually full
Nothing about that list explains a machine that cannot keep a window open. So I went looking at the filesystem instead of the load average.
The number that explained it
Btrfs splits a device into a data block group and a metadata block group, and the
two fill independently. df only ever shows you the data side. The metadata side
is where every inode, every directory entry, and every extent record lives, and
when it fills up, allocation starts failing and the filesystem gets slow even
though you have a third of your disk free.
btrfs filesystem usage -T /home
Data Metadata System
Id Path single DUP DUP Unallocated Total Slack
-- -------------- --------- -------- -------- ----------- --------- -------
1 /dev/nvme0n1p2 - - - 476.44GiB 476.44GiB 3.00KiB
-- -------------- --------- -------- -------- ----------- --------- -------
Total 452.42GiB 12.00GiB 8.00MiB 476.44GiB 476.44GiB 3.00KiB
Used 391.47GiB 11.17GiB 80.00KiB
Data ratio is 1.00 and metadata ratio is 2.00, so the 12.00 GiB "Total" line
is two mirrored copies of roughly 6 GiB each. The Used figure of 11.17 GiB is
also both copies, which puts the real metadata space at 93% full.
That is the whole story of the hang. 60 GiB of free data space sitting there while the part of the filesystem that has to be updated on every single file operation had almost no room left.
I want to correct one number from my own issue report before going further. I originally attributed about 8 GB of that metadata to one application. That does not survive arithmetic. 125,000 files at roughly 2 KB of metadata each is about 250 MB. What the application really had was 33 GB of data and 125,000 files. The symptom was real, the attribution was wrong.
Finding the writer
The next step was to find what was creating files faster than anything else on the machine.
du -sh ~/.local/share/opencode/*
find ~/.local/share/opencode/snapshot -type f | wc -l
34G /home/swadhin/.local/share/opencode/snapshot
365M .../opencode.db
38M .../log/opencode.log
16M .../tool-output
12M .../shell
125829
125,829 files, 3,107 directories, 33 GB. Everything else in that directory was rounding error. Thirteen session directories, and every one of them was a Git repository.
What the snapshot layer actually does
OpenCode keeps a hidden Git repository so it can undo file changes inside a session. The docs are upfront that it is a Git repository and that it can be a problem on large trees:
For large repositories or projects with many submodules, the snapshot system can cause slow indexing and significant disk usage as it tracks all changes using an internal git repository.
The docs also give the switch. In opencode.json:
{
"$schema": "https://opencode.ai/config.json",
"snapshot": false
}
I did not have that set, on OpenCode v2.0.18.
The interesting part is what the repository looked like when I got to it.
cd ~/.local/share/opencode/snapshot/global/4fc815f20b5b347c8529f0d3cd0a0b48750fb6ac
cat HEAD
git count-objects -vH
git log --oneline | wc -l
ref: refs/heads/master
count: 98935
size: 32.96 GiB
in-pack: 0
packs: 0
0
HEAD points at refs/heads/master and refs/ is empty. There are 98,935 loose
objects and not a single commit.
The config file in that same directory names the worktree it was built from:
[core]
repositoryformatversion = 0
filemode = true
bare = false
logallrefupdates = true
worktree = /home/swadhin/BDFASHION
So for that one project, OpenCode hashed essentially the entire tree into loose Git objects, wrote a 7.8 MB index, and then never wrote a commit. Twelve of the thirteen session repos were in the same state.
Why zero commits is the bad case
A Git repository with commits has a garbage collector. Objects get packed, unreferenced ones get pruned, and a repo that grew too big eventually shrinks. A repository with 98,935 loose objects and no refs has none of that. Every object is unreferenced, and unreferenced loose objects are never cleaned up because there is nothing that would ever want them gone.
The size distribution explains where the 33 GB went:
| Object size | Count |
|---|---|
| under 1 KB | 5,977 |
| 1 KB to 64 KB | 11,596 |
| 64 KB to 1 MB | 73,439 |
| 1 MB and up | 7,924 |
Those 7,924 files over 1 MB are the disk cost. The other 90,000 are the metadata cost, and they are the part that hurts on Btrfs.
The mechanism is straightforward:
graph TD
A[Session starts] --> B[OpenCode creates a shadow git repo]
B --> C[git hash-object per file in the tree]
C --> D[objects/xx/yyyy loose object files]
D --> E[No commit, no refs]
E --> F[Nothing prunes them]
Each of those objects is a real file on disk. On ext4 or XFS that is unpleasant but survivable. On Btrfs, where the metadata pool is shared and finite and the copy-on-write log has to be updated on every allocation, it is the thing that takes the filesystem down.
The fix
du -sh ~/.local/share/opencode/snapshot
rm -rf ~/.local/share/opencode/snapshot
btrfs filesystem usage -T /home
I also set snapshot: false in ~/.config/opencode/opencode.json. The cost is
that I lose the revert button in the TUI for files the agent changed. I would
rather have that than a filesystem that hangs the whole machine, and on my setup
I almost never need a session-level undo because every real project is already a
Git repository with its own history.
What I would tell past-me
Check the metadata block group, not df, the first time a Btrfs box starts
misbehaving with free space on the display. It takes one command.
If you run a coding agent on a large repository, set snapshot: false before
you need it. The cost is one feature. The alternative is a hidden repository
under your home directory that grows with no upper bound and no cleanup.
And when something writes more than its share of a filesystem, go read its
config files. The answer was in a worktree line in a bare Git config inside
~/.local/share, and it took ten minutes once I knew to look there.
No comments yet.