Deleted Is Not Gone
One-line summary
Deleting a file and committing means only that the file is gone from now on; the objects left in the history remain as they are. To really remove it, you have to rewrite the history (filter-branch), strip the backup references, empty the reflog, and even run gc --prune=now. Even then, copies that have already spread remain.
Why this is needed
Two accidents have the same shape.
One is a secret. You accidentally committed a .env containing a production token, noticed, and deleted it. If you open the repository, that file is not there. But git log -- conf/prod.env still produces two commits, and git cat-file -p <커밋>:conf/prod.env prints the token in plain text. Git is a content-addressed store, so once a blob has come in, it does not disappear as long as a commit pointing to it remains.
The other is size. If you commit a data file or build output of several hundred MB once, even if you delete it afterward, everyone who clones keeps downloading that object. The answer to "why does cloning our repository take 10 minutes" is usually here.
Both are solved not by going forward but by fixing past history.
How it works
git filter-branch recreates the commits of the specified refs one by one, applying a filter to each commit. To lift a file out, --index-filter, which fixes only the index, is the fastest — because it does not lay out the working tree.
FILTER_BRANCH_SQUELCH_WARNING=1 git filter-branch --index-filter \
'git rm -r --cached --ignore-unmatch conf/prod.env assets/blob.bin' \
--prune-empty --tag-name-filter cat -- --all
Each option has a reason.
| Option | Why it is needed |
|---|---|
--ignore-unmatch |
So that git rm does not fail on commits without that file and stop everything |
--prune-empty |
So that commits that only added that file do not remain as empty commits |
--tag-name-filter cat |
So that tags do not remain pointing at the old commits |
-- --all |
Rewrite all refs, not just one branch |
FILTER_BRANCH_SQUELCH_WARNING |
To skip the 10-second warning git shows |
This is where most people stop. After the command finishes, git log shows the file is gone, but git rev-list --objects --all still lists it. The reason is refs/original/. filter-branch backs up all the original refs under it so that you can undo, and since those are refs too, they hold on to the old commits and old blobs.
So there is an order.
- Delete
refs/original/— list them withgit for-each-ref --format='%(refname)' refs/originaland rungit update-ref -don each. - Empty the reflog —
git reflog expire --expire=now --expire-unreachable=now --all. The reflog is also a reference. It remembers every place HEAD has passed through. - Delete the unreachable objects with
git gc --prune=now. Without--prune=now, they are deleted only after the default grace period (2 weeks) has passed.
If you measure before and after with git count-objects -vH, what this process actually did shows up in numbers.
What it looks like in the field
When the rewrite is done, every commit hash changes. A commit hash is calculated including its parent and tree, so if you fix one thing at the very start of the history, everything after it becomes different. So the following things come with it.
- Everyone on the team must fetch the force-pushed branch. If each person runs
git pull, the old history and the new history get mixed and the commits become two copies. Right after a rewrite, "everyone clone afresh" is the safest guidance. - Every open PR and CI pipeline breaks. Because the commits they pointed to are gone.
- If there is a parent that pins this repository as a submodule, that pin is broken.
And the most important thing — the objects remain as they are in copies that have already been cloned. They remain on colleagues' laptops, in CI caches, in forked repositories and in backups. So when a secret has gone into the history, the first thing to do is not cleaning the history but revoking that secret and issuing a new one. Cleaning the history is second. If you change the order, a live credential stays out there for the days you spend cleaning.
In a large repository, git filter-branch is slow. The official documentation also recommends tools such as git filter-repo or BFG. But those tools have to be installed separately, and what they do and the order of the cleanup afterward are exactly the same — rewrite, remove backup references, expire the reflog, gc.
What you will do in the next lab
You build a history that mixes a configuration file containing a token with a 900 KB chunk, and first confirm that even when you delete it from the working tree, the original text comes out with git cat-file. Then you run the rewrite, find for yourself the reason the objects are still there (refs/original) and strip it away, and measure that the size actually shrinks with git count-objects -vH. Finally you show that the secret is alive as it is in a colleague's copy that was taken beforehand, and end with what you must do first because of that.
The official documentation is git-filter-branch, git-reflog and git-gc.