9 minutes
How Git Actually Works: A Tour of Git Internals
Meet git, from 10,000 feet
Git is not a diff tool. It doesn’t store “what changed between commit A and commit B” and reconstruct history from a chain of patches. Every commit is a full snapshot of your project, and git figures out diffs on demand, for display purposes only, when you ask for them. That single idea, snapshots instead of diffs, is why almost everything else in git works the way it does: branching is cheap, history never has to “replay” anything to check out an old version, and, as you’ll see below, git can implement its entire history as nothing more than a tiny key-value store plus a handful of pointers.
This post is for people who already use add, commit, branch, and merge day to day and want to see what’s actually happening underneath. It takes that snapshot idea and walks down through four altitudes, from “what git is for” down to “the literal bytes on disk”, with a hands-on demo and a short animation at each level. The first two altitudes stay close to the commands you already know; the last one gets into git’s plumbing.
7,000 ft: the three trees
Everything you do with git moves content between three places:
- Working directory — the actual files on disk, the ones your editor sees
- Staging area (a.k.a. the index) — a draft of what your next commit will look like
- Repository (
.git) — the permanent, committed history
mkdir three-trees-demo && cd three-trees-demo
git init -b main
echo "v1" > notes.txt
git status --short
?? notes.txt
?? means: git sees the file in your working directory, it’s not staged, and it’s not committed. git add copies it into the staging area:
git add notes.txt
git status --short
A notes.txt
A in the first column means it’s staged (that column always describes the staging area vs. the repository). git commit then takes whatever is staged and seals it into the repository as a new snapshot:
git commit -m "Add notes"
git status --short
git status --short now prints nothing at all, the working directory matches the last commit exactly. Edit the file again without staging it, and git diff compares working directory against the staging area; git diff --staged compares the staging area against the last commit:
echo "v2" >> notes.txt
git diff --stat
notes.txt | 1 +
1 file changed, 1 insertion(+)
git add notes.txt
git diff --stat
Empty output again, nothing outside the staging area differs from what’s staged. Now compare the staging area against the last commit instead:
git diff --staged --stat
notes.txt | 1 +
1 file changed, 1 insertion(+)
Every git command you use day to day is really just moving a snapshot between these three trees. Here’s what that looks like end to end, init through a few commits and a branch:
5,000 ft: commits are a graph, branches are pointers
Zoom out one level: a repository isn’t a single line of commits, it’s a directed graph. Each commit stores a pointer to the commit(s) before it, its “parent(s)”. Walk those parent pointers backward from any commit and you get its full history, no separate “history log” is stored anywhere, the graph is the log.
A branch is nothing more than a movable pointer at a specific commit, and HEAD is a pointer that (usually) points at a branch, telling git “this is the branch I’m currently on.” git commit doesn’t just create a commit, it also drags the current branch pointer forward to the new commit. git checkout <branch> doesn’t touch any commit at all, it just moves HEAD to point at a different branch.
This is why merging comes in two flavors:
- Fast-forward merge — if the branch you’re merging into hasn’t moved since you branched off, there’s no real merging to do, git just slides that branch’s pointer forward to match the tip of the other branch.
- Three-way merge — if both branches gained commits since they diverged, git finds their common ancestor, compares both tips against it, and creates a brand new commit with two parents, one from each branch.
mkdir merge-demo && cd merge-demo
git init -b main
git commit --allow-empty -m "M0"
git branch feature
git checkout feature
git commit --allow-empty -m "F0"
git checkout main
git merge feature
git log --oneline --graph --decorate
* 5bcf3ae (HEAD -> main, feature) F0
* f85470e M0
One line, no merge commit, main simply caught up to feature. Fast-forward. Do it again but commit on main too before merging:
git checkout -b feature2
git commit --allow-empty -m "F1"
git checkout main
git commit --allow-empty -m "M1"
git merge feature2 -m "Merge feature2 into main"
git log --oneline --graph --decorate
* cb3e949 (HEAD -> main) Merge feature2 into main
|\
| * 0018ee9 (feature2) F1
* | 3c45f86 M1
|/
* 5bcf3ae (feature) F0
* f85470e M0
Now there’s a real merge commit with two parents. Here are both cases animated, first the fast-forward:
And the three-way merge, where both sides moved on before merging:
1,000 ft: the object database
This is where “snapshots, not diffs” stops being an abstract claim. Every commit, every directory listing, every version of every file you’ve ever committed is stored as an object, and every object is named by the SHA-1 hash of its own content. Same content anywhere in history, same hash, same object, stored once. This is a content-addressable store, git is really just a key-value database with three kinds of values on top of it: blobs, trees, and commits.
Let’s build one commit and look at exactly what git wrote to disk.
mkdir objects-demo && cd objects-demo
git init -b main
echo "Hello, Git!" > hello.txt
Nothing git-related has happened yet, that’s just a file. Ask git what the hash of its content would be, without storing anything:
git hash-object hello.txt
670a245535fe6316eb2316c1103b1a88bb519334
Now actually store it as a blob object:
git hash-object -w hello.txt
670a245535fe6316eb2316c1103b1a88bb519334
Same hash, because the hash is purely a function of the content, git computed it and wrote it to .git/objects this time. Confirm it’s really there, and what type of object it is:
git cat-file -t 670a245535fe6316eb2316c1103b1a88bb519334
git cat-file -p 670a245535fe6316eb2316c1103b1a88bb519334
blob
Hello, Git!
That’s it, that’s the entire blob: the raw file content, nothing else. No filename, no permissions, no history. git add does exactly this hash-object step for you and records the result in the staging area. Commit it:
git add hello.txt
git commit -m "Add hello.txt"
git rev-parse HEAD
aba6a2103ae9d634340f848fa05c12eb8e672aab
Look inside that commit object:
git cat-file -p aba6a2103ae9d634340f848fa05c12eb8e672aab
tree d3ec8a0f5950fb1f73ce0d1ed55cd6fa7afcdeb9
author Vlad Flore <flore.vlad@gmail.com> 1789553052 +0200
committer Vlad Flore <flore.vlad@gmail.com> 1789553052 +0200
Add hello.txt
A commit is just a small text record: which tree represents the project’s root directory at this point, who wrote it and when, and the message. No parent line here, this is the first commit. Look inside that tree:
git cat-file -p d3ec8a0f5950fb1f73ce0d1ed55cd6fa7afcdeb9
100644 blob 670a245535fe6316eb2316c1103b1a88bb519334 hello.txt
A tree is a directory listing: mode, object type, hash, and filename, one line per entry. For a file, that entry points at a blob. For a subdirectory, it would point at another tree, trees can nest inside trees. Three tiny object types, and together they can represent an entire filesystem at a single point in time: blob = file content, tree = directory listing, commit = a tree plus metadata plus parent(s).
Now change the file and commit again:
echo "Hello, Git! v2" > hello.txt
git add hello.txt
git commit -m "Update hello.txt"
git cat-file -p HEAD
tree a6b9f4683657af5aff36fbfecfafe61e14a169e0
parent aba6a2103ae9d634340f848fa05c12eb8e672aab
author Vlad Flore <flore.vlad@gmail.com> 1789553052 +0200
committer Vlad Flore <flore.vlad@gmail.com> 1789553052 +0200
Update hello.txt
A new tree, a new parent line pointing straight at the previous commit, and, underneath, a whole new blob (14f5f0b...) for the new file content. Nothing about the first commit changed, nothing got overwritten, git just added three new objects (blob, tree, commit) and pointed the newest one at the previous one. That’s the parent-pointer chain that git log walks:
find .git/objects -type f | sort
.git/objects/14/f5f0b4869f4268560680ceb02ce8df831a74dc
.git/objects/67/0a245535fe6316eb2316c1103b1a88bb519334
.git/objects/a2/d9e32136d0bca38251ad19a71f284d4acd2995
.git/objects/a6/b9f4683657af5aff36fbfecfafe61e14a169e0
.git/objects/ab/a6a2103ae9d634340f848fa05c12eb8e672aab
.git/objects/d3/ec8a0f5950fb1f73ce0d1ed55cd6fa7afcdeb9
Six objects for two commits: 2 blobs, 2 trees, 2 commits. (The first two characters of the hash become the subdirectory name, purely so .git/objects doesn’t end up with tens of thousands of files in one flat directory.)
Here’s that exact sequence animated, blob, tree, commit, and the second generation branching off the first:
Branches and HEAD are just files
Given all that, what is a branch, mechanically? Look inside .git:
cat .git/refs/heads/main
a2d9e32136d0bca38251ad19a71f284d4acd2995
That’s the whole thing. main is a 41-byte text file containing one commit hash. git commit advances the current branch by overwriting this file with the new commit’s hash. git branch other-name creates a new file with the same content. git branch -D other-name deletes a file. There’s no special “branch” object type, a branch is a ref, a named pointer to a commit, living as a plain file under .git/refs/heads/.
cat .git/HEAD
ref: refs/heads/main
And HEAD is one level of indirection on top of that: normally it doesn’t hold a commit hash at all, it holds the name of the branch you’re on. git checkout other-branch rewrites this one line. When you check out a specific commit instead of a branch (a “detached HEAD”), HEAD briefly holds a raw commit hash instead of a ref: line, same file, same idea.
Why .git doesn’t grow forever
If every commit adds new objects and nothing ever gets deleted, doesn’t .git balloon indefinitely? For a while, yes, each object above got written as its own file (a “loose object”). Periodically (and automatically, when there are enough loose objects lying around) git runs garbage collection, which packs many objects together into a single compressed packfile, storing similar objects as deltas against each other instead of full copies. This is why cloning a repository with years of history is usually a fraction of the size you’d expect: git gc and the packfile format are doing real compression work behind the scenes, without changing anything about the object model above. It’s still blobs, trees, and commits, just stored more efficiently.
Conclusion
Zoom back out: git moves your files between a working directory, a staging area, and a repository (7,000 ft). Inside that repository, commits form a graph via parent pointers, and branches are just movable labels on that graph (5,000 ft). And underneath all of that, the entire thing is a content-addressable key-value store of three object types, blobs, trees, and commits, glued together with hashes, plus a couple of tiny text files for refs and HEAD (1,000 ft, ground level). Every git command you’ll ever run is doing some combination of: write an object, update a ref, or move HEAD. Once that clicks, git stops being a wall of memorized commands and starts being a small, well-defined mechanism you can actually reason about, and the commands that manipulate history directly, git reflog, git worktree, and git filter-repo, stop looking like magic.