AI and machine-learning projects often need to version model weights, datasets, images, or audio. Add a multi-gigabyte .safetensors file directly with git add and the repository soon grows, clones slow down, or the hosting provider rejects the push. Git LFS (Large File Storage) addresses this problem. This guide covers its internals, practical use, migration, CI, hosting, self-hosting, and alternatives.
1. Why use Git LFS?
1.1 Git’s difficulty with large files
Git is a content-addressed snapshot system. Each version of each file is stored as a blob under .git/objects, indexed by a content hash. This works well for source code, but large binary files present several problems:
- History retains old versions. Once a version is committed, it remains in history. Deleting the file later does not remove those objects from existing clones. Ten versions of a 2 GB model can occupy roughly 20 GB before compression.
- Delta compression is often ineffective. Git attempts to delta-compress related blobs in packfiles. It works well for text, but compressed binary formats such as images, video, and model weights often change too much for useful deltas.
- A normal clone downloads history. By default,
git clonefetches reachable history and its objects, including versions you may not need. - Hosting limits apply. Many providers limit individual files. For example, GitHub blocks regular Git files larger than 100 MiB.
1.2 The core idea behind Git LFS
Git LFS keeps a small text pointer file in the Git repository and stores the real file content on an LFS server.
An LFS server is not built into Git. Git handles version control; a separate service implements the Git LFS API and stores the LFS objects. Commonly, the server is part of the hosting platform: GitHub, GitLab, Gitea/Forgejo, Bitbucket, and Hugging Face provide LFS support alongside their repositories and access controls. The client can usually derive the LFS endpoint from the Git remote URL (see section 2.3). Alternatively, you can self-host a Git service with LFS support or run a separate LFS server.
A bare repository accessed only over SSH, such as ssh://server/srv/repo.git, provides Git but not an LFS server. LFS pushes to it will fail unless you provide an LFS endpoint separately.
- When a file is added, Git LFS stores its content in the local LFS cache and gives Git a pointer. The LFS object is uploaded when you push.
- When a pointer is checked out, Git LFS retrieves the corresponding content and writes it into the working tree. Whether this happens immediately depends on the client configuration; you can skip it and run
git lfs pulllater.
Git history therefore stores small pointers instead of large blobs, reducing ordinary Git object downloads and repository size. LFS data remains separately stored and has its own transfer and storage costs.
2. How it works
2.1 Pointer format
A file tracked by LFS is represented in the Git object database by a small text file like this:
version https://git-lfs.github.com/spec/v1
oid sha256:4d7a214614ab2935c943f9e0ff69d22eadbb8f32b1258daaa5e2ca24d17e2393
size 12345
| Field | Meaning |
|---|---|
version | Pointer specification version, currently spec/v1. |
oid | SHA-256 hash of the original file; also the unique LFS object identifier. |
size | Size of the original file in bytes. |
Because the object ID is a content hash, LFS deduplicates identical files across paths and commits. A one-byte change creates a new complete object; this matters for quotas and alternatives discussed later.
Inspect the pointer stored in Git with:
# Show the actual content stored in Git for this path at HEAD (the pointer)
git cat-file -p HEAD:models/model.safetensors
# Or parse the LFS pointer
git lfs pointer --file=models/model.safetensors
2.2 Clean and smudge filters
Git LFS does not modify Git. It uses Git’s built-in filter mechanism to transform content between the working tree and Git’s object database. A filter driver has two directions:
| Direction | When it runs | Purpose |
|---|---|---|
| clean | git add: working tree to index/object database | Converts working-tree content into the form stored in Git. |
| smudge | checkout: object database to working tree | Converts stored content into the form used in the working tree. |
Both commands read file content from standard input and write transformed content to standard output. Git invokes them but does not know what transformation they perform.
A filter needs two things:
- A filter command in Git configuration (
.git/configor~/.gitconfig). - A
.gitattributesrule that selects paths for that filter.
For example, this unrelated filter replaces a password in the Git copy while leaving the local file alone:
# .git/config
[filter "hidepass"]
clean = sed -e 's/^password = .*/password = REDACTED/'
smudge = cat
# .gitattributes
config.ini filter=hidepass
The filter command is not distributed with the repository: .gitattributes is committed, but .git/config is not. If a collaborator has no matching filter configured, Git may add the file unchanged as a regular blob instead of a pointer. Install Git LFS and run git lfs install on each user account that needs it.
Git compares and stores the clean-filter output. The working-tree content itself is not what the Git object database records.
How Git LFS uses filters
Git LFS uses clean to replace a large file with a pointer and smudge to replace the pointer with the original content. After git lfs install, global Git configuration includes entries like these:
[filter "lfs"]
clean = git-lfs clean -- %f
smudge = git-lfs smudge -- %f
process = git-lfs filter-process
required = true
The repository’s .gitattributes selects paths:
*.safetensors filter=lfs diff=lfs merge=lfs -text
- clean: On
git add,git-lfs cleancalculates the SHA-256 hash, saves the original file under.git/lfs/objects/<first-two-oid-chars>/<next-two-oid-chars>/<oid>, and writes the pointer for Git to store as a blob. - smudge: On checkout,
git-lfs smudgereads the pointer, checks the local cache, downloads the object if needed, then writes the original content into the working tree. process: Keeps a filter process running so Git does not need to start a newgit-lfsprocess for every file.required = true: If a filter configured on this machine fails, Git should fail the operation. This cannot configure Git LFS for a collaborator who has not installed it.diff=lfs: Uses the LFS diff behavior rather than attempting to diff large binary data.merge=lfs: Uses the LFS merge driver where applicable.-text: Disables line-ending conversion that could corrupt binary files.
git lfs install also installs repository hooks, especially pre-push. Before Git updates the remote ref, the hook uploads LFS objects referenced by the commits that the server does not have. In an existing clone, run the command in that repository to install or update its hooks.
2.3 Transfer and the Batch API
LFS transfers use an HTTP API separate from the Git protocol. The endpoint is usually derived from the Git remote URL, for example:
https://git.example.com/user/repo.git → https://git.example.com/user/repo.git/info/lfs
A repository .lfsconfig can set lfs.url explicitly, for example when Git and LFS use different hosts.
The central endpoint is the Batch API. A download typically follows this flow:
Git client LFS server Object storage (S3, etc.)
│ │ │
│ 1. checkout encounters pointer│ │
│──── POST /objects/batch ────▶│ │
│ {operation: download, │ │
│ objects: [{oid,size}]} │ │
│ │ 2. authorize and locate object│
│◀──── 200 actions.download ───│ │
│ {href, header, expires_in} │ │
│ │ │
│──────── 3. GET href (possibly a signed URL) ────────────────▶│
│◀──────────────────────── 4. file content ──────────────────│
│ │ │
│ 5. verify SHA-256; save in local LFS cache and working tree │
A request looks roughly like this:
POST /user/repo.git/info/lfs/objects/batch
Accept: application/vnd.git-lfs+json
Content-Type: application/vnd.git-lfs+json
{
"operation": "download",
"transfers": ["basic"],
"ref": { "name": "refs/heads/main" },
"objects": [
{ "oid": "4d7a2146...", "size": 12345 }
]
}
For each object, the server returns actions:
- A download action includes an
href, any requiredheadervalues, and an expiry. - An upload action may be omitted if the server already has the object. Some servers also return a
verifyaction for post-upload confirmation.
This design batches requests and separates the control plane from the data plane. The LFS server authorizes requests and issues URLs; file bytes can flow directly between client and object storage or a CDN.
Additional details:
transfersnegotiates the transfer method.basicuses ordinary HTTP PUT/GET; custom transfer agents can implement alternatives such as multipart object-storage uploads.- With an SSH remote, the client can run
git-lfs-authenticateover SSH to obtain HTTP credentials. Git LFS 3 also supports a pure SSH transfer protocol when the server implementsgit-lfs-transfer. - For troubleshooting,
GIT_TRACE=1 GIT_CURL_VERBOSE=1 git lfs pullshows request details.
3. Everyday use and migration
3.1 Install and track files
Install Git LFS through your platform’s package manager:
# macOS
brew install git-lfs
# Debian / Ubuntu
sudo apt install git-lfs
# Fedora / RHEL
sudo dnf install git-lfs
# Windows: included with Git for Windows
Then initialize its Git configuration:
git lfs install
This command configures Git to use Git LFS; it does not install the program itself. It writes the filter configuration to the user’s global Git config and, when run inside a repository, installs or updates repository hooks. Run it once for each user account to configure filters, and in an existing repository when its hooks need installation or repair.
Check the global filter configuration with:
git config --global --get-regexp '^filter\.lfs\.'
Common variants:
| Command | Scope |
|---|---|
git lfs install | Current user (writes ~/.gitconfig). |
git lfs install --system | All users on the machine (system Git config; usually requires administrator/root). |
git lfs install --local | Current repository only (.git/config). |
git lfs install --skip-repo | Global configuration only; do not install hooks in the current repository. |
git lfs uninstall | Remove the corresponding configuration and hooks. |
In CI, configure Git LFS in the runner image or working repository instead of depending on an interactive user’s global configuration.
Track files by extension or directory:
# Quote the pattern so the shell does not expand *
git lfs track "*.safetensors"
git lfs track "*.parquet"
# Track a directory
git lfs track "datasets/**"
# Show current rules
git lfs track
git lfs track writes rules to .gitattributes. Commit that file so collaborators know which paths use LFS:
git add .gitattributes
git add models/model.safetensors
git commit -m "Add model weights via LFS"
git push
Useful inspection commands:
# List LFS files at the checked-out revision (* means content is present; - means pointer only)
git lfs ls-files
# Include object sizes
git lfs ls-files --size
# Inspect whether staged files are stored as LFS pointers
git lfs status
# Check local LFS objects and pointers
git lfs fsck
# Show endpoint, cache path, and environment details
git lfs env
3.2 Common .gitattributes pitfalls
Add the rule before adding the file. A tracking rule affects future git add operations. If a file was already committed as a normal blob, adding a rule later does not convert its old history. To convert only the current version:
git lfs track "*.bin"
git add --renormalize .
git commit -m "Move *.bin to LFS"
This does not shrink old history; use git lfs migrate in section 3.3 for that.
Patterns differ from .gitignore. .gitattributes patterns do not support negation (!). A leading slash is relative to that .gitattributes file. Use ** to match directories at arbitrary depth:
# CSV files directly under root data/
/data/*.csv filter=lfs diff=lfs merge=lfs -text
# CSV files at any depth
*.csv filter=lfs diff=lfs merge=lfs -text
# Everything below assets/
assets/** filter=lfs diff=lfs merge=lfs -text
Patterns are case-sensitive: *.png does not match IMAGE.PNG. On case-insensitive macOS and Windows filesystems, this can be easy to miss; add both cases where necessary.
Avoid tracking small text files such as configuration, JSON, or sample CSV files. Git handles those well and provides readable diffs. Do not track .gitattributes itself: an overly broad rule such as * can capture it and break the setup.
3.3 Migrate an existing repository
If large files already exist as regular Git blobs, git lfs migrate can rewrite history and convert them to LFS.
First, inspect the repository:
# Summarize the largest file types across all refs and history
git lfs migrate info --everything --top=10
Then import selected files, rewriting history:
# Convert these patterns across all branches and tags
git lfs migrate import --everything --include="*.safetensors,*.bin,*.parquet"
# Or select files above a size threshold
git lfs migrate import --everything --above=50MB
Rewriting creates new commits, so affected commit IDs change. Force-push the rewritten refs after checking the result:
git push --force --all
git push --force --tags
Before rewriting history:
- Make a backup, such as a
git clone --mirrorcopy. - Notify collaborators. Their old branches will not merge cleanly with rewritten history; recloning may be simplest.
- Open pull/merge requests may need to be recreated.
- Old server-side objects may continue to consume storage until the host performs garbage collection or support intervenes.
Import without rewriting existing history:
# Add a new commit that converts the selected tracked file
git lfs migrate import --no-rewrite path/to/large.bin
This keeps old commits but leaves old versions in Git. It can be useful when a team adopts LFS without disrupting existing branches.
Export from LFS: If you stop using LFS, convert matching pointers back to ordinary Git blobs:
git lfs migrate export --everything --include="*.png"
4. Advanced techniques
4.1 Selective downloads
If a repository contains many LFS objects but you need only a subset, skip the initial downloads and fetch files on demand:
GIT_LFS_SKIP_SMUDGE=1 git clone https://git.example.com/user/models.git
cd models
# The working tree contains pointers; fetch only the requested files
git lfs pull --include="llama-8b/*"
Persistent include and exclude patterns apply to later pull and checkout operations:
git config lfs.fetchinclude "configs/*,small-model/*"
git config lfs.fetchexclude "datasets/raw/*"
git lfs fetch --recent also downloads objects referenced by recent branches and commits. Tune the time ranges with lfs.fetchrecentrefsdays and lfs.fetchrecentcommitsdays; this can prepare branches you expect to switch to while offline.
Git’s partial clone can further reduce normal Git object downloads:
git clone --filter=blob:none https://git.example.com/user/repo.git
4.2 Clean the local cache
.git/lfs/objects can retain versions you have checked out, and eventually take more space than the working tree. git lfs prune removes objects no longer needed locally. It retains the current checkout, recent references, stashes, other worktrees, and unpushed objects; exact behavior depends on recency settings and the remote.
Preview the cleanup first, then verify objects exist on the remote before deleting local copies:
# Preview the proposed cleanup
git lfs prune --dry-run --verbose
# Verify reachable objects exist on the remote
git lfs prune --verify-remote
Pruning only clears the local cache; it does not reduce server storage. Server-side retention and garbage collection depend on the hosting platform. Even after history is rewritten, old objects may remain until the platform collects them; some hosts require repository deletion or a support request to release the space.
4.3 File locking
Binary files cannot be merged like source code. If two people edit the same Photoshop file or Unity scene, one person’s work may be lost. Git LFS locking can help coordinate edits:
# Mark matching paths as lockable
git lfs track "*.psd" --lockable
The lockable attribute makes checked-out files read-only until a lock is acquired:
git lfs lock design/banner.psd # Acquire a lock; the file becomes writable
git lfs locks # List locks
git lfs unlock design/banner.psd # Unlock after editing and pushing
git lfs unlock --force design/banner.psd # Force-unlock as an administrator
Lock verification requires server support for the Locking API. Configure lfs.<url>.locksverify according to the Git LFS documentation. When enabled, the pre-push hook asks the server whether files being pushed are locked by someone else. Confirm that your host supports locking before depending on this workflow. It is often unnecessary for code or ML projects, but useful for game and design assets.
4.4 Use LFS in CI
CI jobs may download LFS objects repeatedly. Enable LFS checkout only for jobs that need those files. GitHub Actions’ actions/checkout provides an lfs: true input:
steps:
- uses: actions/checkout@v4
with:
lfs: true
If you cache .git/lfs, use a cache key that reflects the objects needed by the job. A key based only on branch or filename can reuse stale content after a file changes. Keep checkout and cache steps in the right order, and verify restored content with git lfs fsck or a content-hash check in the workflow. For lint and documentation jobs that do not need binaries, skip LFS downloads.
Other ideas:
- Download only the files the job needs, for example
git lfs pull --include="tests/fixtures/*". - Do not pull LFS objects in jobs that do not use them.
- A self-hosted runner can keep a persistent cache and point
lfs.storageat it.
5. Hosting and self-hosting
5.1 Hosted platforms
GitHub
GitHub’s per-file limits and storage/bandwidth allowances depend on the plan and can change. Downloads count against the repository owner’s bandwidth; fork traffic may also be charged to the upstream repository. Check GitHub’s current LFS quota and limit documentation, and account for repeated CI downloads.
GitLab
GitLab supports LFS. Self-Managed instances can configure LFS object storage, including an external object-storage service. GitLab.com storage and traffic rules depend on plan and namespace settings; consult the current LFS object-storage documentation.
Hugging Face Hub
The Hub now uses Xet as its large-file storage backend while maintaining a Git and Git-LFS-pointer compatibility path. Git LFS identifies complete files; Xet supports chunk-level deduplication. To use Xet’s transfer optimizations, install and configure a client that supports it as described in the Hub documentation. Behavior varies by client and version.
5.2 Self-host an LFS server
Self-hosting can avoid a provider’s particular quotas, but storage, backups, bandwidth, availability, and operations still have costs. Common options include:
| Option | When it fits |
|---|---|
| Gitea / Forgejo | Lightweight deployment with built-in LFS and configurable object storage. |
| GitLab Self-Managed | You need GitLab integration and can operate its LFS storage. |
| Standalone LFS server | Your Git host lacks LFS and you can integrate authentication, storage, and backups yourself. |
Use the LFS and object-storage documentation for the exact deployed version. Configuration names, storage backends, and direct-download behavior vary by product and version; do not copy an app.ini example from a different release. Before going live, test clone, checkout, push, access control, and backup restoration. Confirm that clients can reach the download URLs the server generates.
5.3 Reverse-proxy considerations
Self-hosted LFS failures often come from the reverse proxy in front of the Git service. Large HTTP requests, request-size limits, timeouts, buffering, and upstream connection settings can all affect transfers. Configure the proxy for the actual server and test with files close to the expected maximum size.
Nginx:
server {
listen 443 ssl;
server_name git.example.com;
# Remove the request-body size limit (default is 1 MB)
client_max_body_size 0;
location / {
proxy_pass http://127.0.0.1:3000;
# Stream requests rather than buffering the whole body first
proxy_request_buffering off;
proxy_buffering off;
# Large uploads can take time
proxy_read_timeout 3600s;
proxy_send_timeout 3600s;
proxy_set_header Host $host;
proxy_set_header X-Forwarded-Proto $scheme;
}
}
If Caddy or another proxy has a request-body limit, make sure the LFS upload path allows the expected size. Check its timeout and buffering behavior as well.
For a CDN or tunnel, check the plan’s limits for request size, connection duration, and upload behavior. If the proxy is not suitable for large objects, configure a separate LFS endpoint or direct object-storage transfers using a method supported by your Git service. Verify generated URLs, authentication, and TLS, and do not expose a private storage endpoint to unauthorized clients.
Troubleshooting clues:
413on upload: request-size limit.- Upload interrupted with
504: timeout. - Download URL points to a private address: check the public URL or object-storage endpoint configuration.
- Authentication failure: check forwarding of the
Authorizationheader and scheme consistency forX-Forwarded-Proto.
6. Troubleshooting and alternatives
6.1 Common problems
After cloning, a file contains pointer text instead of its content
Git LFS may not be installed or initialized on this machine, or the checkout may have intentionally skipped downloads:
git lfs install
git lfs pull
If GIT_LFS_SKIP_SMUDGE=1 was used, git lfs pull downloads the content. If the pull fails, check credentials, the LFS endpoint, whether the object was uploaded, and any quota messages from the host.
smudge filter lfs failed
The checkout failed to download an LFS object. Check:
- Network and credentials: use
git lfs envto check the endpoint andGIT_TRACE=1 git lfs pullfor details. - Whether the object exists on the server. A pointer may have been committed without a successful LFS upload, for example from an environment without Git LFS or after bypassing
pre-push. - Whether a storage or bandwidth quota was reached.
To finish checkout and retry later, temporarily skip smudging:
GIT_LFS_SKIP_SMUDGE=1 git checkout <branch>
Encountered N file(s) that should have been pointers, but weren't
A path selected by .gitattributes exists as a regular blob rather than an LFS pointer. One common cause is a commit made on a machine without Git LFS. Convert the current version without rewriting history:
git add --renormalize .
git commit -m "Fix files that should be LFS pointers"
Or rewrite history to convert old versions as well:
git lfs migrate import --everything --include="<pattern>"
A large file was committed as a regular Git object and already pushed
Deleting it in a later commit does not shrink the repository. Rewrite history with git lfs migrate import (section 3.3). If it has not been pushed, use git reset --soft HEAD~1, add the tracking rule, then add and recommit the file.
Pushes are very slow or appear stuck
- Use
git lfs push --dry-run origin mainto inspect the object count and size. - Adjust parallel transfers with
git config lfs.concurrenttransfers 8. - For self-hosted deployments, inspect reverse-proxy buffering (section 5.3).
Useful diagnostic commands:
| Command | Purpose |
|---|---|
git lfs env | Show version, endpoint, cache path, and configuration. |
git lfs status | Show LFS status for the index and working tree. |
git lfs ls-files --size | List LFS files and their sizes. |
git lfs fsck | Check local object integrity. |
git lfs logs last | Show the most recent error log. |
GIT_TRACE=1 GIT_CURL_VERBOSE=1 | Print HTTP request details. |
6.2 When not to use LFS
Git LFS works best for a moderate number of binary files that need to be versioned alongside code, such as game assets, test fixtures, or small models. Consider other options when:
| Option | Good fit | Characteristics |
|---|---|---|
| Git LFS | Moderate binary files that must version alongside code | Mature ecosystem, broad host support, transparent workflow. |
| git-annex | Many files, several storage backends, distributed backups | Uses symlinks and tracks where content is stored; powerful but has a steeper learning curve. |
| DVC | ML datasets and experiment pipelines | Git stores .dvc metadata; data lives in S3/GCS or another remote. Includes pipelines and experiment tracking and is independent of Git hosting. |
| Hugging Face Hub (Xet) | Sharing public or team models and datasets | Chunk-level deduplication can reduce storage and transfer for frequently updated large files. |
| Object storage plus a manifest | Terabyte-scale data without fine-grained version history | Store URLs and hashes in the repository; a script downloads and verifies the data. |
A few guidelines:
- For files that change frequently by small amounts, such as successive model checkpoints, LFS stores a complete new object each time. Chunk-deduplicating storage such as Xet, or keeping only selected checkpoints, may fit better.
- At hundreds of gigabytes, hosted LFS pricing may exceed direct object storage. Compare DVC or object storage plus a manifest.
- If you need data lineage and experiment management, DVC provides more than storage.
- Do not version files that do not need version history, such as build outputs; use releases or a package registry instead.
References
- Git LFS documentation and command manuals
- Git LFS Batch API
- Git LFS locking API
- Git LFS migrate manual
- Git LFS prune manual
- GitHub LFS storage and bandwidth limits
- GitHub Actions checkout
lfsinput - GitLab LFS object storage
- Gitea Git LFS setup
- Hugging Face Hub Xet storage
- Git attributes and filter drivers
Conclusion
Git LFS uses Git’s filter mechanism to replace large files with pointers, then transfers the content through a separate HTTP API. Understanding clean/smudge filters and the Batch API helps explain most usage and deployment issues. Add .gitattributes rules before adding files, limit downloads in CI, and watch platform quotas. If you self-host, account for storage, backups, bandwidth, reverse proxies, and access control: self-hosting changes how capacity and cost are managed, but it does not remove those costs.