Oxen
End-to-end infrastructure for fine-tuning and inference
How do you keep track of the data that makes a model yours?
Extra training that adapts an already-trained model to a narrower task or dataset. takes a model that was trained for one job and adapts it to a different, usually narrower one. In practice that means more training, on new data, starting from The numbers inside a neural network that training adjusts; together they are the model. that already know a lot. You can retrain the whole network, or freeze most of its layers and change only a few. Tuning the full model often gives better results, but it costs more compute.
A cheaper trick is Low-rank adaptation: fine-tuning that freezes a model and trains small add-on matrices instead., short for low-rank adaptation. It freezes the original weights and slips small trainable matrices into each layer, so only those get updated. On GPT-3 it cut the number of trainable parameters by 10,000 times and the GPU memory needed by 3 times, and unlike some other adapters it adds no delay when the model runs. Think of it as tailoring a suit instead of weaving new cloth.
Either way, the result depends on the data you feed it, and data changes. For code, the answer to that problem is Software that records every change to a set of files so you can compare, undo and share versions.. Git stores each version of a project as a snapshot, and every clone carries the full history.
Further reading Fine-tuning (deep learning) (Wikipedia)LoRA: Low-Rank Adaptation of Large Language Models (arXiv)What is Git? (Pro Git (git-scm.com))About Version Control (Pro Git (git-scm.com))
- i.
Git was built for text
Carrying the full history everywhere is fine for source files and ruinous for a pile of images. GitHub warns on files over 50 MiB, blocks files over 100 MiB, and recommends repositories stay under 1 GB, ideally, and under 5 GB at most. A training set of a million photos blows through that before lunch.
- ii.
Pointers are not enough
The usual patch is Git Large File Storage, an add-on that swaps big files for pointers and keeps their contents on a separate server., which keeps a reference to each large file in the repository and stores the file itself somewhere else. It works, slowly. It still hashes and copies every file into the hidden .git directory. In a benchmark that pushed the more than a million images of ImageNet to remote storage, Git-LFS took more than 20 hours.
- iii.
Data nobody can see
So datasets tend to travel as tarballs or zip files, or get dumped into cloud storage, with little way to see what is inside. Once that happens, working out which change to which rows produced which model is detective work.
Further reading What is Git? (Pro Git (git-scm.com))About large files on GitHub (GitHub Docs)About Git Large File Storage (GitHub Docs)1 Million Files Benchmark (Oxen.ai Docs)
Oxen starts from a simple idea: the data behind a model deserves the same care as code, organised, versioned and shared. Its motto is that "your models are only as good as the data behind them."
It has grown into a pipeline for teams that want to own their models rather than rent generic ones, with the means to train them, version them, deploy them and improve them on their own terms. Fine-tuning is the step to reach for when prompting falls short, for instance when outputs need to be reliably correct, when private data is the edge, or when a big general model is too slow or too expensive.
Further reading Oxen.ai (Oxen.ai Docs)Fine-Tuning Models on Oxen.ai (Oxen.ai Docs)
- Step 1: Git's manners, new engine
Oxen's commands feel like git (add, commit, push), but it was built from the ground up for large data rather than bolted onto Git. Under the hood it uses A tree of hashes in which each node summarises everything beneath it, so a change can be spotted without rereading all the data. and fast hashing to cut how much a repository stores, plus block-level Storing identical chunks of data only once, however many files or versions contain them. and compression. It is open source and meant to scale to millions of files and terabytes of data.
- Step 2: No full clone needed
A remote server can take new files directly. To add one image to a million-file dataset, you stage it in a workspace on the server and commit there, without downloading everything first.
- Step 3: Diffs for tables
For CSV and Parquet files, Oxen compares versions row by row and column by column, down to individual cells, so a change to a dataset reads more like a change to code.
- Step 4: Train from the repo
Point a fine-tune at a dataset in a repository and Oxen provisions the GPUs, runs the job, writes the new weights back to the repository and shuts the GPUs down. The weights and the data are versioned together, so you can always tell which data trained which model, and you can deploy it to an endpoint or download it.
Further reading Version Control (Oxen.ai Docs)1 Million Files Benchmark (Oxen.ai Docs)Oxen.ai (Oxen.ai Docs)Dataset Diffs (Oxen.ai Docs)Fine-Tuning Models on Oxen.ai (Oxen.ai Docs)
How much data does fine-tuning really need?
One study fine-tuned a 65B-parameter model on only 1,000 carefully curated examples and got strong results, which suggests almost all of a model's knowledge comes from pretraining. That puts the weight on quality over quantity, and makes it matter even more which thousand examples you picked.
Cheap adapters or the full model?
In coding and maths, LoRA substantially underperformed full fine-tuning, but it held on better to what the base model could do outside the target domain. Fine-tuning of any kind can also leave a model shakier when real data drifts away from its training data. Where the right trade-off sits depends on the task, and nobody has a general rule.
Further reading LIMA: Less Is More for Alignment (arXiv)LoRA Learns Less and Forgets Less (arXiv)Fine-tuning (deep learning) (Wikipedia)
Oxen is an end-to-end platform for fine-tuning, deploying, and scaling AI models on proprietary data. It started with a simple insight: datasets deserve the same version control and collaboration infrastructure that code has. That foundation has expanded into a full production pipeline — dataset management, fine-tuning, and batch inference — for teams that want to own their models rather than depend on generic ones.
Oxen is particularly well-suited to creative and media workflows, where teams iterate on datasets locally and then scale jobs to production.
- Fine-tuning
- Extra training that adapts an already-trained model to a narrower task or dataset.
- weights
- The numbers inside a neural network that training adjusts; together they are the model.
- LoRA
- Low-rank adaptation: fine-tuning that freezes a model and trains small add-on matrices instead.
- version control
- Software that records every change to a set of files so you can compare, undo and share versions.
- Git LFS
- Git Large File Storage, an add-on that swaps big files for pointers and keeps their contents on a separate server.
- Merkle trees
- A tree of hashes in which each node summarises everything beneath it, so a change can be spotted without rereading all the data.
- deduplication
- Storing identical chunks of data only once, however many files or versions contain them.
- 1Fine-tuning (deep learning) · Wikipedia
- 2LoRA: Low-Rank Adaptation of Large Language Models · arXiv
- 3What is Git? · Pro Git (git-scm.com)
- 4About Version Control · Pro Git (git-scm.com)
- 5About large files on GitHub · GitHub Docs
- 6About Git Large File Storage · GitHub Docs
- 71 Million Files Benchmark · Oxen.ai Docs
- 8Oxen.ai · Oxen.ai Docs
- 9Version Control · Oxen.ai Docs
- 10Dataset Diffs · Oxen.ai Docs
- 11Fine-Tuning Models on Oxen.ai · Oxen.ai Docs
- 12LIMA: Less Is More for Alignment · arXiv
- 13LoRA Learns Less and Forgets Less · arXiv