← PortfolioCompute and infrastructure

Oxen

End-to-end infrastructure for fine-tuning and inference

Team
Greg SchoeningerCEO
Scott Howardco-founder
Founded
2022
Invested
2023
The problem

How do you keep track of the data that makes a model yours?

Click an image to edit it, or an empty slot to add one, and commit. Point at a dot in the history to look at an older version.A folder of images kept as a Merkle tree, where each box holds a hash of everything beneath it. Editing or adding one image changes only the hashes on the path up to the top, so a new commit stores those few boxes and points at everything else it already has. The faint outlines behind a box are its older versions, still in the history.An illustration, not real data.
How fine-tuning works

takes a model that was trained for one job and adapts it to a different, usually narrower one. In practice that means more training, on new data, starting from that already know a lot. You can retrain the whole network, or freeze most of its layers and change only a few. Tuning the full model often gives better results, but it costs more compute.

A cheaper trick is , short for low-rank adaptation. It freezes the original weights and slips small trainable matrices into each layer, so only those get updated. On GPT-3 it cut the number of trainable parameters by 10,000 times and the GPU memory needed by 3 times, and unlike some other adapters it adds no delay when the model runs. Think of it as tailoring a suit instead of weaving new cloth.

Either way, the result depends on the data you feed it, and data changes. For code, the answer to that problem is . Git stores each version of a project as a snapshot, and every clone carries the full history.

Further reading Fine-tuning (deep learning) (Wikipedia)LoRA: Low-Rank Adaptation of Large Language Models (arXiv)What is Git? (Pro Git (git-scm.com))About Version Control (Pro Git (git-scm.com))

Why it is hard
  1. i.

    Git was built for text

    Carrying the full history everywhere is fine for source files and ruinous for a pile of images. GitHub warns on files over 50 MiB, blocks files over 100 MiB, and recommends repositories stay under 1 GB, ideally, and under 5 GB at most. A training set of a million photos blows through that before lunch.

  2. ii.

    Pointers are not enough

    The usual patch is , which keeps a reference to each large file in the repository and stores the file itself somewhere else. It works, slowly. It still hashes and copies every file into the hidden .git directory. In a benchmark that pushed the more than a million images of ImageNet to remote storage, Git-LFS took more than 20 hours.

  3. iii.

    Data nobody can see

    So datasets tend to travel as tarballs or zip files, or get dumped into cloud storage, with little way to see what is inside. Once that happens, working out which change to which rows produced which model is detective work.

Further reading What is Git? (Pro Git (git-scm.com))About large files on GitHub (GitHub Docs)About Git Large File Storage (GitHub Docs)1 Million Files Benchmark (Oxen.ai Docs)

What Oxen is after

Oxen starts from a simple idea: the data behind a model deserves the same care as code, organised, versioned and shared. Its motto is that "your models are only as good as the data behind them."

It has grown into a pipeline for teams that want to own their models rather than rent generic ones, with the means to train them, version them, deploy them and improve them on their own terms. Fine-tuning is the step to reach for when prompting falls short, for instance when outputs need to be reliably correct, when private data is the edge, or when a big general model is too slow or too expensive.

Further reading Oxen.ai (Oxen.ai Docs)Fine-Tuning Models on Oxen.ai (Oxen.ai Docs)

How they go at it
  1. Step 1: Git's manners, new engine

    Oxen's commands feel like git (add, commit, push), but it was built from the ground up for large data rather than bolted onto Git. Under the hood it uses and fast hashing to cut how much a repository stores, plus block-level and compression. It is open source and meant to scale to millions of files and terabytes of data.

  2. Step 2: No full clone needed

    A remote server can take new files directly. To add one image to a million-file dataset, you stage it in a workspace on the server and commit there, without downloading everything first.

  3. Step 3: Diffs for tables

    For CSV and Parquet files, Oxen compares versions row by row and column by column, down to individual cells, so a change to a dataset reads more like a change to code.

  4. Step 4: Train from the repo

    Point a fine-tune at a dataset in a repository and Oxen provisions the GPUs, runs the job, writes the new weights back to the repository and shuts the GPUs down. The weights and the data are versioned together, so you can always tell which data trained which model, and you can deploy it to an endpoint or download it.

Further reading Version Control (Oxen.ai Docs)1 Million Files Benchmark (Oxen.ai Docs)Oxen.ai (Oxen.ai Docs)Dataset Diffs (Oxen.ai Docs)Fine-Tuning Models on Oxen.ai (Oxen.ai Docs)

Still open
  • How much data does fine-tuning really need?

    One study fine-tuned a 65B-parameter model on only 1,000 carefully curated examples and got strong results, which suggests almost all of a model's knowledge comes from pretraining. That puts the weight on quality over quantity, and makes it matter even more which thousand examples you picked.

  • Cheap adapters or the full model?

    In coding and maths, LoRA substantially underperformed full fine-tuning, but it held on better to what the base model could do outside the target domain. Fine-tuning of any kind can also leave a model shakier when real data drifts away from its training data. Where the right trade-off sits depends on the task, and nobody has a general rule.

Further reading LIMA: Less Is More for Alignment (arXiv)LoRA Learns Less and Forgets Less (arXiv)Fine-tuning (deep learning) (Wikipedia)

About Oxen

Oxen is an end-to-end platform for fine-tuning, deploying, and scaling AI models on proprietary data. It started with a simple insight: datasets deserve the same version control and collaboration infrastructure that code has. That foundation has expanded into a full production pipeline — dataset management, fine-tuning, and batch inference — for teams that want to own their models rather than depend on generic ones.

Oxen is particularly well-suited to creative and media workflows, where teams iterate on datasets locally and then scale jobs to production.

Words used here
Fine-tuning
Extra training that adapts an already-trained model to a narrower task or dataset.
weights
The numbers inside a neural network that training adjusts; together they are the model.
LoRA
Low-rank adaptation: fine-tuning that freezes a model and trains small add-on matrices instead.
version control
Software that records every change to a set of files so you can compare, undo and share versions.
Git LFS
Git Large File Storage, an add-on that swaps big files for pointers and keeps their contents on a separate server.
Merkle trees
A tree of hashes in which each node summarises everything beneath it, so a change can be spotted without rereading all the data.
deduplication
Storing identical chunks of data only once, however many files or versions contain them.
Sources