# Lemurian Labs

Universal compiler and runtime for AI

- Team: Jay Dawani (CEO), Vassil Dimitrov (co-founder), Chris Vick (VP)
- Founded: 2018
- Invested: 2025
- Links: [Website](https://www.lemurianlabs.com)
- From our blog: [On Computers](https://pebblebed.com/blog/computing-is-bigger)
- Field: Compute and infrastructure

## The problem: How do you write AI code once and run it fast on any chip?

### How an ML compiler works

A model written in PyTorch or JAX is, underneath, a graph: matrix multiplies, additions and other operations wired together. Something has to turn that graph into instructions a particular chip can run. That something is an ML compiler. XLA, a widely used one, takes model graphs from frameworks like PyTorch, TensorFlow and JAX and compiles them into machine instructions for various architectures.

The work happens in stages, a process called lowering. First come passes that don't care about the hardware, such as fusing operations together and planning where intermediate results live in memory. Then a backend adds what it knows about the target chip, fuses further, and generates code, often through LLVM. It's a little like translating a recipe first into plain steps, then into instructions for one particular kitchen.

At the bottom sit kernels, the small, tightly tuned programs that do one operation on one kind of chip. On NVIDIA GPUs those are usually written with CUDA, Nvidia's proprietary platform for running general-purpose code on its GPUs. Most frameworks lean on libraries of vendor-specific kernels.

### Why it is hard

**Hand-tuned sets the bar.** Historically, compilers couldn't beat hand-written kernels, so people wrote kernels by hand, chip by chip. The gap can be large. FlashAttention, an attention algorithm designed around how data moves between a GPU's memory levels, trained GPT-2 about 3 times faster than existing baselines. Any compiler that wants to replace hand-tuning has to match work like that, automatically.

**Code that won't travel.** Kernels are tied to the hardware they were written for. Frameworks optimise for a narrow range of server GPUs, and moving a workload to new platforms, from phones to custom accelerators, takes significant manual effort. There are translation layers, such as AMD's HIP, which comes with a tool for importing CUDA code, but each is one more thing to keep in step with the original.

**Waiting on memory.** Over two decades, peak server compute grew about 3.0 times every two years, while memory bandwidth grew 1.6 times and interconnect bandwidth 1.4 times. Memory has overtaken compute as the main bottleneck in AI, especially when serving models. Making each operation faster on its own does little if the numbers can't get to it in time.

### What Lemurian Labs is after

Lemurian's complaint is that hardware-specific kernels have left AI infrastructure fragmented and brittle, and locked to particular hardware, so porting to another vendor or cloud is costly or close to impossible. In its words, "kernels have become the new assembly language."

The goal is a software stack, called Tachyon, that runs an AI workload on any hardware at any scale while matching or beating hand-tuned kernels. Write the code once, then treat one chip or a whole cluster as a single machine.

### How they go at it

**Start from parallel code.** Ordinary compilers try to dig parallelism out of step-by-step code. Tachyon starts from a pythonic tensor language and an intermediate representation that spells out concurrency and data flow from the start, so the compiler can reason about fusion and splitting work directly instead of guessing.

**Plan around data movement.** Kernel libraries make single operations fast. Tachyon's graph optimiser looks across operations instead, restructuring the computation to reuse data, move less of it and cut synchronisation. It uses the machine's memory sizes, bandwidth and topology to split work across one device or a thousand.

**Hardware as a parameter.** The code generator takes hardware characteristics as inputs, so the same pipeline can target NVIDIA, AMD, Intel and others. The compiler reasons about each chip explicitly, down to vector instruction shapes and cache behaviour.

**A runtime that adapts.** Big systems don't run on a fixed schedule; latencies and execution times vary. Tachyon's runtime manages data movement itself, overlaps communication with computation, stages data before it is needed and adjusts as conditions change.

### Still open

Can a compiler keep up with hand-tuning? There is real progress. TVM has shown performance competitive with hand-tuned libraries on low-power CPUs, mobile GPUs and server GPUs, and XLA aims to match ops that were fused by hand. But new model ideas keep arriving, and a clever hand-written kernel usually shows up first.

Which layer becomes the common ground? MLIR was built to tackle software fragmentation and cut the cost of building compilers for new hardware. XLA was designed so a new backend can be slotted in for novel chips. Many projects are trying to be the portable layer at once, and which one ends up shared across the industry is still open.

### Words used here

- **ML compiler**: Software that turns a model's graph of operations into efficient instructions for a specific chip.
- **lowering**: Translating a program step by step from a high-level description into ever more hardware-specific code.
- **kernels**: Small, heavily optimised programs that perform one operation, such as a matrix multiply, on a particular chip.
- **CUDA**: Nvidia's proprietary platform and programming interface for running general-purpose code on its GPUs.
- **LLVM**: A widely used open-source toolkit that compilers use to optimise code and generate machine instructions.
- **memory bandwidth**: How much data per second can move between a chip and its memory.
- **intermediate representation**: The internal form a compiler translates a program into so it can analyse and optimise it.

## About Lemurian Labs

Lemurian Labs is building the software foundation for the autonomy economy, solving the single biggest constraint on AI progress by creating a universal software layer that makes every chip, in every configuration, in every cloud, at every scale, usable as if it were one machine.

The company addresses a critical problem: AI is trapped behind walls of proprietary hardware and fragmented software. Lemurian's breakthrough lies in a first principles rethink of compilers and runtimes that makes heterogeneous systems feel like one seamless machine.

Developers write code once and it runs anywhere—faster, cheaper, and without lock-in—while simplifying workflows and delivering 2-30x performance gains.

## Sources

1. [XLA architecture](https://openxla.org/xla/architecture), OpenXLA
2. [CUDA](https://en.wikipedia.org/wiki/CUDA), Wikipedia
3. [TVM: An Automated End-to-End Optimizing Compiler for Deep Learning](https://arxiv.org/abs/1802.04799), arXiv
4. [FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness](https://arxiv.org/abs/2205.14135), arXiv
5. [AI and Memory Wall](https://arxiv.org/abs/2403.14123), arXiv
6. [MLIR: A Compiler Infrastructure for the End of Moore's Law](https://arxiv.org/abs/2002.11054), arXiv
7. [Tachyon by Lemurian Labs](https://www.lemurianlabs.com), Lemurian Labs
8. [Tachyon Technology](https://www.lemurianlabs.com/technology), Lemurian Labs
