← PortfolioCompute and infrastructure

Lemurian Labs

Universal compiler and runtime for AI

Team
Jay DawaniCEO
Vassil Dimitrovco-founder
Chris VickVP
Founded
2018
Invested
2025
Links
From our blog
The problem

How do you write AI code once and run it fast on any chip?

Click a machine on the left to re-plan the same program for it. Hold the button down to run one operation at a time instead, and watch the trips to memory multiply.The same three-step program, planned again for each machine from what that machine is: how much fast memory each chip has, and how many chips there are and how they connect. The work is cut into tiles that fit, shared out across the chips, and each tile goes through all three steps before it leaves fast memory. Doing one step at a time instead would send each tile back and forth three times.An illustration, not real data.
How an ML compiler works

A model written in PyTorch or JAX is, underneath, a graph: matrix multiplies, additions and other operations wired together. Something has to turn that graph into instructions a particular chip can run. That something is an . XLA, a widely used one, takes model graphs from frameworks like PyTorch, TensorFlow and JAX and compiles them into machine instructions for various architectures.

The work happens in stages, a process called . First come passes that don't care about the hardware, such as fusing operations together and planning where intermediate results live in memory. Then a backend adds what it knows about the target chip, fuses further, and generates code, often through . It's a little like translating a recipe first into plain steps, then into instructions for one particular kitchen.

At the bottom sit , the small, tightly tuned programs that do one operation on one kind of chip. On NVIDIA GPUs those are usually written with , Nvidia's proprietary platform for running general-purpose code on its GPUs. Most frameworks lean on libraries of vendor-specific kernels.

Further reading XLA architecture (OpenXLA)CUDA (Wikipedia)TVM: An Automated End-to-End Optimizing Compiler for Deep Learning (arXiv)

Why it is hard
  1. i.

    Hand-tuned sets the bar

    Historically, compilers couldn't beat hand-written kernels, so people wrote kernels by hand, chip by chip. The gap can be large. FlashAttention, an attention algorithm designed around how data moves between a GPU's memory levels, trained GPT-2 about 3 times faster than existing baselines. Any compiler that wants to replace hand-tuning has to match work like that, automatically.

  2. ii.

    Code that won't travel

    Kernels are tied to the hardware they were written for. Frameworks optimise for a narrow range of server GPUs, and moving a workload to new platforms, from phones to custom accelerators, takes significant manual effort. There are translation layers, such as AMD's HIP, which comes with a tool for importing CUDA code, but each is one more thing to keep in step with the original.

  3. iii.

    Waiting on memory

    Over two decades, peak server compute grew about 3.0 times every two years, while grew 1.6 times and interconnect bandwidth 1.4 times. Memory has overtaken compute as the main bottleneck in AI, especially when serving models. Making each operation faster on its own does little if the numbers can't get to it in time.

Further reading Tachyon by Lemurian Labs (Lemurian Labs)FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness (arXiv)TVM: An Automated End-to-End Optimizing Compiler for Deep Learning (arXiv)CUDA (Wikipedia)AI and Memory Wall (arXiv)

What Lemurian Labs is after

Lemurian's complaint is that hardware-specific kernels have left AI infrastructure fragmented and brittle, and locked to particular hardware, so porting to another vendor or cloud is costly or close to impossible. In its words, "kernels have become the new assembly language."

The goal is a software stack, called Tachyon, that runs an AI workload on any hardware at any scale while matching or beating hand-tuned kernels. Write the code once, then treat one chip or a whole cluster as a single machine.

Further reading Tachyon by Lemurian Labs (Lemurian Labs)Tachyon Technology (Lemurian Labs)

How they go at it
  1. Step 1: Start from parallel code

    Ordinary compilers try to dig parallelism out of step-by-step code. Tachyon starts from a pythonic tensor language and an that spells out concurrency and data flow from the start, so the compiler can reason about fusion and splitting work directly instead of guessing.

  2. Step 2: Plan around data movement

    Kernel libraries make single operations fast. Tachyon's graph optimiser looks across operations instead, restructuring the computation to reuse data, move less of it and cut synchronisation. It uses the machine's memory sizes, bandwidth and topology to split work across one device or a thousand.

  3. Step 3: Hardware as a parameter

    The code generator takes hardware characteristics as inputs, so the same pipeline can target NVIDIA, AMD, Intel and others. The compiler reasons about each chip explicitly, down to vector instruction shapes and cache behaviour.

  4. Step 4: A runtime that adapts

    Big systems don't run on a fixed schedule; latencies and execution times vary. Tachyon's runtime manages data movement itself, overlaps communication with computation, stages data before it is needed and adjusts as conditions change.

Further reading Tachyon Technology (Lemurian Labs)

Still open
  • Can a compiler keep up with hand-tuning?

    There is real progress. TVM has shown performance competitive with hand-tuned libraries on low-power CPUs, mobile GPUs and server GPUs, and XLA aims to match ops that were fused by hand. But new model ideas keep arriving, and a clever hand-written kernel usually shows up first.

  • Which layer becomes the common ground?

    MLIR was built to tackle software fragmentation and cut the cost of building compilers for new hardware. XLA was designed so a new backend can be slotted in for novel chips. Many projects are trying to be the portable layer at once, and which one ends up shared across the industry is still open.

Further reading TVM: An Automated End-to-End Optimizing Compiler for Deep Learning (arXiv)XLA architecture (OpenXLA)MLIR: A Compiler Infrastructure for the End of Moore's Law (arXiv)

About Lemurian Labs

Lemurian Labs is building the software foundation for the autonomy economy, solving the single biggest constraint on AI progress by creating a universal software layer that makes every chip, in every configuration, in every cloud, at every scale, usable as if it were one machine.

The company addresses a critical problem: AI is trapped behind walls of proprietary hardware and fragmented software. Lemurian's breakthrough lies in a first principles rethink of compilers and runtimes that makes heterogeneous systems feel like one seamless machine.

Developers write code once and it runs anywhere—faster, cheaper, and without lock-in—while simplifying workflows and delivering 2-30x performance gains.

Words used here
ML compiler
Software that turns a model's graph of operations into efficient instructions for a specific chip.
lowering
Translating a program step by step from a high-level description into ever more hardware-specific code.
kernels
Small, heavily optimised programs that perform one operation, such as a matrix multiply, on a particular chip.
CUDA
Nvidia's proprietary platform and programming interface for running general-purpose code on its GPUs.
LLVM
A widely used open-source toolkit that compilers use to optimise code and generate machine instructions.
memory bandwidth
How much data per second can move between a chip and its memory.
intermediate representation
The internal form a compiler translates a program into so it can analyse and optimise it.
Sources