← PortfolioMachines that move

OpenMind

Operating system and coordination layer for intelligent robots

Team
Jan Liphardtfounder
Founded
2024
Invested
2024
Links
The problem

How do you get robots from different makers to run the same brain?

Move inside the camera view to walk the person around the robot, and click there to have them say hello. Click a body at the bottom to plug it in: everything above the line stays the same.What the robot sees and hears is turned into short sentences, posted on a shared bus and fused into one paragraph. A fast model and a slower cloud planner answer from that paragraph at their own speeds, while a mentor writes notes for the planner. Underneath, the same decision goes to whichever body is plugged in, and each body’s plugin turns it into its own motor commands.An illustration, not real data.
How robot software stacks up

Underneath most robots sits . The best known is ROS, the Robot Operating System, which, despite the name, is a set of software frameworks that runs on top of one. It handles , low-level device control, message-passing between processes and package management. It was never a system, which matters for a machine that has to react quickly, and ROS 2 added support for real-time code.

Hardware abstraction is an old trick. Software talks to a family of similar devices through one set of commands, even though each device is wired differently underneath. A joystick is the textbook case: one interface for moving and firing, whoever made the stick.

Once robots leave the lab they also have to share space. Open-RMF, an open-source project, exists to let multiple fleets of robots work alongside each other and alongside building infrastructure like doors and elevators.

Further reading Robot Operating System (Wikipedia)Hardware abstraction (Wikipedia)Open-RMF (Open Source Robotics Alliance)

Why it is hard
  1. i.

    Every body is different

    A joystick has a few buttons. A robot dog and a humanoid have very different bodies, and each robot has its own capabilities, so the code that carries out an action is typically specific to the hardware. Learning has the same problem: the field has tended to train a separate model for every application, every robot and even every environment.

  2. ii.

    Thinking at two speeds

    Language models are good at working out what to do and slow at doing it. A small, fast model can answer in about 300 milliseconds; a large cloud model doing long-term planning takes around two seconds. A robot about to walk into a chair can't wait two seconds, so something has to split the fast reflexes from the slow thinking.

  3. iii.

    Robots without papers

    Robots can't get passports, it isn't clear which rules apply to them, and they can't use the ordinary banking system. Most are also "trapped in single-vendor ecosystems that limit collaboration", so when machines from different companies meet in a warehouse or a hospital corridor, one has little way to check who the other is.

Further reading Actions (OpenMind)Open X-Embodiment: Robotic Learning Datasets and RT-X Models (arXiv)Architecture (OpenMind)ERC-7777: Governance for Human Robot Societies (Ethereum Improvement Proposals)Connecting the bots: US firm builds tool to help humanoid robots work together (Interesting Engineering)

What OpenMind is after

OpenMind builds the software that makes robots useful. Its runtime, OM1, is meant to make it easy to build robots that are easy to upgrade and reconfigure for different physical form factors.

The idea is one AI persona that can run in the cloud or on physical hardware such as quadrupeds, TurtleBot 4 and humanoids, with a coordination layer on top so machines from any manufacturer can work together.

Further reading OM1 (OpenMind)OM1 (GitHub)Introduction (OpenMind)Connecting the bots: US firm builds tool to help humanoid robots work together (Interesting Engineering)

How they go at it
  1. Step 1: Senses into sentences

    OM1 turns sensor data into language. A vision language model, or , describes what the cameras see, speech recognition turns audio into text, and all of it lands on a . A fuser then rolls those snippets into one paragraph describing the robot's world, along the lines of "You see a human, 3.2 meters to your left."

  2. Step 2: A committee of models

    Decisions come from several models at once: a fast one for time-critical actions, a slower cloud model for reasoning and planning, and a mentor model that critiques the robot's interaction with people every 30 seconds and passes notes to the planner. The models themselves are swappable, from OpenAI and Gemini to local ones through Ollama.

  3. Step 3: Down to the motors

    A hardware abstraction layer turns a decision like "pick up the red apple with your left hand" into gripper servo commands, often through existing ROS2, CycloneDDS or Zenoh middleware. New robots join through hardware-specific plugins. For robots that need more onboard muscle, the BrainPack is a module that mounts on the robot and adds mapping, object recognition, remote control and self-charging.

  4. Step 4: Robots with ID

    FABRIC is the coordination network: it lets robots identify themselves, verify locations and share knowledge with machines they haven't met. Identity is built on an Ethereum standard OpenMind co-wrote, ERC-7777, which gives robots and lets a physical robot prove who it is with cryptographic signatures from secure hardware.

Further reading Architecture (OpenMind)Introduction (OpenMind)Actions (OpenMind)BrainPack (OpenMind)Connecting the bots: US firm builds tool to help humanoid robots work together (Interesting Engineering)ERC-7777: Governance for Human Robot Societies (Ethereum Improvement Proposals)

Still open
  • Where should language stop and motor control begin?

    Vision-language-action models go straight from an image and an instruction to low-level robot actions. Layered designs keep language upstream and hand off to movement policies and ROS2 code. Which scales better is an open argument.

  • Who writes the rules for mixed crowds of humans and robots?

    ERC-7777 proposes machine-readable charters that humans and robots can join, update and leave, but the standard is still being peer-reviewed, and nobody yet knows what a robot's identity will be accepted for in practice.

Further reading Vision-language-action model (Wikipedia)Architecture (OpenMind)ERC-7777: Governance for Human Robot Societies (Ethereum Improvement Proposals)

About OpenMind

OpenMind is building the system software that allows robots to operate autonomously in the real world. The company is developing a unified coordination layer that integrates perception, language models, task planning, and low-level control, enabling robots to translate high-level goals into reliable physical action.

As foundation models improve, the bottleneck in robotics has shifted from intelligence to coordination. Robots can now perceive environments and understand tasks, but lack a secure, interoperable system for managing long-horizon behavior across heterogeneous hardware. OpenMind addresses this gap by providing the control and execution layer that legacy robotics stacks never delivered.

Founded by Jan Liphardt, a Stanford professor whose background spans physics, biophysics, machine learning, and distributed systems, OpenMind launched commercially in early 2024 and is building toward a world where any robot can be reliably orchestrated at scale.

Words used here
middleware
Software that sits between a robot's hardware and its applications, passing messages and hiding low-level details.
hardware abstraction
A layer that lets software use different devices through the same commands.
real-time
Guaranteed to respond within a fixed deadline, which control loops need.
VLM
A vision language model: a language model that can also take images as input.
Natural Language Data Bus
OM1's shared channel where every sensor's output is posted as short text descriptions.
on-chain identities
Identities recorded on a public blockchain, which anyone can check without asking the maker.
Sources