← PortfolioMaking and talking

KREA

AI-first tools for professional creators

Team
Victor PerezCEO
Diego RodriguezCTO
Founded
2022
Invested
2023
Links
From our blog
The problem

How do you make an AI image model feel like a pencil?

Draw on the left with the mouse held down. The picture on the right keeps up, four steps at a time.A sketch on one side and the picture on the other, with no render button in between. Each new stroke sends the area around it back to noise, and it resolves again in four denoising steps while the rest of the picture stays put.An illustration, not real data.
How image models work

Most modern image generators are . In training, the network learns to undo the process of adding noise to a picture. To make a new one, it starts from pure static and runs the network over and over, taking a little noise away each time, like a sculptor chipping a statue out of TV snow.

A language model turns your prompt into a , which steers the denoising. To save compute, most models do all of this in a compressed rather than on raw pixels, with an autoencoder translating in and out.

Most leading video models denoise every frame of a clip at the same time, and each frame gets to look at all the others, including the ones that come after it.

Further reading Diffusion model (Wikipedia)Text-to-image model (Wikipedia)Krea Realtime 14B: Real-Time, Long-Form AI Video Generation (Krea)

Why it is hard
  1. i.

    One pass at a time

    Each denoising step is a full run of the network, and it has to wait for the one before. Good samples used to take hundreds or thousands of model evaluations. chips away at this: a student model learns to match its teacher in half the steps, and doing that again and again gets you down to as few as 4 steps without losing much quality. Consistency models go further and map noise straight to an image, with extra steps on offer if you want some quality back.

  2. ii.

    Words are a blunt tool

    Try describing a complicated layout, a pose or a shape in a text prompt and you'll see the problem. Getting close to the picture in your head usually means editing the prompt, looking at the result and editing again. If every round means waiting for a render, that loop gets old fast.

  3. iii.

    The AI look

    You know it when you see it: overly blurry backgrounds, waxy skin, boring composition. Part of the blame may lie with the data. LAION Aesthetics, a scorer commonly used to pick training images, favours blurry backgrounds, soft textures and bright images, and that bias ends up baked into the model. Tuning on existing preference datasets has side effects too, including sliding back toward the same look.

  4. iv.

    Video that can't stream

    If you denoise a whole clip at once, you can't show any of it until all of it is done, so real-time streaming is out. The alternative is generation: finish a frame, show it, make the next. The catch is that the model now builds on frames it made itself instead of real footage. That mismatch is called , and small flaws can snowball.

Further reading Progressive Distillation for Fast Sampling of Diffusion Models (arXiv)Consistency Models (arXiv)Adding Conditional Control to Text-to-Image Diffusion Models (arXiv)Releasing Open Weights for FLUX.1 Krea (Krea)Krea Realtime 14B: Real-Time, Long-Form AI Video Generation (Krea)

What KREA is after

KREA is going after two gaps. The first is taste. The goal it set for its first image model was refreshingly blunt: "Make AI images that don't look AI." With Krea 2 it wants style to be something you can guide, mix, turn up and turn down, rather than a vague word in a prompt.

The second is speed. The bet is that creative AI tools only feel steerable and responsive if they work in real time. It wants a loop with no queue and no render button.

Further reading Releasing Open Weights for FLUX.1 Krea (Krea)What Is Krea 2? Krea's Foundation Image Model (Krea)Krea Realtime 14B: Real-Time, Long-Form AI Video Generation (Krea)Krea Realtime: live AI canvas for images and video (Krea)

How they go at it
  1. Step 1: Start from a raw base

    For FLUX.1 Krea, KREA started from flux-dev-raw, a pre-trained 12B-parameter diffusion transformer from Black Forest Labs that is free of the AI aesthetic. Most of a model's look, the thinking goes, is learned in . So it fine-tuned on hand-picked images, then ran several rounds of preference optimisation on data collected with a clear, opinionated art direction.

  2. Step 2: Teach video to stream

    Krea Realtime 14B is distilled from Wan 2.1 14B using Self-Forcing, a way of turning video diffusion models into autoregressive ones. It trains the model on its own earlier outputs, the way it will actually run, which goes straight at exposure bias. One stage of training cuts sampling from about 30 steps to 4. It runs at 11 frames per second, 4 steps a frame, on a single NVIDIA B200 GPU.

  3. Step 3: Make style a dial

    Krea 2 is KREA's first foundation image model built from scratch. It carries the style of a reference image into the output, lets you choose how strongly, and can blend styles together. In the Realtime canvas, rough strokes, simple shapes and blocks of colour are enough to steer what comes out.

Further reading Releasing Open Weights for FLUX.1 Krea (Krea)Krea Realtime 14B: Real-Time, Long-Form AI Video Generation (Krea)What Is Krea 2? Krea's Foundation Image Model (Krea)Krea Realtime: live AI canvas for images and video (Krea)

Still open
  • How long can a streamed video hold together?

    An autoregressive model feeds its own mistakes back in as context, and this gets worse the longer the context grows. The Self Forcing authors report quality that matches slower non-causal models, but a model that's fine on short clips can still fall apart on the endless generation a live tool needs.

  • How do you measure a good image?

    Common benchmarks mostly check whether the model did what the prompt said: spatial relationships, object counts and the like. Older aesthetic scorers aren't good enough for today's models, and preference may be too personal to squeeze into one number anyway. How one model, or one metric, serves many tastes is still open.

Further reading Krea Realtime 14B: Real-Time, Long-Form AI Video Generation (Krea)Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion (arXiv)Releasing Open Weights for FLUX.1 Krea (Krea)

About KREA

KREA has been spearheading AI-centric tools since before the ChatGPT revolution. KREA has specialized in rapidly transferring new research results to tasteful, practically useful visual tools. Their ICP — visual professionals in architecture, advertising, and game development — find KREA an irreplaceable part of the modern toolchain: focused on flows where the pixels are pushed around by models instead of mouse movements, with the professional polish and integration they expect of a state-of-the-art enterprise creative tool.

KREA is used by visual professionals in architecture, advertising, and game development who need a creative tool built around models — not one that bolts AI on as an afterthought.

Words used here
diffusion models
Generative models that make an image by starting from random noise and removing it step by step.
latent space
A compressed numerical version of an image that a model works in instead of raw pixels.
text embedding
A list of numbers that a language model produces to represent the meaning of a prompt.
Distillation
Training a faster student model to reproduce what a slower teacher model does.
autoregressive
Generating output one piece at a time, with each new piece built on the ones already made.
exposure bias
The gap that arises when a model trained on real data must, in use, build on its own imperfect output.
post-training
Further training after the broad first stage, used to steer a model toward the outputs its makers want.
Sources