← PortfolioMaking and talking

ArtCraft

AI-powered voice synthesis for video creators

Team
Brandon ThomasCEO
Founded
2022
Invested
2023
The problem

How do you direct an AI film instead of just describing it?

Move over the set on the left to place the camera, and click to take a shot. Every shot is the same room and the same pose.A small film set, from above and through the camera. The backdrop, props, a foreground pillar and a posed mannequin are laid out in 3D first, and moving the camera reframes the shot, with the layers sliding past each other in depth. Every shot is taken in the same room with the same pose, and only the final rendering would be handed to a model.An illustration, not real data.
How you steer a generator

Image and video generators mostly take orders in words. You type a , the model renders something that fits it, and if it isn't what you meant you type again. Words are fine for mood. For precise layouts, poses and shapes they are clumsy, and getting close to the picture in your head can take many rounds of trial and error.

Researchers have opened a second channel: pictures as instructions. ControlNet adds spatial controls to large pretrained text-to-image , so a model can follow an edge map, a human pose skeleton, a segmentation map or a depth map alongside the prompt. Draw a stick figure with its arm raised and you get a person with their arm raised.

Film has solved a version of this problem before. is how directors plan camera angles, staging and effects before shooting, from hand-drawn storyboards to digital 3D simulations. It lets them try out camera placement and movement without incurring actual production costs. , a term borrowed from theatre, is the precise arrangement of actors in the scene.

Further reading Adding Conditional Control to Text-to-Image Diffusion Models (arXiv)Previsualization (Wikipedia)Blocking (stage) (Wikipedia)

Why it is hard
  1. i.

    Words get misread

    Video models may misread a prompt and drift from what you meant, partly because they struggle to capture the context around the words. Fine details come out garbled too, the familiar distorted hands and unreadable text. Every retry costs another render.

  2. ii.

    Same face, next shot

    A story needs the same character from one shot to the next, and diffusion models find that hard, especially for subjects with complex details. Users of these models struggle to generate consistent characters. The usual fixes rely on several existing pictures of the character or on labour-intensive manual work.

  3. iii.

    One photo, many angles

    To move a camera around an object you need its shape, and a single photo doesn't show the back. Turning one image into new viewpoints is an under-constrained problem, so models fill the gaps with geometric priors learned from huge numbers of natural images. That's an educated guess, and it can guess wrong.

  4. iv.

    Long clips are expensive

    Text-to-video models are very computationally heavy, which limits how long and how good their outputs can be. Quality tends to decline as a generated video gets longer. So a film gets made from short clips that have to match each other.

Further reading Text-to-video model (Wikipedia)StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video Generation (arXiv)The Chosen One: Consistent Characters in Text-to-Image Diffusion Models (arXiv)Zero-1-to-3: Zero-shot One Image to 3D Object (arXiv)

What ArtCraft is after

ArtCraft's pitch fits in a line from its homepage: "Text prompting is neat, but artists crave control." The goal is to let an artist compose the shot, with real depth and a set camera, and hand only the final rendering to an AI model.

It also goes after how these tools are sold. ArtCraft is free and open source, with no subscription and no aggregator sitting between the artist and the models.

Further reading ArtCraft: Controllable AI for Artists (ArtCraft)

How they go at it
  1. Step 1: Build the set first

    Backdrops, foreground elements and props are layered in 3D to make compositions with depth. An image can be turned into a 3D object to position, rotate and frame, and 3D asset kits can be combined, which modellers call , to control camera angles, object placement and depth.

  2. Step 2: Pose, then shoot

    Characters get posed and the camera gets set before anything is generated. A posed mannequin can guide where a character stands and how. Image to Location puts virtual actors into one consistent environment, so several shots can be filmed in the same room "without things disappearing."

  3. Step 3: Bring any model

    The studio is a front end for many models, a catalog of 62 covering images, video, music and sound, 3D meshes and whole worlds, some of them built from . It runs on macOS, Windows and the web, and the whole studio is on GitHub.

Further reading ArtCraft: The IDE for artists (GitHub)Kitbashing (Wikipedia)ArtCraft: Controllable AI for Artists (ArtCraft)

Still open
  • Can a character stay the same across a whole film?

    Some methods now boost consistency across a series of generated images without retraining, and extend it to longer video with smooth transitions and consistent subjects. Others build a consistent identity from nothing but a text prompt, trading off how closely each image follows the prompt against how consistent the character stays.

  • How much 3D can you get from flat pictures?

    Gaussian splatting turns several images of a scene into a 3D representation that can be rendered from new angles. From a single image, diffusion-based methods now beat older single-view reconstruction, helped by Internet-scale pre-training, but the problem stays under-constrained, and what hides behind the object is still a guess.

Further reading StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video Generation (arXiv)The Chosen One: Consistent Characters in Text-to-Image Diffusion Models (arXiv)Gaussian splatting (Wikipedia)Zero-1-to-3: Zero-shot One Image to 3D Object (arXiv)

About ArtCraft

ArtCraft (formerly Storyteller) provides AI-powered storytelling tools for video content creation. Founded in 2022 and based in Atlanta, the platform enables users to create synthetic voices and audio with features including text-to-speech, voice-to-voice conversion, and voice designer capabilities.

The company's technology combines text-to-speech and deep fake technology to transform speech into unique characters, making it easy for creators to generate custom voice content for storytelling and video production.

Words used here
prompt
The text description a user types to tell a generative model what to make.
conditioning
Extra input, such as a sketch, pose or depth map, that a model must follow while generating.
diffusion models
Generative models that make an image by starting from random noise and removing it step by step.
Previsualization
Planning shots, staging and camera moves before filming, with storyboards or rough 3D scenes.
Blocking
Deciding exactly where each actor stands and moves within a scene.
kitbashing
Building a new model by combining parts taken from existing model kits.
Gaussian splats
A way of representing a 3D scene as many soft, coloured blobs that can be rendered quickly from any angle.
Sources