# ArtCraft

AI-powered voice synthesis for video creators

- Team: Brandon Thomas (CEO)
- Founded: 2022
- Invested: 2023
- Links: [Website](https://artcraft.ai), [LinkedIn](https://www.linkedin.com/in/possibilistic), [GitHub](https://github.com/echelon)
- Field: Making and talking

## The problem: How do you direct an AI film instead of just describing it?

### How you steer a generator

Image and video generators mostly take orders in words. You type a prompt, the model renders something that fits it, and if it isn't what you meant you type again. Words are fine for mood. For precise layouts, poses and shapes they are clumsy, and getting close to the picture in your head can take many rounds of trial and error.

Researchers have opened a second channel: pictures as instructions. ControlNet adds spatial conditioning controls to large pretrained text-to-image diffusion models, so a model can follow an edge map, a human pose skeleton, a segmentation map or a depth map alongside the prompt. Draw a stick figure with its arm raised and you get a person with their arm raised.

Film has solved a version of this problem before. Previsualization is how directors plan camera angles, staging and effects before shooting, from hand-drawn storyboards to digital 3D simulations. It lets them try out camera placement and movement without incurring actual production costs. Blocking, a term borrowed from theatre, is the precise arrangement of actors in the scene.

### Why it is hard

**Words get misread.** Video models may misread a prompt and drift from what you meant, partly because they struggle to capture the context around the words. Fine details come out garbled too, the familiar distorted hands and unreadable text. Every retry costs another render.

**Same face, next shot.** A story needs the same character from one shot to the next, and diffusion models find that hard, especially for subjects with complex details. Users of these models struggle to generate consistent characters. The usual fixes rely on several existing pictures of the character or on labour-intensive manual work.

**One photo, many angles.** To move a camera around an object you need its shape, and a single photo doesn't show the back. Turning one image into new viewpoints is an under-constrained problem, so models fill the gaps with geometric priors learned from huge numbers of natural images. That's an educated guess, and it can guess wrong.

**Long clips are expensive.** Text-to-video models are very computationally heavy, which limits how long and how good their outputs can be. Quality tends to decline as a generated video gets longer. So a film gets made from short clips that have to match each other.

### What ArtCraft is after

ArtCraft's pitch fits in a line from its homepage: "Text prompting is neat, but artists crave control." The goal is to let an artist compose the shot, with real depth and a set camera, and hand only the final rendering to an AI model.

It also goes after how these tools are sold. ArtCraft is free and open source, with no subscription and no aggregator sitting between the artist and the models.

### How they go at it

**Build the set first.** Backdrops, foreground elements and props are layered in 3D to make compositions with depth. An image can be turned into a 3D object to position, rotate and frame, and 3D asset kits can be combined, which modellers call kitbashing, to control camera angles, object placement and depth.

**Pose, then shoot.** Characters get posed and the camera gets set before anything is generated. A posed mannequin can guide where a character stands and how. Image to Location puts virtual actors into one consistent environment, so several shots can be filmed in the same room "without things disappearing."

**Bring any model.** The studio is a front end for many models, a catalog of 62 covering images, video, music and sound, 3D meshes and whole worlds, some of them built from Gaussian splats. It runs on macOS, Windows and the web, and the whole studio is on GitHub.

### Still open

Can a character stay the same across a whole film? Some methods now boost consistency across a series of generated images without retraining, and extend it to longer video with smooth transitions and consistent subjects. Others build a consistent identity from nothing but a text prompt, trading off how closely each image follows the prompt against how consistent the character stays.

How much 3D can you get from flat pictures? Gaussian splatting turns several images of a scene into a 3D representation that can be rendered from new angles. From a single image, diffusion-based methods now beat older single-view reconstruction, helped by Internet-scale pre-training, but the problem stays under-constrained, and what hides behind the object is still a guess.

### Words used here

- **prompt**: The text description a user types to tell a generative model what to make.
- **conditioning**: Extra input, such as a sketch, pose or depth map, that a model must follow while generating.
- **diffusion models**: Generative models that make an image by starting from random noise and removing it step by step.
- **Previsualization**: Planning shots, staging and camera moves before filming, with storyboards or rough 3D scenes.
- **Blocking**: Deciding exactly where each actor stands and moves within a scene.
- **kitbashing**: Building a new model by combining parts taken from existing model kits.
- **Gaussian splats**: A way of representing a 3D scene as many soft, coloured blobs that can be rendered quickly from any angle.

## About ArtCraft

ArtCraft (formerly Storyteller) provides AI-powered storytelling tools for video content creation. Founded in 2022 and based in Atlanta, the platform enables users to create synthetic voices and audio with features including text-to-speech, voice-to-voice conversion, and voice designer capabilities.

The company's technology combines text-to-speech and deep fake technology to transform speech into unique characters, making it easy for creators to generate custom voice content for storytelling and video production.

## Sources

1. [Adding Conditional Control to Text-to-Image Diffusion Models](https://arxiv.org/html/2302.05543v3), arXiv
2. [Previsualization](https://en.wikipedia.org/wiki/Previsualization), Wikipedia
3. [Blocking (stage)](https://en.wikipedia.org/wiki/Blocking_(stage)), Wikipedia
4. [Text-to-video model](https://en.wikipedia.org/wiki/Text-to-video_model), Wikipedia
5. [StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video Generation](https://arxiv.org/abs/2405.01434), arXiv
6. [The Chosen One: Consistent Characters in Text-to-Image Diffusion Models](https://arxiv.org/abs/2311.10093), arXiv
7. [Zero-1-to-3: Zero-shot One Image to 3D Object](https://arxiv.org/abs/2303.11328), arXiv
8. [ArtCraft: Controllable AI for Artists](https://artcraft.ai), ArtCraft
9. [ArtCraft: The IDE for artists](https://github.com/storytold/artcraft), GitHub
10. [Kitbashing](https://en.wikipedia.org/wiki/Kitbashing), Wikipedia
11. [Gaussian splatting](https://en.wikipedia.org/wiki/Gaussian_splatting), Wikipedia
