ArtCraft
AI-powered voice synthesis for video creators
How do you direct an AI film instead of just describing it?
Image and video generators mostly take orders in words. You type a The text description a user types to tell a generative model what to make., the model renders something that fits it, and if it isn't what you meant you type again. Words are fine for mood. For precise layouts, poses and shapes they are clumsy, and getting close to the picture in your head can take many rounds of trial and error.
Researchers have opened a second channel: pictures as instructions. ControlNet adds spatial Extra input, such as a sketch, pose or depth map, that a model must follow while generating. controls to large pretrained text-to-image Generative models that make an image by starting from random noise and removing it step by step., so a model can follow an edge map, a human pose skeleton, a segmentation map or a depth map alongside the prompt. Draw a stick figure with its arm raised and you get a person with their arm raised.
Film has solved a version of this problem before. Planning shots, staging and camera moves before filming, with storyboards or rough 3D scenes. is how directors plan camera angles, staging and effects before shooting, from hand-drawn storyboards to digital 3D simulations. It lets them try out camera placement and movement without incurring actual production costs. Deciding exactly where each actor stands and moves within a scene., a term borrowed from theatre, is the precise arrangement of actors in the scene.
Further reading Adding Conditional Control to Text-to-Image Diffusion Models (arXiv)Previsualization (Wikipedia)Blocking (stage) (Wikipedia)
- i.
Words get misread
Video models may misread a prompt and drift from what you meant, partly because they struggle to capture the context around the words. Fine details come out garbled too, the familiar distorted hands and unreadable text. Every retry costs another render.
- ii.
Same face, next shot
A story needs the same character from one shot to the next, and diffusion models find that hard, especially for subjects with complex details. Users of these models struggle to generate consistent characters. The usual fixes rely on several existing pictures of the character or on labour-intensive manual work.
- iii.
One photo, many angles
To move a camera around an object you need its shape, and a single photo doesn't show the back. Turning one image into new viewpoints is an under-constrained problem, so models fill the gaps with geometric priors learned from huge numbers of natural images. That's an educated guess, and it can guess wrong.
- iv.
Long clips are expensive
Text-to-video models are very computationally heavy, which limits how long and how good their outputs can be. Quality tends to decline as a generated video gets longer. So a film gets made from short clips that have to match each other.
Further reading Text-to-video model (Wikipedia)StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video Generation (arXiv)The Chosen One: Consistent Characters in Text-to-Image Diffusion Models (arXiv)Zero-1-to-3: Zero-shot One Image to 3D Object (arXiv)
ArtCraft's pitch fits in a line from its homepage: "Text prompting is neat, but artists crave control." The goal is to let an artist compose the shot, with real depth and a set camera, and hand only the final rendering to an AI model.
It also goes after how these tools are sold. ArtCraft is free and open source, with no subscription and no aggregator sitting between the artist and the models.
Further reading ArtCraft: Controllable AI for Artists (ArtCraft)
- Step 1: Build the set first
Backdrops, foreground elements and props are layered in 3D to make compositions with depth. An image can be turned into a 3D object to position, rotate and frame, and 3D asset kits can be combined, which modellers call Building a new model by combining parts taken from existing model kits., to control camera angles, object placement and depth.
- Step 2: Pose, then shoot
Characters get posed and the camera gets set before anything is generated. A posed mannequin can guide where a character stands and how. Image to Location puts virtual actors into one consistent environment, so several shots can be filmed in the same room "without things disappearing."
- Step 3: Bring any model
The studio is a front end for many models, a catalog of 62 covering images, video, music and sound, 3D meshes and whole worlds, some of them built from A way of representing a 3D scene as many soft, coloured blobs that can be rendered quickly from any angle.. It runs on macOS, Windows and the web, and the whole studio is on GitHub.
Further reading ArtCraft: The IDE for artists (GitHub)Kitbashing (Wikipedia)ArtCraft: Controllable AI for Artists (ArtCraft)
Can a character stay the same across a whole film?
Some methods now boost consistency across a series of generated images without retraining, and extend it to longer video with smooth transitions and consistent subjects. Others build a consistent identity from nothing but a text prompt, trading off how closely each image follows the prompt against how consistent the character stays.
How much 3D can you get from flat pictures?
Gaussian splatting turns several images of a scene into a 3D representation that can be rendered from new angles. From a single image, diffusion-based methods now beat older single-view reconstruction, helped by Internet-scale pre-training, but the problem stays under-constrained, and what hides behind the object is still a guess.
Further reading StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video Generation (arXiv)The Chosen One: Consistent Characters in Text-to-Image Diffusion Models (arXiv)Gaussian splatting (Wikipedia)Zero-1-to-3: Zero-shot One Image to 3D Object (arXiv)
ArtCraft (formerly Storyteller) provides AI-powered storytelling tools for video content creation. Founded in 2022 and based in Atlanta, the platform enables users to create synthetic voices and audio with features including text-to-speech, voice-to-voice conversion, and voice designer capabilities.
The company's technology combines text-to-speech and deep fake technology to transform speech into unique characters, making it easy for creators to generate custom voice content for storytelling and video production.
- prompt
- The text description a user types to tell a generative model what to make.
- conditioning
- Extra input, such as a sketch, pose or depth map, that a model must follow while generating.
- diffusion models
- Generative models that make an image by starting from random noise and removing it step by step.
- Previsualization
- Planning shots, staging and camera moves before filming, with storyboards or rough 3D scenes.
- Blocking
- Deciding exactly where each actor stands and moves within a scene.
- kitbashing
- Building a new model by combining parts taken from existing model kits.
- Gaussian splats
- A way of representing a 3D scene as many soft, coloured blobs that can be rendered quickly from any angle.
- 1Adding Conditional Control to Text-to-Image Diffusion Models · arXiv
- 2Previsualization · Wikipedia
- 3Blocking (stage) · Wikipedia
- 4Text-to-video model · Wikipedia
- 5StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video Generation · arXiv
- 6The Chosen One: Consistent Characters in Text-to-Image Diffusion Models · arXiv
- 7Zero-1-to-3: Zero-shot One Image to 3D Object · arXiv
- 8ArtCraft: Controllable AI for Artists · ArtCraft
- 9ArtCraft: The IDE for artists · GitHub
- 10Kitbashing · Wikipedia
- 11Gaussian splatting · Wikipedia