Loading knowledge network

Patman's Neural Network

AI Technology

Audio drives video. Not the other way around.

Lightricks and ElevenLabs bring audio-to-video generation. Sound determines what happens on screen.

Published on 20 January 2026

Translated from German

Video tools have a problem. They treat sound as an afterthought. Visuals first, then sound at some point. Lightricks is turning that around.

Audio-to-video is now live. With ElevenLabs as an exclusive launch partner. Broader access arrives on January 27.

This is not text-to-video with retrofitted audio. It begins with sound. Audio becomes the control layer. Voice, music, and sound effects dictate timing, motion, and performance from the very first frame. Not just later as decoration.

The structure of the video emerges from the audio itself. Speech rhythm determines pacing. Musical energy influences motion and camera behavior. Scene changes happen where the sound demands them. Not where a prompt guesses they might fit.

This is particularly powerful at the start of the creative process. When teams want to test ideas quickly. When they need a feel for how something lands before investing time into details.


The Problem with Previous Tools

For years, video tools have treated sound as a separate step. Even advanced generative systems bring sound into play late in the game. After scenes, shots, and movements have already been determined.

Want visuals that truly fit a voice or a piece of music? You have to translate sound into something else. Prompts, timestamps, camera notes, or subsequent cuts.

These workarounds have become so normal that we no longer question them. But they don't work well. Audio already contains intent. It carries timing, emphasis, rhythm, and emotion. When it isn't allowed to lead, videos feel less natural.

Audio-to-video starts from a simple idea. Stop translating sound. Let it steer generation directly.


Partnership with ElevenLabs

The launch is rolling out exclusively with ElevenLabs, a global AI audio provider, during the initial release phase.

ElevenLabs builds world-class audio that tells stories. LTX-2 turns it into video. Directly. The audio layer becomes the visual story.

Luke Harries, Growth at ElevenLabs

“Offering LTX's audio-to-video capabilities exclusively to our users gives our community access to incredible creativity. They can create professional videos quickly. We are thrilled about this partnership because we have always believed that AI should help creators overcome technical hurdles.”

Daniel Berkovitz, Chief Product Officer at Lightricks

“Accelerating the creative process with audio-to-video capabilities makes ElevenLabs a natural partner. Starting from sound gives creators precise control over pacing, performance, and structure. An approach that has long been used in animation and is now becoming accessible for all video creation.”


Built for Real Workflows

Audio-to-video is available starting January 20 in LTX and ElevenLabs Image & Video. API and open-source access follow on January 27.

Users provide an audio file. Voice, dialogue, music, or sound effects. That is the primary input. Optionally, an image can anchor a character or a scene. A short text prompt can steer the visual style. But audio remains in control.

The result is a single Full HD video clip whose length and motion are driven by the audio. For longer sequences, clips can be chained together. Teams can thus construct complete videos modularly without abandoning the audio-first approach.

This is infrastructure, not a demo. Built for platforms, developers, and studios that are building products and pipelines. Not just experimenting.


LTX as Gateway for Creators

Audio-to-video is part of the broader LTX ecosystem. Lightricks builds the underlying models and infrastructure. LTX is the access point for developers and creators.

LTX applies the same technology to real creative workflows. It helps shape how these tools are used in practice. Every layer informs the others. Research, platform, and production move together.

For Platforms

True audio-driven control enables new audio-first products. Without complex orchestration.

For Developers

It reduces the gap between intent and output.

For Creative Teams

It turns sound into something you can start with. Not something you fix later.

Sound Tells Stories

Sound has always done the heavy lifting of storytelling. Now it finally gets to lead.

Source: https://ltx.studio/blog/ltx-audio-to-video-generation-with-elevenlabs

Return to network