Audio and Speech¶
Algan lets you synchronize animations directly with sound files and voice-over
narration. Each Scene maintains its own
AudioManager, audio tracks, and speech
source.
Audio Contexts¶
Use Audio to align an
animation block’s runtime to a sound file:
from algan import *
circle = Circle().spawn()
with Audio("music.wav"):
circle.rotate(360, OUT)
circle.scale(2)
Scene.save_video("music_scene.mp4")
This flips the usual animation workflow on its head: instead of manually calculating how many seconds each visual action should take, the audio clip determines the runtime, and all animations inside the block automatically scale to match.
Both Audio and Speech contexts take wait_at_end, a number of extra seconds to hold after the
clip finishes, so you can add longer pauses to parts of the narration.
Speech defaults it to 1 second; Audio defaults it to 0.
Audio accepts string file paths or MoviePy AudioFileClip objects (allowing you
to trim or preprocess audio beforehand):
from moviepy import AudioFileClip
clip = AudioFileClip("music.mp3").subclipped(10, 20)
with Audio(clip):
mob.move(RIGHT)
Speech Contexts¶
Speech is an
Audio context whose clip
is generated from a script segment, so the block runs for exactly as long as
the line takes to say:
from algan import *
title = Text("Gradient descent").spawn()
with Speech("Gradient descent follows the slope downhill."):
title.move(UP)
title.color = BLUE
Scene.save_video("gradient_descent.mp4")
Important
The default speech generator synthesizes through pyttsx3, which drives a
system text-to-speech engine rather than shipping one. macOS and Windows
have one built in (NSSpeechSynthesizer and SAPI5), so Speech works out of
the box there. On Linux (including CI containers and most Docker images)
pyttsx3 falls back to eSpeak, which is not installed by default and
is not a Python dependency Algan can pull in for you. Without it, a
Speech context raises at synthesis time. You
can install it with:
sudo apt install espeak-ng # Debian / Ubuntu
sudo dnf install espeak-ng # Fedora
If you would rather not depend on a system engine at all, supply your own generator (see A custom speech generator) or use recorded narration.
By default, the Scene’s AudioManager uses Algan’s pyttsx3 speech generator. Each
Speech context appends its script to scene.audio_manager.video_transcript.
When a video contains audio, save_video writes that transcript beside the
video as <video_stem>_script.txt.
Using recorded narration¶
For recorded narration, configure the specific Scene’s AudioManager rather than a process-global singleton:
from algan import *
from algan.utils.audio_utils import get_speech_generator_from_file
generator = get_speech_generator_from_file(
audio_file="narration.wav",
transcript_file="narration.txt",
)
Scene.current().audio_manager.set_speech_source(generator)
diagram = Circle().spawn()
with Speech("First we draw a circle."):
diagram.scale(1.5)
Scene.save_video("narrated_diagram.mp4")
get_speech_generator_from_file aligns the transcript to the audio and
returns a callable. Each Speech segment asks that callable for the matching
subclip. The optional audio dependencies used for alignment are available via
Algan’s audio extra.
A custom speech generator¶
A speech generator is any callable accepting a script string and returning a MoviePy audio clip:
from moviepy import AudioFileClip
def speech_generator(script):
# Select or synthesize a clip for this exact script segment.
return AudioFileClip("prepared_segment.wav")
Scene.current().audio_manager.set_speech_source(speech_generator)
The generator is Scene-local. Two Scenes can use different voices or recorded sources in the same process without interfering with one another.
Composing narration and sound effects¶
Audio contexts nest like other animation contexts. For example, a sound effect can run in parallel with a visual change inside a narration segment:
with Speech("The object now transforms into a triangle."):
with Sync():
with Audio("whoosh.wav"):
pass
mob.become(Triangle(add_to_scene=False))
See Also¶
Combining Animations – the animation contexts
AudioandSpeechextend.Text and Mathematics – putting on screen what the narration is saying.
Multi-Scene Projects – installing one
speech_sourceacross every scene of a video, and where transcripts land.Saving Videos and Images – the
audio_codecandffmpeg_paramsarguments tosave_video.