KaiCalls needed vertical demo videos. Not one video — dozens. One for plumbers, one for electricians, one for HVAC techs, one for roofers. Same format, different industry hooks.
Building one video in After Effects takes hours. Building 20 takes weeks. The obvious answer: generate them programmatically. Write code once, render any industry.
Remotion made this possible. React components compile to MP4. Every element — text, animations, audio — lives in TypeScript. Change a config, get a new video.
The first version was terrible. Version 9 finally worked. Here's what broke between them.
The Stack
Remotion turns React into video. Components render at 30fps into frames, FFmpeg stitches them into MP4. Standard React patterns — props, state, hooks — all work.
ElevenLabs generates the voices. Three characters:
- Brian — Deep narrator voice. "The call that changes everything."
- Eric — The Kai AI agent. Calm, professional, solution-oriented.
- Sarah — The panicked caller. Pipe burst, power out, AC dead.
remotion-subtitle handles word-by-word captions with bounce animations. No manual keyframing.
The 9 Iterations
Each version fixed something the previous one broke. Here's the progression:
1 Audio desync
Narrator audio started before visuals. The voiceover said "pipe bursts" while the screen still showed the logo. Fixed by adding J-cuts — audio leads visual by 8 frames, not 60.
2 Caption overlap
Text rendered dead center, directly over the animated orb. Moved captions to bottom third with translateY(40%) offset.
3 Dead air gaps
Silence between voiceover segments. Added a continuous ambient bed at 0.15 volume — never stops, fills every gap.
4 Spring jitter
Animations bounced at the end. Looked cheap. Switched to overdamped springs: { stiffness: 150, damping: 20, mass: 0.6 }. No jiggle. Premium feel.
5 Corrupted audio files
One dialogue clip was 22 seconds instead of 4. ElevenLabs glitched during generation. The render stacked audio on top of itself. Added duration validation before render.
6 Caption/VO mismatch
Captions said "burst pipe" while audio said "flooding basement." The config file and the generated audio drifted apart. Fixed by regenerating all VO from the single source of truth: industries.ts.
7 Memory kills
Renders crashed at 85% completion. SIGKILL from the kernel. The 16GB VPS couldn't handle concurrent Chrome processes. Set --concurrency=1 and cleaned stale processes before each render.
8 Wrong dialogue order
Kai said the solution first. Then the caller said "Oh thank god." Backwards. Emergencies work differently: panic first, then relief. Swapped the sequence — caller states the problem, then Kai responds.
9 Timing polish
Final pass: tightened transitions, trimmed dead time from 60s to 30s, added Ken Burns effect on the closing logo (1.0 → 1.05 scale over 150 frames).
The dialogue order fix was the breakthrough. v1–v7 had Kai speak first. Felt robotic. Real emergency calls start with panic: "My basement is flooding!" Then the solution: "I'm dispatching someone now." That sequence change made the whole video click.
The Final Result
30 seconds. Vertical format (1080×1920). Industry-specific hook, dialogue, and data extractions. One render script, configurable per industry.
Compare that to an early iteration without the dialogue fix:
Config-Driven Industries
One system renders all industries. The industries.ts config defines everything that changes:
plumber: {
captions: {
hook: "Pipe bursts at 2am. By the time you call back,
they've already found someone.",
reveal: "You're not ignoring calls. You're under a house.
Or driving. Or asleep.",
proof: "It answers. Gets the details. Wakes you up
only if it's real."
},
dialogue: {
caller: "My basement is flooding! There's a pipe that
just burst — I need someone NOW.",
kai: "I'm dispatching someone right now. They'll be
there within the hour. Hang tight."
},
extractions: {
caller: "Sarah Mitchell",
issue: "Emergency — Burst pipe flooding basement",
scheduled: "ASAP — En route"
}
}
The render script reads the config, generates VO with ElevenLabs, and outputs the video. Adding a new industry takes 10 minutes — write the config, run the script.
Current industries: plumber, electrician, HVAC, roofer, and a generic home services fallback.
What I'd Do Differently
- Validate audio duration upfront. One corrupted file cost hours of debugging. Add a pre-render check.
- Test dialogue order early. The caller-first fix should have been iteration 2, not iteration 8.
- Render on a bigger machine. Memory pressure on a 16GB VPS caused random failures. Cloud render would be more reliable.
Try It Yourself
Remotion's documentation is solid: remotion.dev. The key insight: treat video like any other React app. Components, props, state. The timeline is just frame count.
ElevenLabs handles voice generation. Their eleven_multilingual_v2 model produces natural dialogue. Cost: roughly $0.30 per minute of audio.
Total cost per video: about $2 in API calls, 5 minutes of render time.
The pitch for programmatic video: You write the system once. Every variation after that is a config change. 50 industries? 50 configs. No timeline scrubbing. No export-wait-preview loops.
Nine iterations from broken to polished. Most of the fixes were obvious in retrospect — audio sync, memory limits, dialogue order. The work was finding them.