Media AI pipeline
AI video generation pipeline
A system that turns a brand brief into a finished video automatically, using several AI models behind the scenes and built so any of them can be swapped as better ones appear.
- Role
- Architect and developer
- Client
- BrandRizz
The problem
Generating brand videos programmatically means stitching together several generative models (video, voice, image) and a renderer, then running a job that can take minutes and fail halfway. The models themselves change every few months.
What I built, and what I deliberately didn't
Built:
- A provider-agnostic adapter layer so each stage of the pipeline (video generation, voice, post-processing) talks to a common interface. Veo, Kling, ElevenLabs and WaveSpeed sit behind it and can be swapped as the market moves, without touching the pipeline.
- Remotion rendering to compose generated clips, voice and brand elements into a finished video from React components.
- Durable workflow orchestration with Vercel Workflow, so long-running renders survive restarts, retry failed steps and resume where they left off.
Didn't build: a dependency on any single model vendor, or a bespoke job queue. The adapters absorb vendor churn and the workflow runtime handles durability.
For the technical readerUnder the hood: production details, architecture and stack
Architecture
- BriefBrand inputs
- AdaptersVeo · Kling · ElevenLabs · WaveSpeed
- RemotionComposition and render
- WorkflowDurable steps, retries
Stack
Next.js, TypeScript, Remotion, Vercel Workflow, ElevenLabs, Veo, Kling, WaveSpeed.