Customers/AI visual dubbing company
OverviewAll customer stories
Anonymized customer · fictional logo

A leader in AI visual dubbing for film industrializes dataset collection

Training visual-dubbing models for cinema takes immense, well-structured datasets of talking-face video. Tornado API delivered an academically referenced corpus of nearly one hundred thousand videos straight into their S3 bucket, reconciled, verified, in MKV.

AI visual dubbing company
Products used
Ingestion API
Managed Datasets
AI & ML
High volume
Delivers to AWS S3

How it started

They found Tornado through a Google search while looking for a way to build their training corpus, and reached out directly to the founder.

Challenge

Visual dubbing models learn from talking-face video, and they need it at research scale, in high quality, structured against academic references like SpeakerVid and TalkVid. Building a corpus of ~100,000 referenced videos in-house meant months of engineering on a problem entirely outside their core business.

Every week spent on scrapers, proxies and retry logic was a week not spent on the models themselves.

Solution

Tornado ran the delivery as a managed pipeline: ingestion of the referenced corpora, reconciliation of every file against the academic manifests, and direct delivery in MKV to their S3 bucket with completeness reporting along the way.

No proxies, no retry logic, no infra on their side, one manifest in, one bucket filling up.

Results

98,663 videos delivered, ~32.65 TB in MKV, every file reconciled against the reference manifests. The dataset team went back to model work the same week.

Since the backfill, incremental additions run through the same pipeline automatically.

We needed a 30 TB training corpus delivered to one bucket. It landed in days, organized, verified, with zero babysitting from our side.Dataset engineering lead, AI visual dubbing company
More customer stories:Taja AIVideo podcast hostContent repurposing platform