Capabilities
Metadata Enrichment
Tagging, speaker ID, object detection, categorization.

How it works
Process
Upload or select a video. Metadata enrichment runs alongside other tasks.
AI analysis
Speech, visual content, and audio are analyzed to extract structured metadata: speakers, objects, brands, topics, and emotions.
Review
Explore generated tags, speaker labels, object detections, and topic categories.
Export
Export metadata in standard formats or push directly to your CMS/MAM via API.
Key features
- Speaker identification — recognize and label individual speakers.
- Object detection — objects, locations, and visual elements.
- Brand recognition — logos and brand mentions.
- Semantic tagging — topic and category classification.
- Emotion and sentiment — tone and emotional context.
- Sound detection — automatic detection and classification of sounds in audio tracks.
- Ad-ready metadata — contextual signals for ad targeting.
- Available in Turbo mode — full metadata enrichment runs on both Standard and Turbo processing models.
FAQ
Yes. Since v1.44, Turbo processing includes full metadata enrichment — the same analysis as Standard mode, with faster turnaround.
Sound detection automatically identifies and classifies non-speech audio events (applause, music, sirens, etc.) in your content. Sound descriptions appear alongside the transcript and can be included in subtitle exports.

