Company profile
TwelveLabs
Video intelligence platform and API for enterprises.
- Category
- Computer vision
- Headquarters
- San Francisco, California
- Sells to
- Enterprise
- Business model
- Usage-based API
- Deployment
- Cloud / SaaS, On-premise, Hybrid, API
- Pricing
- Not published
- Builds own models
- Yes
- Modalities
- Video, Audio, speech, Text, Image, Multimodal
What TwelveLabs does
TwelveLabs offers a video intelligence platform and API that transforms raw video into searchable, AI-ready data at massive scale. It provides infrastructure for ingesting multimodal data, indexing video, and enabling natural language search across entire video libraries. The platform includes foundation models like Marengo for multimodal embeddings and Pegasus for video language understanding, allowing for advanced search, content segmentation, compliance checks, highlight generation, and insight extraction. It supports various workflows for media, sports, advertising, government, and security sectors, with flexible deployment options and developer-friendly SDKs.
Products
- JockeyThe first video intelligence AI agent, a unified agentic system that reasons across videos and images, planning its own steps to answer complex questions with grounded, cited moments.
- MarengoA multimodal embedding model that turns video into data by generating spatiotemporal embeddings. It analyzes frames, temporal relationships, speech, and sound to enable cross-modal search and any-to-any retrieval tasks across text, audio, image, and video.
- PegasusA powerful video-first language model that integrates visual, audio, and speech information to understand video and generate accurate descriptions, analysis, and creative outputs in natural language.
- TwelveLabs Video Understanding PlatformAn intelligence layer for video that allows developers to search, analyze, generate embeddings, and reason across video content using APIs and SDKs. It supports deep semantic search, dynamic video-to-text generation, and intuitive integration.
Key capabilities
- Multimodal data ingestion through a single pipeline
- Fast indexing (60x real-time speed, 10k+ hours per day)
- Natural language search across video libraries
- Automatic content segmentation
- Compliance and brand safety checks with explainable AI
- Automated highlight generation
- Insight generation and pattern surfacing
- Video-native perception, reasoning, and orchestration
- Spatiotemporal embeddings for video data
- Continuous reasoning over full temporal arc of video assets
- Any-to-any search capabilities
- Rich embeddings for semantic search, hybrid search, anomaly detection
- Video-to-text generation for summaries, chapters, highlights, Q&A
- API and SDK access (Python, Node.js)
- Flexible deployment (on-premise, hybrid, cloud)
- Fine-tuning capabilities for models
- Image-to-video search
- One-time video indexing for multiple tasks
- Contextual vectors across visual, audio, spoken word, and on-screen text
- Domain-specific vocabulary understanding
- RAG pairing for relevant information retrieval
- High-quality training data generation
- Anomaly detection in video data
Use cases
- Search and discover specific actions, scenes, dialogue, and human emotions in video
- Automatically identify natural breaks, scene changes, and pacing shifts in long-form video
- Identify policy risks, sensitive content, and brand safety issues
- Generate rough cuts, thematic clips, and assemble material for editing workflows
- Analyze video at scale to surface patterns and signals for creative and editorial decisions
- Archive monetization by turning historical content into licensable inventory
- Scene detection and structural metadata generation for media workflows
- Creation of promos, trailers, and social clips
- Contextual ad placement and ad break optimization
- Brand safety and suitability verification for advertising
- Creative performance insights for advertising campaigns
- Multi-source evidence search across CCTV, drone, body-worn, and citizen-submitted footage
- Incident response and reporting for government and security
- Pattern detection and trend analysis in archived footage
- Semantic search for moments in video using text or image queries
- Video summarization and text generation (chapters, highlights, Q&A)
- Building custom classifiers using natural language
- Content discovery and recommendation systems
- Multilingual search
- Asset management for large video libraries
- Surveillance and investigation support
- Training models with high-quality video embeddings
- Workplace safety compliance
- Recursive video enhancement
- Influencer identification
- Social media post generation from videos
- Shade finding in videos
- Olympic video classification
- Interview analysis
- Brand sponsorship measurement
AI approach
TwelveLabs pioneers multimodal, video-native AI that sees and understands like humans do. They develop proprietary video foundation models, Marengo and Pegasus, which analyze visual, audio, speech, and temporal relationships in video to enable search, analysis, and text generation. They offer both dedicated APIs for individual tasks (Models) and an agentic system (Jockey) for corpus-level reasoning.
Tech named: Jockey (AI agent), Marengo (Multimodal Embedding Model), Pegasus (Video Language Model), LLMs, spatiotemporal embeddings, deep neural networks, multimodal foundation model, vector embeddings, RAG (Retrieval Augmented Generation)
Industries served
- Media & Entertainment
- Sports & Broadcasting
- Advertising and Marketing
- Public Sector
- Government
- Security
- Automotive
What it says sets it apart
- World's most powerful video intelligence platform
- Video-native AI that sees and understands like humans do
- Multimodal approach surpassing unimodal models (text or images only)
- Infrastructure for video intelligence at massive scale
- Proprietary foundation models (Marengo, Pegasus)
- Ability to reason continuously over the full temporal arc of video assets
- High composite accuracy (78.5% for Marengo, #1 on Video-MME for Pegasus)
- Faster content review and compliance scanning (10x)
- SOC 2 Type II certified and secure by design
- Flexible deployment options including air-gapped configurations for high-sensitivity workloads
- Simplified API integration for a rich set of video understanding tasks
- Natural language use for queries and prompts, more effective than rules or keywords
- Domain-specific understanding for vocabulary and jargon
- Efficient processing (180x run time indexing, 10,000 hours in less than an hour)
- Ability to deploy the entire intelligence stack where data lives
Funding rounds we track
New Enterprise Associates, NAVER Ventures, Amazon, NAVER, Radical Ventures, Korea Investment Partners, Index Ventures, Quadrille Capital, Red Bull Ventures
From the AI funding tracker — rounds as reported by the linked publications.
This profile was compiled from TwelveLabs's own public pages in July 2026 and reflects what the company states about itself — not an endorsement or an independent audit of those claims. Facts are extracted with AI and filtered by an automated check that drops any named product, customer or certification missing from the source pages. Full method. Something out of date? Tell us.