About
TwelveLabs builds enterprise video-native multimodal AI models and APIs that search, analyze, and understand video across vision, audio, and language. It sells to teams in media, sports, advertising, government, security, and other video-intensive industries, differentiating through foundation models designed specifically for video understanding.
Market
Twelve Labs competes in the video-native multimodal AI and video intelligence market, positioning itself as a provider of foundation models and infrastructure that enable machines to interpret visual, audio, and spoken information in video. Its differentiation is a video-specific focus spanning semantic search, summarization, event and object identification, and insight extraction through APIs and multimodal models, rather than general-purpose AI alone.
Twelve Labs primarily targets video-intensive organizations in media and entertainment, security, enterprise knowledge management, and related industries that need programmatic video search, summarization, event detection, and insight extraction. The likely buyers are engineering, AI/platform, and product teams; the available evidence does not specify a particular company-size range.
At a Glance
Problem
Video is a high-value but poorly indexed information source: organizations store it in archives, camera systems, broadcasts, meetings, factories, stadiums, and other repositories, yet typically access it through filenames, folders, captions, transcripts, or human memory. That makes finding a precise moment expensive and slow, while leaving valuable footage underused for analysis, reuse, personalization, and monetization. The pain is especially acute when teams must search across large libraries at scene level rather than merely locate a file.
The clearest killer use case is sports and media production. MLSE reports that Twelve Labs reduced video search and retrieval from 16 hours to 9 minutes, enabling faster highlight creation and more personalized fan content. Similarly, SBS used the technology to search for specific scenes and reuse archived visual-effects footage, replacing manual deep-dives into large volumes of video data with semantic retrieval.
Product / Service
Twelve Labs is an enterprise video-intelligence platform delivered through REST APIs and Python and Node.js SDKs. Customers upload videos and use the platform to search, analyze, generate embeddings, or reason across a knowledge store. Its multimodal models combine visual content, audio, spoken words, and on-screen text, allowing users to search with natural-language or image queries, identify moments and interactions, summarize or analyze a video, answer questions, and feed video embeddings into their own machine-learning systems.
The product is offered as usage-based APIs and platform capabilities including indexing and search, embedding, and analysis. Marengo provides video representations for retrieval and machine-learning workflows, while Pegasus supports video analysis and generation-oriented tasks. The benefit is that customers can turn an archive into a machine-readable, time-addressable knowledge base without building and maintaining separate models for text, image, and audio analysis; Twelve Labs also exposes a pricing calculator and a free-start path for API usage.
Market
Twelve Labs competes in enterprise video intelligence, multimodal video foundation models, and AI infrastructure for making video searchable and usable by applications and agents. Its differentiation is native semantic and temporal understanding across vision, audio, language, and their relationships, rather than relying only on keywords, metadata, or transcripts. The competitive set includes hyperscaler services such as Google Cloud Video Intelligence, Microsoft Azure AI Video Indexer, and Amazon Rekognition Video, which provide automated video recognition, indexing, or stored and streaming video analysis.
The company has meaningful commercial and financing traction rather than appearing to be a pre-revenue concept. Public customer evidence includes MLSE and SBS deployments, and Twelve Labs says its technology is being used across media, entertainment, sports, and advertising workflows. By July 2026 it had announced a $100 million round co-led by NEA and NAVER Ventures, followed by $30 million in strategic investments from Databricks, Snowflake, SK Telecom, HubSpot Ventures, and In-Q-Tel; the available evidence does not disclose revenue, so adoption, customer outcomes, partnerships, and funding are the clearest public traction signals.
Founders & Leadership
Funding History
Index Ventures
Radical Ventures
NVentures, Intel Capital, Samsung NEXT
New Enterprise Associates (NEA), NVentures
Databricks, Snowflake Ventures, SK Telecom, HubSpot Ventures, IQT
New Enterprise Associates (NEA), NAVER Ventures
Recent News
TwelveLabs raised $100 million in Series B funding, co-led by NEA and NAVER Ventures. The company will use the funding to scale its Video Cognition System and advance its Video Superintelligence vision.
TwelveLabs announced Rodeo, its first application-layer product, an AI-powered creative copilot that lets creators find, edit, and assemble video footage using natural language.
VAST Data and TwelveLabs announced a partnership enabling a customer-managed deployment path for TwelveLabs’ video foundation models on the VAST AI Operating System. The collaboration supports video search, analytics, and understanding across on-premises, cloud, and other controlled environments.
TwelveLabs announced Marengo 3.0, a multimodal embedding model for video retrieval. It supports composed queries, multilingual search, and long-form video understanding.
TwelveLabs integrated its Marengo and Pegasus models into Frame.io V4 through Custom Actions. The integration enables creative teams to semantically search video libraries within their workflows.
TwelveLabs introduced its enterprise video AI Partner Program, offering technical enablement, commercial incentives, and go-to-market support for solution partners.
TwelveLabs described an integration of Marengo with Amazon S3 Vectors for cross-modal search, semantic retrieval, and scalable AWS video workflows.
TwelveLabs presented an end-to-end video AI pipeline using Marengo and Pegasus on Amazon Bedrock. The workflow supports video search, analysis, and metadata generation.
Active Roles
17Business Model
TwelveLabs offers free access, usage-based Developer pricing tied to the amount of video processed, and enterprise contracts. Revenue comes from APIs and platform capabilities for video indexing, search, embedding, and analysis.
Products
Customers
Tech Stack
Similar Companies
Competitors
Key Investors
Databricks, Databricks Ventures, Firstman Studios, HubSpot Ventures, InnoWhale Ventures