Twelve Labs
API that indexes video by what happens in it, not filenames
Overview
Twelve Labs builds foundation models (Marengo, Pegasus) that let developers search, classify and summarize video content itself — actions, speech, on-screen text, objects — via a search, analyze and embed API. Free tier covers 600 minutes; paid usage is metered per minute.
Pricing
Pricing shown for reference only. These figures reflect RECATOOLS research as of 12 Jul 2026 and may be out of date or incomplete. This is not financial or purchasing advice — always confirm the current price on the provider’s official website before making any decision.
- 600 combined minutes (index/analyze/segment)
- Up to 100 videos per index
- 90-day index retention
- Search, Embed and Analyze APIs included
- Indexing $0.042/min
- Search $4 per 1,000 queries
- Analyze $0.0292/min input + $0.0075/1k output tokens
- Up to 100,000 videos per index
- Negotiated volume pricing
- Private cloud / on-prem deployment
- Dedicated support
What you can produce with Twelve Labs
- Search API for finding specific moments across a video library
- Analyze API for prompted text generation from video content
- Embed API for multimodal video/audio/image/text vector embeddings
- Marengo model family for content-based indexing
- Pegasus model for structuring video into entities, scenes and time segments
- Cloud, private cloud or on-premise deployment
- Free tier: 600 minutes combined usage, no credit card required
- Official Python and JavaScript SDKs (open source)
ASEAN Perspective
Twelve Labs in Southeast Asia
ASEAN-region availability and pricing notes coming soon. Drop the editorial team a note via /contact/ if you can supply local context (Singapore/Malaysia/Indonesia/Thailand/Vietnam).
Twelve Labs is the sharpest dedicated video-understanding API on the market right now, and the money agrees: a $100M Series B closed July 1, 2026, co-led by NEA and Naver Ventures with Amazon, Radical Ventures and Korea Investment Partners joining, taking total funding to roughly $150M. Marengo handles multimodal embeddings, Pegasus turns footage into structured, LLM-readable data — scene boundaries, entities, timestamps — which is a real gap general-purpose vision LLMs still handle badly.
The free 600-minute tier makes evaluation cheap, and per-minute Developer pricing is transparent rather than sales-gated. Caveats: it's API-only with no consumer product, English-first docs, and USD billing, so ASEAN teams get no regional pricing or localization. A Glassdoor review titled "Sinking ship — avoid" surfaced around the same funding news, worth weighing against the investor list rather than taking at face value. Best fit: teams building video search, moderation or ad-verification products who don't want to train their own vision model.
What people say
Twelve Labs closed a $100M Series B on July 1, 2026, co-led by NEA and Naver Ventures with Amazon, Radical Ventures, Korea Investment Partners, Index Ventures, Quadrille Capital and Red Bull Ventures participating — the release frames it as funding "video superintelligence" and puts cumulative funding around $150M. The San Francisco/Seoul startup also expanded its AWS alliance in the same announcement.
Product-wise, the API is built on two model families: Marengo (multimodal embeddings across video, audio, text) and Pegasus (structures video into machine-readable segments — entities, actions, time codes — so downstream LLMs can reason over it). Three endpoints do the work: Search (find moments across a library), Analyze (generate text from prompts against a video), and Embed (vector output for semantic search or recommendation). It runs on cloud, private cloud or on-prem.
Pricing is public and metered rather than quote-only for the mid tier: Free gives 600 combined indexing/analyze/segment minutes, 90-day index retention, up to 100 videos per index. The paid Developer tier charges $0.042/minute to index, $4 per 1,000 search queries, $0.0292/minute for Analyze input plus $0.0075 per 1,000 output tokens, and $0.0015/minute for embedding infrastructure — unlimited video hours, 100,000 videos per index. Enterprise is custom, committed-use contracts.
Independent review coverage is thin: G2 lists it mainly via a competitor/alternatives page rather than a scored profile, and PeerSpot's product page explicitly has zero collected reviews as of this writing. Employee sentiment on Glassdoor is split — one review calls it "innovative and driven," another (posted around the funding announcement) is titled "Sinking ship — avoid." Given the fresh $100M and strategic investors like Amazon and Naver, that pessimistic take is worth noting but shouldn't be weighted the same as a funded, growing customer base. TechCrunch and Databricks have both profiled the technology positively for depth of video understanding versus generic multimodal LLMs.
Official SDKs (Python, JavaScript) are open source under the twelvelabs-io GitHub org and auto-generated from an OpenAPI spec — contributors are told not to bother submitting manual PRs to the SDK repos since codegen overwrites them each release.
Summary of public user & expert reviews, compiled by RECATOOLS.
About this listing
This entry was compiled from publicly available data including Twelve Labs's official website, press releases, documentation, and reputable third-party publications. RECATOOLS is not affiliated with Twelve Labs unless explicitly stated.
Third-party AI tools update their pricing, features, availability, and policies frequently. Information here may be outdated by the time you read this — we make reasonable efforts to keep listings current, but cannot guarantee absolute accuracy.
For the latest details, please refer to Twelve Labs directly →
Spotted something out of date? Suggest an update →
Alternatives to Twelve Labs
More in Video & Audio