← Back to list
Lex Fridman PodcastPodcast4 Jan 2021Source: lexfridman.comHost: Lex Fridman

#151 – Dan Kokotov: Speech Recognition with AI and Humans

In plain words

This episode covers Rev, a speech-to-text company that uses an 'AI first draft + human polish' model. The founder argues its real edge isn't the AI algorithm, but a data flywheel: every customer payment generates high-quality training data, making the AI smarter. For market view, they're bullish on Rev's hybrid approach beating pure AI (14% error vs. 2-3% human) and big competitors in complex scenarios. Key holdings: Rev (data flywheel moat), Google (criticized for poor creator support), Amazon (Mechanical Turk has terrible UI and small team).

AI SummaryAI-generated · may contain errors · verify against the original

At a Glance

Dan Kokotov, VP of Engineering at Rev.ai, in a conversation with Lex Fridman, deeply analyzed Rev's business model, the current state of AI voice recognition technology, and the future of human-machine collaboration. Core judgment: Rev achieves an optimal balance of accuracy, cost, and scale through a hybrid model of "AI first draft + human correction." Its core moat lies not only in AI algorithms but also in the data flywheel created by its unique business model.

Theme 1: From a Chaotic "Upwork" to a Standardized "Rev" — Business Model Innovation

Dan Kokotov argues that the foundation of Rev's success lies in transforming a chaotic, non-standardized service (voice transcription) into a scalable, one-click product through standardization and a "friction-free" experience.

  • Historical Context and Mechanism Breakdown: Rev was born out of optimizing the Upwork model. On Upwork, buyers must choose from many freelancers, facing decision fatigue and delivery risk; freelancers also spend much time optimizing their profiles. Rev's solution: select verticals that can be standardized, such as "transcription" and "translation," eliminate all selection steps, and allow users to simply provide a file and receive a standardized format result within a fixed timeframe (e.g., 48 hours, later shortened). This "hide all details" experience is what Lex Fridman described as "easy and delightful."
  • Data Support: Lex Fridman's personal experience with Rev was "holy shit, somebody figured out how to do it just really easily." Its pricing model is based on audio duration (approx. $1.25/minute), not task complexity, which itself is a form of standardization.
  • Inference and Challenges: The key to this model is balancing supply and demand in a "two-sided market." If there are too few transcriptionists (Revvers) on the platform, customer wait times are too long; if there are too many, they may not get enough work. Dan notes that "maintaining this balance is a very challenging problem that we've been refining methods for over the years" — this is the core operational difficulty.

Theme 2: Current State of AI Voice Recognition and the "Data Flywheel" — Rev's Core Moat

Dan Kokotov judges that Rev's AI engine (Rev.ai) is already world-leading in general voice recognition, but its advantage comes not from unique algorithms, but from the "best data" and "data flywheel" fostered by its unique business model.

  • Data Chain and Mechanism Breakdown: Dan points out that in the ASR field, "the biggest thing is data. The more data you have, the higher the quality, the more accurate the labels, the better the results." Rev's business model itself is a "paid data labeling" loop: customers pay, Rev's human transcriptionists (Revvers) produce high-quality "gold standard" transcriptions, which are then used to train the AI model. This creates a "magical data flywheel."
  • Current Performance: Its internal test shows a Word Error Rate (WER) of about 14%, far above the human level of 2-3%, indicating a significant gap between human and AI.
  • Competitive Landscape: Dan confidently states that on his own test set (including podcasts, videos, interviews, lectures, etc.), Rev's ASR engine can beat Google, Amazon, Microsoft, and other giants. He also notes that these giants excel in voice assistants (e.g., Siri), but for the unstructured, long-form dialogue scenarios that Rev focuses on, their advantage is not obvious.
  • Inference and Falsification Signals: Dan believes Rev's "data flywheel" is still in its early stages. They not only have the final gold-standard labeled data, but also all the "editing process" data generated when transcriptionists edit the AI's first draft (e.g., which words were changed, how long it took). He believes this data contains richer signals that can be used to train the next generation of models. This will be a key path to improving AI accuracy in the future. Falsification signal: If competitors can acquire equally high-quality, diverse voice data at lower cost, or if a major algorithmic breakthrough occurs, Rev's data advantage will be weakened.

Theme 3: From "Tool" to "Platform" — Rev.ai's Future Vision and Product Ecosystem

Dan Kokotov describes Rev.ai's ultimate vision: not just a transcription tool, but a platform that makes "all conversations as searchable and indexable as notes."

  • Current Product Matrix: Rev has built a complete product line from consumer to developer, from human to AI.
  • Rev.com (Human + AI): For end users, offering high-quality human transcription, priced at $1.25/minute.
  • Temi (Pure AI): For price-sensitive consumers, pure AI transcription, priced at $0.25/minute, with its own editor.
  • Rev.ai (API): For developers, providing the AI engine as an API service, allowing developers to "build anything with it."
  • Mechanism and Inference: Dan compares Rev.ai to AWS (Amazon Web Services), hoping to provide developers with a "speech-to-text" computing module, thereby enabling new application scenarios. He specifically mentions one application: indexing and archiving all company meeting content, making it as easy to search and share as notes. He believes this will fundamentally change the way we work, especially with the increasing prevalence of remote work. Lex Fridman strongly resonated with this and pointed out that this is exactly what the podcast industry urgently needs — making audio content searchable and quotable.
  • Risks and Uncertainties: Dan acknowledges that the platform's success depends on the creativity of the developer community. They need to "see what people build with it, and then learn and try to make it easier to build those applications." This is inherently an uncertain process.

Mentioned Investments

Target Analyst Stance Key Data
Rev (Rev.com / Rev.ai) Bullish (Detailed explanation of its business model, data flywheel, and future vision as a core business) Pricing: $1.25/min (human+AI); $0.25/min (pure AI); AI engine WER 14%; World-leading ASR engine
Google (YouTube) Risk Warning (Compared, its auto-captions, API docs, and creator support are all judged inferior to Rev's) Beaten by Rev in internal tests, and operationally "doesn't care if creators succeed"
Amazon (Mechanical Turk) Risk Warning (Compared, its UI and API are extremely poor, and the operations team is very small) Said its "interface is terrible" and questioned its extremely small team size
Spotify Neutral (Evaluated its exclusive deal with Joe Rogan and expressed hope it will index podcast content like Rev) Not explicitly stated, but expressed hope that it could convert podcast content into text

Key Takeaways

1. Rev's business model is "paid data labeling" (Dan Kokotov): Its core moat is that every time a customer pays, they contribute a high-quality training data point for Rev, forming a unique "data flywheel." This is the fundamental reason for the continuous improvement of its AI capabilities.

2. AI voice recognition still lags significantly behind human levels, but the hybrid model is the optimal solution today (Dan Kokotov): Rev's AI engine has a WER of about 14%, while human experts can achieve 2-3%. The "AI first draft + human correction" model achieves the best balance between cost and accuracy, making it the best path for scalable service.

3. Rev's "editing process" data is an untapped gold mine (Dan Kokotov): The company not only has the final text, but also all operational data from transcriptionists editing the AI's first draft (e.g., time, edits). This data contains richer signals than the final result and can be used to train a smarter next-generation AI.

4. "Eliminating friction" is key to the success of service-oriented products (Dan Kokotov): Rev's success lies in simplifying Upwork's complex selection process into a standardized "drag-and-drop" experience. "Making users not care about the behind-the-scenes details is true ease of use."

5. Good management is "teaching according to aptitude" (Dan Kokotov): He cites the management book First, Break All the Rules, emphasizing that "management should be based on exceptions." That is, there is no one-size-fits-all management template; feedback must be tailored to each employee's personality (some need criticism, others need encouragement).

6. Long-term success should be measured not by short-term "engagement," but by users' "long-term well-being" (Dan Kokotov): He believes that even from a purely commercial standpoint, driving growth by optimizing users' long-term "emotional health" rather than short-term "anger/engagement" is a more sustainable and profitable model. This is a direct critique of the current social media platform business model centered on "pursuing engagement."

~9 min full read
Deep Analysis