← Back to list
Colossus (Invest Like the Best / Business Breakdowns)Podcast7 Jan 2025Source: joincolossus.comHost: Patrick O'Shaughnessy

Gaurav Misra & Dwight Churchill - Building Captions - [Invest Like the Best, EP.405]

In plain words

This piece is about the founders of Captions, an AI video generation company. They argue that making videos is a 'bounded problem' (with a clear end goal) compared to general AI, which is an endless race. They believe video generation can reach Hollywood quality in 18 months and become a permanent asset. Key holdings: Captions (generates hundreds of thousands of videos daily, founder bullish), TikTok/ByteDance (aggressive competitor who copied Captions), and L'Oreal (using GPT internally, a positive enterprise AI example).

AI SummaryAI-generated · may contain errors · verify against the original

This episode of Invest Like the Best invites Captions co-founders Dwight Churchill and Gaurav Misra to explore the key distinction between AI in "bounded problems" (e.g., video generation) and "unbounded problems" (e.g., general intelligence), as well as how to build a sustainable business. The core

~11 min full read · 7 sections
Deep Analysis

At a Glance

Dwight Churchill and Gaurav Misra are co-founders of AI video generation company Captions. This episode focuses on the fundamental impact of "bounded problems" (e.g., video generation) versus "unbounded problems" (e.g., general intelligence) on the business models of AI companies. Core thesis: Gaurav Misra argues that video generation is a "bounded problem" that can reach Hollywood quality within 18 months, and its business model is far superior to pursuing AGI, an "unbounded problem"—because the former has a clear endpoint, becoming a permanent asset once solved, rather than a never-ending capital race.

Bounded vs. Unbounded: The Foundational Classification That Determines AI Company Fortunes

Gaurav Misra believes AI companies should be divided into two categories: those solving "bounded problems" and those solving "unbounded problems," each with fundamentally different business models.

  • Unbounded Problems (AGI Front): These companies (e.g., GPT-like model companies) attempt to solve the unresolved puzzle of "intelligence," which has no ceiling—humans range from brilliant to average, and machines may surpass the smartest humans. The result is an endless capital race: each generation of models is replaced by the next, and the previous one is essentially written off. "We don't know when this race will end; it may never end."
  • Bounded Problems (Media Generation Front): Video, audio, and music generation are essentially "rendering a solved problem." CGI already exists; humans can create any imaginable visual—just at extremely high cost. AI's role is to make an already feasible solution "100 times cheaper, more accessible, and with a larger market." Such problems have a clear endpoint: when video generation reaches "near perfection," the model becomes a permanent asset. "Once it exists, it will continue to generate value without easily depreciating."
  • Key Implication: The training cost for bounded-problem models is high (tens of billions of dollars), but marginal costs decline; in contrast, capital expenditure on the AGI track has no end. Dwight Churchill adds: "Investors are overly focused on the massive capital expenditure on the AGI track, overlooking the business transformation happening in the bounded-problem space, which is much closer to the traditional software business model."

Captions Data Flywheel: From Free Product to Model Monopoly

Gaurav Misra describes how Captions built a data moat that competitors cannot easily replicate, using a strategy of "product as data collector."

  • Cold Start: The product initially launched on the App Store as a "built over a weekend" speech-to-text plus subtitle tool. The next day, without any promotion, it topped the app store charts—600 videos per minute were being generated. Key design: From day one, the app had a built-in data collection mechanism, paving the way for future model training.
  • Flywheel Structure: A free traditional video editing tool (recording + editing) attracts a massive user base, collects a large amount of "talking video" data (people speaking to the camera), which feeds back into the AI video generation model. Model improvements attract more users, creating a closed loop. Dwight emphasizes: "Other companies can only scrape data from the internet; we naturally acquire data through user growth, and the data is directly relevant to video generation models—this gives us a significant advantage."
  • Comparison with Traditional Data Giants: This resembles the Facebook/Google model—a free consumer product collects data to support a paid B2B product. Gaurav notes: "Simply downloading all our training videos would cost $1 million, a completely different magnitude from the cost of text data."
  • Competitive Barriers: Dwight reveals that TikTok/ByteDance has repeatedly attempted to "kill" Captions, including copying its App Store description, brand colors, website copy, and even "verbatim copying in press releases." But "the software they eventually delivered was mediocre, relying on TikTok's distribution channels, while we win with a better product."

Technology Path and Timeline for Video Generation

Gaurav Misra provides a concrete timeline: video generation can achieve object interaction within 6 months and Hollywood quality within 18 months, with inference costs declining at a 10x efficiency rate.

  • Technical Nature: Captions trains diffusion models, not GPT-style next-token prediction. Diffusion models start from "static noise" and, at each step, extract a layer of clarity from the noise conditioned on text (e.g., "a man in a blue shirt"), progressively approaching the target image. Current video diffusion models still have 10–30 billion parameters, far below the 400 billion parameters of text models, indicating significant room for growth.
  • Key Timeline:
  • Within 6 months: Achieve "person-object interaction" (e.g., a specified person holding a specific brand of water bottle), combining image conditioning (providing a photo of the bottle) with text conditioning.
  • Within 18 months: Generated videos may become "almost indistinguishable" from real recordings. Gaurav emphasizes: "This is not the most optimistic estimate, but a reasonable expectation. People have not fully grasped the speed of this technological change."
  • Cost Reduction Path: Dwight Churchill points out that inference costs are dropping at a rate of more than 10x per year. GPU prices have historically been deflationary (H100 → H200 architecture iterations), and model distillation can compress 100 diffusion steps into a few. Gaurav adds: "We are currently in the most inefficient phase; future efficiency will improve by at least an order of magnitude."

Investor Perspective: AI's Investment Value Lies Not in AGI, but in the Mature Business Models of "Bounded Problems"

Dwight Churchill believes investors have a structural bias in understanding AI—overly focused on the capital race for AGI, while ignoring the predictable business transformation happening in the "bounded problem" space.

  • Lessons from Traditional Software Companies: Dwight notes that the 80–90% gross margins of traditional SaaS companies like Salesforce are precisely the best attack targets for "bounded problem" companies. "High margins are an attack point for startups in the early stage and a signal that the moat is beginning to erode in the later stage." For the AI space, he advises investors to focus on the gross margin evolution of "bounded problem" companies—they will ultimately return to the high-margin model of traditional software.
  • Exploring Pricing Power: Gaurav Misra observes that consumers' willingness to pay for AI products is significantly higher than for traditional software. Captions was for a long time a pure paid product (no free tier), yet users were still willing to pay; and "a $25/month subscription fee" is feasible, with some AI video generation companies even charging up to $2,000/month for consumer subscriptions. But Dwight warns: "Don't rush to link pricing with labor costs. CFOs naturally want to lower labor costs; AI will only put further pressure on them, while a subscription model may offer more sustainable value."
  • Data Flywheel vs. Model Race: Gaurav believes the ultimate winner depends on "who has the best model"—and the best models come from the best data flywheels. Captions' strategy: accumulate massive amounts of "talking video" data through a free product, train fully licensed models, and avoid the legal risks of "unauthorized data" that competitors face. Dwight adds: "We focus on the narrow but deep space of 'talking video.' This is the only company training a dedicated foundation model for this niche—the opportunity is that once you excel in this narrow field, you can expand to the entire video ecosystem."

Mentioned Positions

Position Guest Attitude Key Data
Captions Bullish (founder perspective, but emphasizes superiority of its business model) Daily video generation in the "hundreds of thousands"; AI product has covered 1–5% of potential use cases; consumer subscription price can reach $25/month
TikTok/ByteDance Risk alert (powerful but aggressive competitor) Repeatedly attempted to "kill" Captions, including copying its App Store description, brand colors, website copy
Snap Background mention (founder Gaurav's former employer, cited as a case of product innovation and competitive lessons) Once provided Gaurav with product design training, but its "private sharing" DNA failed to adapt to TikTok's public sharing trend
L'Oreal Positive case (as an example of enterprise AI adoption) Deployed GPT-like tools internally for employees to query any question

Judgments Worth Remembering

1. "Bounded vs. Unbounded Problems": Video generation is "rendering," not "intelligence." Gaurav Misra argues that the AGI race has no finish line, while video generation, once it reaches Hollywood quality, becomes a permanent asset—this is the core investment logic of a "bounded problem."

2. Video generation can reach Hollywood quality within 18 months. Gaurav Misra provides a concrete timeline: current video model parameters (about 10–30 billion) are far below text models (about 400 billion), but technology and capital are pouring in, reaching "almost indistinguishable" levels within 18 months. Falsification condition: If after 18 months video generation still has obvious artifacts or cannot handle person-object interaction, the prediction fails.

3. AI video's "product-market fit" already exists in B2C, not B2B. Gaurav Misra points out that Captions' free consumer product is actually a "data collector," whose core value is accumulating training data, not direct monetization. This contrasts sharply with the traditional SaaS "pay first, then use" model.

4. The data flywheel is the core of the long-term moat for AI video companies. Gaurav Misra emphasizes that simply downloading the training videos would cost $1 million. Captions naturally accumulates fully licensed data through a free product, while competitors can only scrape unauthorized data from the internet—this constitutes a structural advantage.

5. Pricing power for AI video may be higher than for traditional software. Gaurav Misra observes that users are willing to pay $25/month (or even up to $2,000/month) for AI video generation, far above the $7.99–$12.99/month range for traditional video editing software. Dwight Churchill adds: "Don't rush to link pricing with labor costs—CFOs always want to lower labor costs, AI will put further pressure on them, and the subscription model may offer more sustainable value."

6. Gross margins of "bounded problem" AI companies will return to traditional software levels. Dwight Churchill believes that model training costs (hundreds of billions of dollars) are finite and predictable, and inference costs are declining at a rate of more than 10x per year, so gross margins will eventually approach the 80–90% level of traditional SaaS. Risk alert: "High margins are an attack point for startups in the early stage and a signal that the moat is beginning to erode in the later stage."

7. The only real competitor in AI video is TikTok, not Meta or Google. Gaurav Misra observes that Facebook has turned to open-source models, becoming the "good guy," while TikTok/ByteDance is executing a "Copy, kill, destroy" strategy, including copying Captions' App Store description, brand colors, and website copy. But "the software they eventually delivered was mediocre—we win with a better product."

8. AI's "intelligence" is already a "bounded problem"—just a translation of programming languages. Gaurav Misra argues that code generation is essentially "translating English into a new programming language," not creating general intelligence. Scott from Cognition's "Dev AI" is an embodiment of this trend. Analogy: From punch cards → assembly language → C++ → Python → English, each step is an upgrade of "translation"—the next "programming language" is natural language.