This podcast discusses a major shift in AI: pre-training scaling is hitting limits, so the industry is moving to 'test-time compute' (models think on the fly). This democratizes AI, letting small teams use Meta's open-source Llama to rival top labs cheaply. Key names: Meta (its Llama model is now the standard for small teams), OpenAI (strong brand but faces free competitors), and Anthropic (great tech but stuck in a tough market spot). The takeaway: AI investing is becoming more rational, less about betting billions on a single model.
This episode of Invest Like the Best features Benchmark partner Chetan Puttagunta and anonymous public market investor Modest Proposal, exploring a critical inflection point in AI development: leading labs are encountering scaling limits and shifting from pre-training to test-time compute. The core
Here is the English translation of the provided Chinese investment research notes, following all specified rules.
This episode features Benchmark partner Chetan Puttagunta and anonymous public market investor Modest Proposal. They discuss a critical turning point in AI development: leading labs are hitting scaling limits in pre-training and are shifting toward test-time compute. The core thesis is that this shift will democratize AI development, enabling small teams to approach frontier model performance at extremely low costs, and fundamentally reshape the public and private market investment landscape, moving AI investing from a capital-intensive "bet on God" model to a more predictable, revenue-aligned rational model.
Chetan Puttagunta believes all major AI labs have hit the bottleneck of pre-training scaling. Over the past two years, model performance improvements relied on a power law: "invest 10x compute for a step-function performance leap." However, this path faces a fundamental challenge: human-generated text data is nearly exhausted, and synthetic data has not driven continued pre-training scaling as expected. This is forcing the industry toward a new paradigm: "test-time compute."
Modest Proposal points out that this shift is significant for public market investors because it better aligns capital expenditure with revenue generation. In the pre-training era, companies had to invest tens of billions of dollars upfront to build compute clusters, wait 9-12 months to train a model, and then hope to generate revenue through inference. Under the test-time compute paradigm, spending is tied directly to actual model usage, which is financially more efficient and scalable. Furthermore, this will force a redesign of the entire network architecture: no longer requiring million-chip superclusters, but rather distributed, low-latency, high-efficiency inference data centers. He believes the public market has yet to digest the potential impact of this change on the investment narrative for power, grids, and related infrastructure.
Chetan Puttagunta observes that during the window of slowing pre-training scaling, small teams are catching up to the frontier at an unprecedented pace. In the past six weeks, Benchmark has engaged with multiple teams of just 2-5 people. By leveraging Meta's open-source model Llama, they have achieved frontier-level performance in specific domains (e.g., code generation) at costs orders of magnitude lower than large labs. This is enabled by the standardization of the Llama ecosystem and the proliferation of efficient algorithmic techniques like distillation and fine-tuning. This "capital-light, rapid technological breakthrough" model is precisely what venture capital favors and signals the return of investment opportunities at the model layer.
Modest Proposal adds that open-source models are a massive tailwind for cloud providers like AWS. If the industry's focus shifts from building a "universal god-like" pre-training model to solving real customer problems at inference time, AWS's traditional cloud service model of "providing tools for developers" becomes incredibly powerful. AWS doesn't need to own a top-tier LLM; it just needs to offer the most reliable, cost-effective API service to be a winner. This fundamentally changes the entire vision for technology deployment.
Modest Proposal provides a sharp analysis of the three frontier model companies:
Chetan Puttagunta adds his perspective on Google. Google has DeepMind's deep expertise in self-play, along with strong consumer and cloud businesses. Theoretically, it has "all the winning pieces." However, the core question is: even if it wins in AI, can it replicate the most successful business model in history – its internet search? This is the classic innovator's dilemma.
Chetan Puttagunta emphasizes that the stabilization of the model layer is a huge boon for the application layer explosion. Application developers were previously hesitant to invest because models underwent step-function changes annually. Now, they can be confident that under the inference paradigm, their "last mile" delivery technologies (e.g., data layers, UI, specific algorithms) will have lasting value. Since November 2022, Benchmark has invested in 25 AI companies, the vast majority of which are application-layer companies, at a pace comparable to the App Store era in 2009 and the internet era in 1995. These AI-native applications are disrupting the enterprise software market with "10x ROI compared to traditional SaaS" and "30-minute decision cycles."
Modest Proposal believes revenue at the application layer is becoming real, alleviating market concerns about AI investment returns. Inference costs are plummeting (100-200x), and GPU utilization is soaring. These factors point to a healthy economic model. He specifically notes that if pre-training spending decreases, capital will shift to inference, making the entire AI ecosystem's economics "much more reasonable." He warns that the market's excessive focus on NVIDIA may overlook other downstream hardware and infrastructure companies, whose probability-weighted return distributions have fundamentally changed.
| Position | Analyst Stance | Key Data |
|---|---|---|
| Meta (Llama) | Bullish on its open-source strategy shaping the ecosystem | Llama 3.1/3.2 series models; Llama 4 training cluster >100,000 H100s |
| OpenAI | Bullish on brand & distribution, but faces "outrun free" challenge | ChatGPT consumer mindshare; AGI IP agreement with Microsoft |
| Anthropic | Risk warning, believes it is "stuck in the middle" | Sonnet 3.5 considered best model; recent $4B funding round |
| xAI | Neutral, differentiation unclear | Plans to build a 200,000 chip cluster |
| Bullish on assets, but questions ability to replicate search business model | DeepMind's self-play capability; cloud business at scale | |
| AWS | Bullish on its advantage under the new paradigm | Largest cloud provider; viewed as a "logistics business" |
| Microsoft | Not explicitly stated | IP agreement with OpenAI is a key variable |
| NVIDIA | Neutral, believes market is overly focused | Referred to as the "largest remaining beneficiary" |
| Cerebras | Bullish on inference performance | Inference speed of 900+ tokens/s on Llama 3.1 405B, 70-75x faster than GPUs |
| DeepSeek | Mentioned as a case study of a small team catching the frontier | No specific data provided |
1. "Pre-training scaling has hit a data wall; the industry is shifting to test-time compute." (Chetan Puttagunta) – Human text data is exhausted, synthetic data hasn't driven continued scaling; the new paradigm uses inference time as the scaling axis.
2. "Test-time compute aligns capital expenditure with revenue generation, making the economic model of AI investment more rational." (Modest Proposal) – Shifts from betting hundreds of billions upfront on a model's success to paying based on actual usage, eliminating the "bet on God" risk.
3. "Small teams can now match frontier model performance in specific domains for under a million dollars." (Chetan Puttagunta) – Open-source models like Llama and efficient algorithmic techniques allow 2-5 person teams to disrupt the model layer with "capital-light, rapid technology."
4. "Anthropic is 'stuck in the middle': lacking consumer mindshare and squeezed by the open-source model Llama on the enterprise side." (Modest Proposal) – Despite top-tier technology, it lacks a viable market strategy, and its funding size hints at a troubled pre-training path.
5. "OpenAI's core challenge is 'can it outrun free?' – Meta and Google are likely to offer free ChatGPT alternatives." (Modest Proposal) – Strong brand and distribution are advantages, but facing free competitors with billions of user touchpoints makes the moat questionable.
6. "Stabilization of the model layer is a boon for the application layer; developers can finally confidently invest in 'last mile' technology." (Chetan Puttagunta) – Previously hesitant due to annual step-function model changes, they can now be confident that investments under the inference paradigm have lasting value.
7. "AI-native applications are disrupting enterprise software with '10x ROI' and '30-minute decision cycles,' compressing sales cycles from months to days." (Chetan Puttagunta) – Application-layer companies are in high demand, experiencing a boom similar to the 2009 App Store and 1995 internet eras.
8. "If pre-training scaling slows, network architecture needs redesign: from million-chip superclusters to low-latency, high-efficiency distributed inference data centers." (Modest Proposal) – This has profound implications for the investment narrative around power, grids, optical networks, etc., but the market has yet to price it in.