This interview argues that the future of AI isn't chatbots but background 'agents' that run for hours, making cost more important than speed. Guest Neil Movva thinks the premium charged by labs like OpenAI for being 3-6 months ahead is unsustainable because open-source models catch up. His company, Sail, buys unwanted chips (like AMD) and power, even accepting occasional outages, to slash AI inference costs. Key mentions: NVIDIA (good short-term but not long-term due to slow performance-per-watt gains), AMD (undervalued, an arbitrage opportunity), and Cerebras (high bandwidth for storing model weights).
Neil Movva, founder of Sail, proposed on the Invest Like the Best podcast the construction of a "token factory" to reduce AI inference costs by 10x. His core argument is that future AI agents will run in the background for extended periods (hours to days), where latency becomes irrelevant and cost b
Here is the English translation of the provided Chinese investment research notes, following all specified rules.
Neil Movva, founder of Sail, proposes building a "token factory" to reduce AI inference costs by 10x. His core thesis is that future AI agents will run in the background for extended periods (hours to days), making latency irrelevant and cost the critical variable. Sail employs a "scavenger strategy" to acquire chips and power that others do not want, optimizing the throughput-latency trade-off in GPUs to minimize per-token cost. Neil Movva believes the premium charged by frontier labs for being 3-6 months ahead is unsustainable, as the diffusion and distillation of open-source model capabilities cannot be stopped.
Neil Movva argues that the future of AI inference lies in long-running background agents, not real-time interactive chatbots. He believes that when users no longer wait for an AI response but instead let it work autonomously in the background, latency becomes irrelevant, and cost becomes the only critical variable.
Neil Movva points out that a GPU is inherently a "throughput machine," but the current AI industry forces it to act like a "private car" in pursuit of low latency, sacrificing efficiency. Sail's strategy is to return the GPU to its "bus" nature, maximizing throughput.
Neil Movva's core strategy is "scavenging"—buying chips and power that no one else wants on the market, making them efficient through software optimization, and thus building an ultra-low-cost inference factory.
Neil Movva holds a contrarian view on NVIDIA: "bullish short-term, bearish long-term." He believes NVIDIA's position is solid in low-latency inference, but the future belongs to high-throughput inference, for which NVIDIA's architecture is not optimal.
Neil Movva believes that the internet, as a source of high-quality text data, is a one-time "subsidy," and future model capability improvements will primarily rely on self-reinforcement learning (RL) on verifiable tasks.
| Position | Analyst Sentiment | Key Data |
|---|---|---|
| NVIDIA | Bullish short-term, Bearish long-term | Performance per watt improvement from Hopper to Rubin is not significant; NVLink is a necessity for low latency but costly (8x hardware for 4-5x speed). |
| AMD | Bullish (arbitrage opportunity exists) | Market perception bias leads to undervaluation; software ecosystem lags NVIDIA, leaving room for optimization by Sail. |
| Cerebras | Bullish (as an accelerator) | Its Wafer Scale Engine 3 offers SRAM bandwidth up to 21 PB/s (NVIDIA HBM ~10 TB/s), suitable for storing model weights, but KV Cache capacity is limited. |
| Anthropic / OpenAI | Neutral (competitive relationship) | They pay a high "premium" for being 3-6 months ahead, but Neil believes this premium is unsustainable. |
| DeepSeek | Positive mention | Achieved "an order of magnitude" progress in KV Cache compression, with continued annual improvements. |
| Intel | Neutral (potential alternative) | In an extreme scenario without TSMC, its process is at most 2x behind in performance per watt, which is not unacceptable. |
1. "The best latency is no latency" (Neil Movva): When AI agents work autonomously in the background, users don't wait, and the latency problem disappears. This is the cornerstone of Sail's entire business model.
2. "There are no bad chips, only bad pricing" (Neil Movva): Every chip has its comparative advantage; the key is finding the right price and use case. Sail's core competency is finding a place for any chip in the inference stack.
3. "The internet is a one-time data subsidy" (Neil Movva): High-quality human text data (~30 trillion tokens) has been "squeezed dry" by models. Future capability improvements will depend on self-reinforcement learning on verifiable tasks.
4. "Security has become a 'proof of work'" (Neil Movva): The security of software depends on how many dollars you spend to have an AI try to break it. This vividly illustrates the application prospects of low-cost, large-scale inference in cybersecurity.
5. "From Hopper to Blackwell to Rubin, the improvement in performance per watt is not significant" (Neil Movva): This is his core technical argument for being bearish on NVIDIA long-term, suggesting the pace of NVIDIA's hardware progress may be slowing.
6. "We are willing to accept 95% availability" (Neil Movva): By accepting reliability far below industry standards, Sail can access extremely low-cost power and data center resources overlooked by mainstream players.
7. "We will not bid against Anthropic or OpenAI" (Neil Movva): Sail's "scavenger strategy" aims to leverage fragmented, non-standard compute resources that major customers "overlook," thereby building a cost advantage.
8. "Curiosity is the only thing you can't teach" (Neil Movva): When hiring, he doesn't look for CUDA or AI experience, but for engineers with a genuine passion for "performance engineering," because technologies change, but curiosity is timeless.