← Back to list
Colossus (Invest Like the Best / Business Breakdowns)Podcast25 Aug 2026Source: colossus.comHost: Patrick O'Shaughnessy

Neil Movva - Making AI 10x Cheaper - [Invest Like the Best, EP.488]

In plain words

This interview argues that the future of AI isn't chatbots but background 'agents' that run for hours, making cost more important than speed. Guest Neil Movva thinks the premium charged by labs like OpenAI for being 3-6 months ahead is unsustainable because open-source models catch up. His company, Sail, buys unwanted chips (like AMD) and power, even accepting occasional outages, to slash AI inference costs. Key mentions: NVIDIA (good short-term but not long-term due to slow performance-per-watt gains), AMD (undervalued, an arbitrage opportunity), and Cerebras (high bandwidth for storing model weights).

AI SummaryAI-generated · may contain errors · verify against the original

Neil Movva, founder of Sail, proposed on the Invest Like the Best podcast the construction of a "token factory" to reduce AI inference costs by 10x. His core argument is that future AI agents will run in the background for extended periods (hours to days), where latency becomes irrelevant and cost b

~11 min full read · 9 sections
Deep Analysis

Here is the English translation of the provided Chinese investment research notes, following all specified rules.

At a Glance

Neil Movva, founder of Sail, proposes building a "token factory" to reduce AI inference costs by 10x. His core thesis is that future AI agents will run in the background for extended periods (hours to days), making latency irrelevant and cost the critical variable. Sail employs a "scavenger strategy" to acquire chips and power that others do not want, optimizing the throughput-latency trade-off in GPUs to minimize per-token cost. Neil Movva believes the premium charged by frontier labs for being 3-6 months ahead is unsustainable, as the diffusion and distillation of open-source model capabilities cannot be stopped.

Topic Sections

1. The Future Belongs to "Background Agents," Not "Chatbots"

Neil Movva argues that the future of AI inference lies in long-running background agents, not real-time interactive chatbots. He believes that when users no longer wait for an AI response but instead let it work autonomously in the background, latency becomes irrelevant, and cost becomes the only critical variable.

  • Argument: This thesis is based on the trend of "test-time compute scaling"—giving a model more time yields better answers. Starting with Opus 4.5, models can already handle tasks lasting about an hour, and task duration is growing exponentially. He predicts that by the end of this year, background and real-time workloads will each account for 50%, eventually evolving to 90% background, 10% real-time.
  • Implication: The token consumption of background agents is "unbounded" because it is no longer limited by human attention. This creates a business opportunity to drastically reduce token costs. Verification Signal: Observe whether the average task execution time of frontier models (e.g., GPT-5, Claude 4) continues to increase.
2. Throughput vs. Latency: The "Bus vs. Car" Debate for GPUs

Neil Movva points out that a GPU is inherently a "throughput machine," but the current AI industry forces it to act like a "private car" in pursuit of low latency, sacrificing efficiency. Sail's strategy is to return the GPU to its "bus" nature, maximizing throughput.

  • Mechanism Breakdown: GPUs process tasks in parallel through "batching." The larger the batch, the higher the throughput, but the longer the latency for a single request. He analogizes this to a "bus" (throughput-priority, slow route but carries many passengers) vs. a "private car" (latency-priority, fast route but carries few passengers). The current industry (e.g., customers like Cursor) forces all inference companies to optimize for the "private car" model, but Neil believes this is a thing of the past.
  • Data Chain: NVIDIA's NVLink interconnect technology is a "necessity" for low-latency inference because it allows matrix multiplication to be split across multiple GPUs for parallel computation, reducing single-operation time. However, Neil notes this requires 8x the hardware investment for only a 4-5x speed improvement (sub-linear scaling), which is not cost-effective for a throughput-focused company like Sail.
  • Implication: Sail will adopt different parallelism schemes (e.g., expert parallelism, pipeline parallelism) and is willing to accept very low output speeds of 1-10 tokens/second in exchange for the lowest possible per-token cost. Falsification Condition: If the mainstream form of future AI applications remains real-time interaction rather than background agents, Sail's strategy will lose its market.
3. The "Scavenger Strategy": Full-Stack Arbitrage from Chips to Power

Neil Movva's core strategy is "scavenging"—buying chips and power that no one else wants on the market, making them efficient through software optimization, and thus building an ultra-low-cost inference factory.

  • Chip Scavenging: Neil believes "there are no bad chips, only bad pricing." He is willing to buy any chip (AMD, Cerebras, Google TPU, AWS Trainium, etc.) as long as the price is right. The key is that suppliers of these chips (e.g., AMD) have less mature software ecosystems than NVIDIA, leaving significant room for optimization (alpha) for a "software expert" like Sail. He explicitly states that AMD chips are undervalued due to market perception bias, which plays right into his hands.
  • Power & Data Center Scavenging: Neil plans to utilize distributed small-scale data centers (1 megawatt level, about the size of 8 refrigerator-sized racks), rather than the industry-standard hundred-megawatt mega data centers. He is willing to accept 95% availability (i.e., 5% downtime), which would be fatal for real-time applications but acceptable for background agents. He even considers using intermittent power sources like solar and wind, dynamically migrating workloads based on weather forecasts.
  • Implication: Through scavenging, Sail can avoid direct competition for compute resources with major customers like Anthropic and OpenAI, acquiring resources they "overlook." Verification Signal: Observe whether Sail can consistently secure a stable supply of non-NVIDIA chips and whether the operating costs of its distributed data center network are significantly below the industry average.
4. A Contrarian View on NVIDIA: Bullish Short-Term, Bearish Long-Term

Neil Movva holds a contrarian view on NVIDIA: "bullish short-term, bearish long-term." He believes NVIDIA's position is solid in low-latency inference, but the future belongs to high-throughput inference, for which NVIDIA's architecture is not optimal.

  • Bullish Short-Term: NVIDIA's NVLink and powerful software ecosystem make it irreplaceable for low-latency inference. Neil admits that for the lowest latency, NVIDIA's Blackwell chip is the only choice.
  • Bearish Long-Term: Neil's core argument is that from Hopper to Blackwell to Rubin, the improvement in NVIDIA's "performance per watt" is not significant. This means that when the market shifts towards latency-insensitive background inference, NVIDIA's moat (NVLink, low latency) will lose its value. At that point, other chips with better "flops per dollar" will become more competitive.
  • Data Chain: Neil also offers a contrarian view on chip manufacturing. He argues that losing access to TSMC would be a major shock but not a disaster, as the best Western process (e.g., Intel) is at most 2x behind in performance per watt, a much smaller gap than market panic suggests.
5. The Future of Data and Models: From Internet "Subsidy" to Self-Play

Neil Movva believes that the internet, as a source of high-quality text data, is a one-time "subsidy," and future model capability improvements will primarily rely on self-reinforcement learning (RL) on verifiable tasks.

  • Data Chain: High-quality text data amounts to roughly 30 trillion tokens; relaxing the standard yields about 300 trillion tokens. Models have essentially "read" it all. Going forward, the value of random user feedback is less than expert feedback because the models themselves already surpass average users.
  • Mechanism Breakdown: Future data will come from an "environment gym" (RL environment), where models engage in self-play on verifiable tasks (e.g., coding, math), improving through trial and error and feedback. This is seen as the path to AGI: continuously stacking "specialized intelligence" until there are no gaps.
  • Implication: This trend benefits Sail, as both RL training and inference require massive, low-cost compute resources. Falsification Condition: If model capability improvements stagnate without new human data, Neil's "self-play" hypothesis could be disproven.

Position Moves

Position Analyst Sentiment Key Data
NVIDIA Bullish short-term, Bearish long-term Performance per watt improvement from Hopper to Rubin is not significant; NVLink is a necessity for low latency but costly (8x hardware for 4-5x speed).
AMD Bullish (arbitrage opportunity exists) Market perception bias leads to undervaluation; software ecosystem lags NVIDIA, leaving room for optimization by Sail.
Cerebras Bullish (as an accelerator) Its Wafer Scale Engine 3 offers SRAM bandwidth up to 21 PB/s (NVIDIA HBM ~10 TB/s), suitable for storing model weights, but KV Cache capacity is limited.
Anthropic / OpenAI Neutral (competitive relationship) They pay a high "premium" for being 3-6 months ahead, but Neil believes this premium is unsustainable.
DeepSeek Positive mention Achieved "an order of magnitude" progress in KV Cache compression, with continued annual improvements.
Intel Neutral (potential alternative) In an extreme scenario without TSMC, its process is at most 2x behind in performance per watt, which is not unacceptable.

Memorable Judgments

1. "The best latency is no latency" (Neil Movva): When AI agents work autonomously in the background, users don't wait, and the latency problem disappears. This is the cornerstone of Sail's entire business model.

2. "There are no bad chips, only bad pricing" (Neil Movva): Every chip has its comparative advantage; the key is finding the right price and use case. Sail's core competency is finding a place for any chip in the inference stack.

3. "The internet is a one-time data subsidy" (Neil Movva): High-quality human text data (~30 trillion tokens) has been "squeezed dry" by models. Future capability improvements will depend on self-reinforcement learning on verifiable tasks.

4. "Security has become a 'proof of work'" (Neil Movva): The security of software depends on how many dollars you spend to have an AI try to break it. This vividly illustrates the application prospects of low-cost, large-scale inference in cybersecurity.

5. "From Hopper to Blackwell to Rubin, the improvement in performance per watt is not significant" (Neil Movva): This is his core technical argument for being bearish on NVIDIA long-term, suggesting the pace of NVIDIA's hardware progress may be slowing.

6. "We are willing to accept 95% availability" (Neil Movva): By accepting reliability far below industry standards, Sail can access extremely low-cost power and data center resources overlooked by mainstream players.

7. "We will not bid against Anthropic or OpenAI" (Neil Movva): Sail's "scavenger strategy" aims to leverage fragmented, non-standard compute resources that major customers "overlook," thereby building a cost advantage.

8. "Curiosity is the only thing you can't teach" (Neil Movva): When hiring, he doesn't look for CUDA or AI experience, but for engineers with a genuine passion for "performance engineering," because technologies change, but curiosity is timeless.