This episode covers Databricks, a private data and AI company valued at $130 billion, which evolved from an open-source project (Apache Spark, a free tool for processing large datasets). Guest Alan Tu says its underrated strength is hitting two home runs: making the open-source tech mainstream, then building a paid product far better than the free version. He's bullish because Databricks solves the painful 'data cleaning' step needed before any AI project, and AI hype actually boosts its core business. Key holdings: Databricks (ARR over $4B, cash-flow positive); Snowflake (main rival, but Databricks is winning by moving from unstructured to structured data); Microsoft Azure (frenemy—Databricks avoids being killed by ensuring customers also spend on Azure infrastructure).
Databricks is a privately held software company valued at $130 billion. Its core business helps enterprises collect, store, and process massive amounts of data for analytics and machine learning model training. In this episode, Alan Tu, a portfolio manager at WCM Investment Management, explains that
Guest: Alan Tu, Portfolio Manager and Analyst at WCM Investment Management (WCM invested in Databricks in December 2024). Main theme: Databricks has evolved from an academic open-source project into a private data and AI platform valued at $130 billion, with its core strengths lying in the "data lakehouse" architecture and a long-termist culture. The most impactful judgment in the episode: Alan Tu believes that Databricks' most underappreciated capability is not technology, but "hitting two home runs at once" — first achieving mainstream adoption of the open-source technology Apache Spark, then building proprietary products worth paying for, while most open-source companies cannot even hit the first home run.
Alan Tu believes that the fundamental problem Databricks solves is "transforming messy data into an analyzable format," which is the most time-consuming and foundational step in all data and AI projects.
Alan Tu notes: "Databricks helps lead the industry toward where they believe it should go — they identify pain points and propose solutions, rather than looking at the existing market and saying, 'We can make something similar too.'"
Alan Tu emphasizes that Databricks' founding team (seven researchers from the Berkeley AMP Lab) made three key bets in 2009: the cloud would grow, data would grow, and open source was a good way to build a business. The hardest part, however, was "making money on top of open source."
Key data: Databricks currently has an ARR of over $4 billion, with approximately $1 billion (one-quarter) coming from AI-related revenue.
Alan Tu believes that the competition between Databricks and Snowflake is not a "winner-takes-all" scenario, but Databricks' expansion from "unstructured data processing" to "structured data warehousing" has been more successful than Snowflake's reverse expansion.
Alan Tu notes: "Many companies have great products and strong momentum, but once Microsoft deems them too strategic, it will go all out to crush them. Databricks has never positioned itself in a way that makes cloud vendors 100% determined to eliminate it."
Alan Tu believes that AI's impact on Databricks is multi-layered, with the core being the structural tailwind from the consensus that "without a data strategy, there is no AI strategy."
1. Tailwind for Core Data Processing: As enterprises recognize that AI requires clean, well-organized data, this directly drives demand for Databricks' core products. This demand does not depend on whether AGI is achieved or how the next OpenAI model performs—as long as companies believe AI matters, they will invest in data engineering.
2. Usage by AI-Native Companies: AI-native companies, including major AI labs, also use Databricks internally.
3. New Product Opportunities: Databricks is building a full-stack product for agentic applications (Agent Bricks, Lake Base), helping enterprises create intelligent agent applications that automate specific tasks. This involves a full suite of tools including RAG (Retrieval-Augmented Generation), vector databases, and model evaluation.
Alan Tu adds: "Databricks' AI-related revenue has already reached $1 billion, but perhaps more important is the demand for data engineering driven by 'AI awareness' itself—this is a more durable and less volatile growth engine."
Alan Tu believes that Databricks' biggest risk is not competition, but "whether it can sustain execution and a long-termist culture as it scales."
Alan Tu concludes: "Every time you make a short-term-oriented decision, you open a leak somewhere. Whether Databricks can maintain a 'first principles' way of thinking is what I care about most."
| Position | Analyst View | Key Data |
|---|---|---|
| Databricks | Bullish (WCM invested in December 2024) | ARR > $4 billion; AI-related revenue ~$1 billion; data warehouse product annualized revenue approaching $1 billion; net dollar expansion rate > 140%; free cash flow positive |
| Snowflake | Neutral (primary competitor) | No specific data provided; directly competes with Databricks in the data warehouse space |
| Microsoft Azure | Neutral (strategic partner/co-opetition) | Offers Azure Databricks branded product; has a long-term strategic partnership with Databricks |
| AWS | Neutral (infrastructure provider/co-opetition) | Customers consume AWS infrastructure when using Databricks |
| Red Hat | Background reference (pioneer in open-source commercialization) | Built a service and support business based on Linux |
| Cloudera | Background reference (previous-generation data lake company) | Based on Hadoop technology, went public but technology was not strong enough |
1. Alan Tu believes the core challenge of open-source commercialization is "needing to hit two home runs simultaneously" — first achieving mainstream adoption of the open-source technology, then building a proprietary product worth paying for. Most companies fail to hit even the first, but Databricks succeeded. Supporting evidence: They started from first principles, did not copy the Red Hat model, and instead created proprietary implementations that far outperform the open-source version.
2. Alan Tu points out that Databricks' most underestimated capability is "market education" — they coined the category name "Lakehouse," which was mocked at launch but is now an industry-recognized architecture. Supporting evidence: This is not just product execution but also the marketing and commercialization ability to "make the entire market understand why this architecture is the future."
3. Alan Tu believes AI's most enduring driving force for Databricks is not AI products themselves, but the consensus that "without a data strategy, there is no AI strategy" — this creates a long-term tailwind for the core data processing business, independent of whether AGI is achieved. Supporting evidence: As long as enterprises believe AI is important, they will invest in data engineering, which is Databricks' core value.
4. Alan Tu emphasizes that Databricks' long-termist culture is proven by specific "trade-offs" — they never released an on-premise version (even when the cloud was not yet widely accepted), named the company "Databricks" instead of "Spark" (forgoing short-term brand dividends), and deliberately refrained from charging for certain features (sacrificing short-term revenue). Supporting evidence: These decisions all carried clear short-term costs at the time, but the team adhered to long-term judgment.
5. Alan Tu points out that Databricks' "co-opetition" with cloud providers is an underestimated strategic advantage — customers using Databricks also consume cloud provider infrastructure, so cloud providers have no 100% incentive to eliminate Databricks. Supporting evidence: Many growth-stage software companies have been killed by giants like Microsoft for being "too strategic," but Databricks avoided this fate through aligned interests.
6. Alan Tu believes Databricks' cost structure is healthier than it appears on the surface — core data processing is CPU-based and does not require large-scale GPU procurement; at a $4 billion ARR scale, it has already achieved positive free cash flow. Supporting evidence: Even though model inference involves GPU costs, the magnitude is completely different from AI training companies.
7. Alan Tu points out that Databricks' "data lakehouse" architecture has been more successful expanding from unstructured to structured data than Snowflake's reverse expansion — the data warehouse product reached nearly $1 billion in annualized revenue within two years, while Snowflake's expansion into data engineering is far smaller in scale. Supporting evidence: This is not just a technical advantage but also the result of strategic decisions such as "open format" and "no storage-based pricing."
8. Alan Tu believes Databricks' biggest risk is not competition, but "whether it can maintain execution and long-termist culture as it scales" — every short-term-oriented decision opens a loophole, and the Agentic application space is still in the "category creation" phase, requiring a repeat demonstration of product execution plus market education. Supporting evidence: Some companies in the industry have paid the price for "distraction" or "slow product launches"; Databricks is advancing multiple product lines simultaneously, and sustained execution is key.