← Back to list
Colossus (Invest Like the Best / Business Breakdowns)Podcast8 Jan 2026Source: joincolossus.comHost: Colossus

Databricks: From Data to Decisions - [Business Breakdowns, EP.238]

In plain words

This episode covers Databricks, a private data and AI company valued at $130 billion, which evolved from an open-source project (Apache Spark, a free tool for processing large datasets). Guest Alan Tu says its underrated strength is hitting two home runs: making the open-source tech mainstream, then building a paid product far better than the free version. He's bullish because Databricks solves the painful 'data cleaning' step needed before any AI project, and AI hype actually boosts its core business. Key holdings: Databricks (ARR over $4B, cash-flow positive); Snowflake (main rival, but Databricks is winning by moving from unstructured to structured data); Microsoft Azure (frenemy—Databricks avoids being killed by ensuring customers also spend on Azure infrastructure).

AI SummaryAI-generated · may contain errors · verify against the original

Databricks is a privately held software company valued at $130 billion. Its core business helps enterprises collect, store, and process massive amounts of data for analytics and machine learning model training. In this episode, Alan Tu, a portfolio manager at WCM Investment Management, explains that

~14 min full read · 9 sections
Deep Analysis

Databricks: From Data to Decisions - [Business Breakdowns, EP.238]

At a Glance

Guest: Alan Tu, Portfolio Manager and Analyst at WCM Investment Management (WCM invested in Databricks in December 2024). Main theme: Databricks has evolved from an academic open-source project into a private data and AI platform valued at $130 billion, with its core strengths lying in the "data lakehouse" architecture and a long-termist culture. The most impactful judgment in the episode: Alan Tu believes that Databricks' most underappreciated capability is not technology, but "hitting two home runs at once" — first achieving mainstream adoption of the open-source technology Apache Spark, then building proprietary products worth paying for, while most open-source companies cannot even hit the first home run.


1. Databricks’ Core Value: Solving the Most Painful Step — Data Preparation

Alan Tu believes that the fundamental problem Databricks solves is "transforming messy data into an analyzable format," which is the most time-consuming and foundational step in all data and AI projects.

  • Analogy: When given an Excel spreadsheet with inconsistent data formats (e.g., a price column containing currency symbols or text), 80%-90% of the time is spent on "data cleaning" rather than analysis itself. Databricks moves this process to the cloud and scales it to massive datasets.
  • Scale Difference: The data sources enterprises face include structured data (row-and-column tables) and unstructured data (log files, images, videos, clickstream data). Databricks handles "a completely different order of magnitude."
  • Core Product Evolution: From Apache Spark (distributed computing engine) → MLflow (machine learning lifecycle management) → Delta (introducing ACID transaction guarantees) → data warehouse product (directly competing with Snowflake). The data warehouse product, launched about two years ago, has already reached an annualized revenue of nearly $1 billion, demonstrating TAM expansion through "multiple products → multiple personas."

Alan Tu notes: "Databricks helps lead the industry toward where they believe it should go — they identify pain points and propose solutions, rather than looking at the existing market and saying, 'We can make something similar too.'"


2. Open-Source Commercialization: Two Grand Slams from "Academic Project" to "Commercial Platform"

Alan Tu emphasizes that Databricks' founding team (seven researchers from the Berkeley AMP Lab) made three key bets in 2009: the cloud would grow, data would grow, and open source was a good way to build a business. The hardest part, however, was "making money on top of open source."

  • The open-source dilemma: Open source can drive massive adoption and mindshare, but "the free version becomes your biggest competitor"—competitors with more distribution channels and customer relationships can better monetize the same technology.
  • Databricks' solution: Rather than copying Red Hat's model (providing support services for Linux), they approached it from first principles: they had to create a "better product worth paying for." They built a proprietary implementation of Spark that far surpassed the open-source version in performance, reliability, and scalability, and made it clear: "If you want this version, you have to pay."
  • Cultural trait: "When you become popular because of open-source success and suddenly announce a competitive product that charges a fee, the community will see it as a betrayal. You need to be willing to be the 'villain.'" Alan Tu believes that the academic background meant the team had "no preconceived business notions," allowing them to make purer product decisions.

Key data: Databricks currently has an ARR of over $4 billion, with approximately $1 billion (one-quarter) coming from AI-related revenue.


3. Competitive Landscape: Data Lakehouse vs. Snowflake, and the "Co-opetition" with Cloud Vendors

Alan Tu believes that the competition between Databricks and Snowflake is not a "winner-takes-all" scenario, but Databricks' expansion from "unstructured data processing" to "structured data warehousing" has been more successful than Snowflake's reverse expansion.

  • Market Reality: Enterprises typically use multiple vendors simultaneously. A common scenario is "first process data with Databricks, then store the processed data in Snowflake's data warehouse."
  • Lakehouse: A category name coined by Databricks, merging the data lake (unstructured) with the data warehouse (structured). At launch, it was mocked as "too clever," but has since become an industry-recognized architectural category. Alan Tu believes this is an underappreciated strength of Databricks—not only executing the product but also educating the entire market.
  • Co-opetition with Cloud Vendors: Databricks has had a strategic partnership with Microsoft Azure from day one (Azure Databricks). Key principle: Never give cloud vendors a 100% incentive to "kill" Databricks—customers using Databricks also consume cloud vendor infrastructure (compute and storage), creating aligned interests.
  • Pricing Strategy: Charges based on usage (compute consumption), but does not charge for storage—this contrasts with Snowflake. Databricks deliberately adopts open formats, allowing customers to "store data anywhere without moving it into Databricks," which is an "architectural disruption."

Alan Tu notes: "Many companies have great products and strong momentum, but once Microsoft deems them too strategic, it will go all out to crush them. Databricks has never positioned itself in a way that makes cloud vendors 100% determined to eliminate it."


4. Positioning in the AI Era: Data Strategy is a Prerequisite for AI Strategy

Alan Tu believes that AI's impact on Databricks is multi-layered, with the core being the structural tailwind from the consensus that "without a data strategy, there is no AI strategy."

  • Three Layers of Impact:

1. Tailwind for Core Data Processing: As enterprises recognize that AI requires clean, well-organized data, this directly drives demand for Databricks' core products. This demand does not depend on whether AGI is achieved or how the next OpenAI model performs—as long as companies believe AI matters, they will invest in data engineering.

2. Usage by AI-Native Companies: AI-native companies, including major AI labs, also use Databricks internally.

3. New Product Opportunities: Databricks is building a full-stack product for agentic applications (Agent Bricks, Lake Base), helping enterprises create intelligent agent applications that automate specific tasks. This involves a full suite of tools including RAG (Retrieval-Augmented Generation), vector databases, and model evaluation.

  • Cost Structure Advantage: Core data processing workloads are based on CPUs, not GPUs. Databricks does not need to procure GPUs on the same massive scale as AI training companies. While model serving involves GPU costs, the scale is entirely different.
  • Financial Health: With $4 billion in ARR, the company has already achieved positive free cash flow. Its main costs are traditional software company expenses like personnel and R&D, with no need for heavy capital expenditure on hardware.

Alan Tu adds: "Databricks' AI-related revenue has already reached $1 billion, but perhaps more important is the demand for data engineering driven by 'AI awareness' itself—this is a more durable and less volatile growth engine."


5. Risk and the Long-Termist Culture

Alan Tu believes that Databricks' biggest risk is not competition, but "whether it can sustain execution and a long-termist culture as it scales."

  • Execution Risk: The market is highly dynamic, and the speed of innovation is critical. Some companies in the industry have paid the price for "distraction" or "delayed product launches." Databricks is advancing multiple product lines simultaneously, and sustained execution capability is key.
  • Uncertainty in AI Products: The agentic application space is still in the "category creation" phase—the industry has yet to reach a consensus on tool sets and product forms. Databricks needs to once again demonstrate its ability in product execution and market education, much like it did when defining "lakehouse."
  • Cultural Risk: A long-termist culture requires concrete "trade-offs" to prove itself. Alan Tu cites several examples:
  • Naming the company "Databricks" instead of "Spark"—forgoing short-term brand equity to leave room for a multi-product platform in the future.
  • Never offering an on-premises version—even when the cloud was not yet widely accepted, the team firmly believed the cloud was the future.
  • Deliberately not charging for certain features (e.g., governance layer, open formats)—sacrificing short-term revenue to build an ecosystem moat over the long term.
  • Private vs. Public Trade-off: During the 2022 growth tech stock correction, Databricks, by remaining private, was able to "keep attacking," while many public companies were forced to retrench. This reinforced management's preference for staying private. However, Alan Tu notes that large-scale fundraising (e.g., the recent multi-billion-dollar rounds) is primarily used to address tax issues related to employee equity incentives, rather than operational needs.

Alan Tu concludes: "Every time you make a short-term-oriented decision, you open a leak somewhere. Whether Databricks can maintain a 'first principles' way of thinking is what I care about most."


Referenced Positions

Position Analyst View Key Data
Databricks Bullish (WCM invested in December 2024) ARR > $4 billion; AI-related revenue ~$1 billion; data warehouse product annualized revenue approaching $1 billion; net dollar expansion rate > 140%; free cash flow positive
Snowflake Neutral (primary competitor) No specific data provided; directly competes with Databricks in the data warehouse space
Microsoft Azure Neutral (strategic partner/co-opetition) Offers Azure Databricks branded product; has a long-term strategic partnership with Databricks
AWS Neutral (infrastructure provider/co-opetition) Customers consume AWS infrastructure when using Databricks
Red Hat Background reference (pioneer in open-source commercialization) Built a service and support business based on Linux
Cloudera Background reference (previous-generation data lake company) Based on Hadoop technology, went public but technology was not strong enough

Judgments Worth Remembering

1. Alan Tu believes the core challenge of open-source commercialization is "needing to hit two home runs simultaneously" — first achieving mainstream adoption of the open-source technology, then building a proprietary product worth paying for. Most companies fail to hit even the first, but Databricks succeeded. Supporting evidence: They started from first principles, did not copy the Red Hat model, and instead created proprietary implementations that far outperform the open-source version.

2. Alan Tu points out that Databricks' most underestimated capability is "market education" — they coined the category name "Lakehouse," which was mocked at launch but is now an industry-recognized architecture. Supporting evidence: This is not just product execution but also the marketing and commercialization ability to "make the entire market understand why this architecture is the future."

3. Alan Tu believes AI's most enduring driving force for Databricks is not AI products themselves, but the consensus that "without a data strategy, there is no AI strategy" — this creates a long-term tailwind for the core data processing business, independent of whether AGI is achieved. Supporting evidence: As long as enterprises believe AI is important, they will invest in data engineering, which is Databricks' core value.

4. Alan Tu emphasizes that Databricks' long-termist culture is proven by specific "trade-offs" — they never released an on-premise version (even when the cloud was not yet widely accepted), named the company "Databricks" instead of "Spark" (forgoing short-term brand dividends), and deliberately refrained from charging for certain features (sacrificing short-term revenue). Supporting evidence: These decisions all carried clear short-term costs at the time, but the team adhered to long-term judgment.

5. Alan Tu points out that Databricks' "co-opetition" with cloud providers is an underestimated strategic advantage — customers using Databricks also consume cloud provider infrastructure, so cloud providers have no 100% incentive to eliminate Databricks. Supporting evidence: Many growth-stage software companies have been killed by giants like Microsoft for being "too strategic," but Databricks avoided this fate through aligned interests.

6. Alan Tu believes Databricks' cost structure is healthier than it appears on the surface — core data processing is CPU-based and does not require large-scale GPU procurement; at a $4 billion ARR scale, it has already achieved positive free cash flow. Supporting evidence: Even though model inference involves GPU costs, the magnitude is completely different from AI training companies.

7. Alan Tu points out that Databricks' "data lakehouse" architecture has been more successful expanding from unstructured to structured data than Snowflake's reverse expansion — the data warehouse product reached nearly $1 billion in annualized revenue within two years, while Snowflake's expansion into data engineering is far smaller in scale. Supporting evidence: This is not just a technical advantage but also the result of strategic decisions such as "open format" and "no storage-based pricing."

8. Alan Tu believes Databricks' biggest risk is not competition, but "whether it can maintain execution and long-termist culture as it scales" — every short-term-oriented decision opens a loophole, and the Agentic application space is still in the "category creation" phase, requiring a repeat demonstration of product execution plus market education. Supporting evidence: Some companies in the industry have paid the price for "distraction" or "slow product launches"; Databricks is advancing multiple product lines simultaneously, and sustained execution is key.