← Back to list
Colossus (Invest Like the Best / Business Breakdowns)Podcast12 Jun 2018Source: traffic.libsyn.comHost: Patrick O'Shaughnessy

Michael Recce – Tim Cook’s Dashboard - [Invest Like the Best, EP.91]

In plain words

This piece explains how to use machine learning and alternative data to rebuild a company's business dashboard for an information edge. The author says ML is useless for predicting stock prices but very effective for predicting business growth. Three key holdings: Blue Apron (IPO looked good on surface, but cohort analysis revealed declining customer retention and rising acquisition costs—a risk); Lululemon (inferred rising male customer share, seen as positive); Walmart (9% of customers drive 50% of revenue, with concentration increasing—a warning sign).

AI SummaryAI-generated · may contain errors · verify against the original

At a Glance Neuberger Berman’s Chief Data Scientist, Michael Recce, explores how data and machine learning can be leveraged to create an information edge in investing. His core argument: if the world’s best Apple analyst were equipped with Tim Cook’s private business dashboard, its value would be im

~10 min full read · 9 sections
Deep Analysis

Michael Recce – Tim Cook's Dashboard - [Invest Like the Best, EP.91]

At a Glance

Michael Recce is the Chief Data Scientist at Neuberger Berman, previously serving as Chief Data Scientist at Point72 and GIC (Singapore's sovereign wealth fund), and founded companies in anti-money laundering analytics and online ad targeting. The main thread of this episode: how to use machine learning and alternative data to rebuild enterprise-level dashboards and create an information advantage. The most impactful judgment in the entire episode: machine learning is "useless" for predicting stock prices, but extremely effective for "relatively stable" problems like predicting business growth — because stock prices are non-stationary, while business growth is stationary.


Theme 1: Data Is Not Information — The Key Lies in "Rebuilding the Corporate Dashboard"

Recce believes that equipping the world’s best Apple analyst with Tim Cook’s private business dashboard would be invaluable. His core goal is to rebuild similar dashboards for multiple companies — not by simply aggregating data, but by drilling down to the granularity of product lines, geographic regions, and customer cohorts.

  • Mechanism Breakdown: Recce views corporate data as "digital residue" — transaction-level information from low-cost electronic devices, such as credit card records, web browsing behavior, and mobile phone location data. The key is that this data is not used directly; it must undergo a process of "information extraction."
  • Data ≠ Information: Recce emphasizes that data itself has become commoditized ("everyone can buy credit card data"), but information is a different matter. "People have been picking up shiny pebbles on the beach surface without truly using data to build models." The real advantage lies in how data is processed — for example, by inferring consumer demographics (gender, age, income level) to observe whether Lululemon’s appeal to male customers is increasing.
  • Specific Case: When Blue Apron went public, the surface-level data showed rising customer numbers and revenue. However, after breaking it down by customer cohort, it became clear that each new cohort exhibited declining customer retention and rising customer acquisition costs — the underlying business fundamentals were completely different. "A single level of detail changed the entire picture of the business."

Recce compares this process to "building a Zillow for the stock market." Zillow automatically values properties; while it may not match the precision of the top appraisers, it excels in automation and scalability. Similarly, the approach is to first automatically construct corporate valuation models using data, then compare them with market prices.


Theme 2: The Boundary Between Useful and Useless Machine Learning — "Predicting Stock Prices Is a Disaster, Predicting Business Growth Is a Powerful Tool"

Recce draws a clear boundary: machine learning is "useless" for predicting stock prices, but "very useful" for predicting corporate business growth. The core reason lies in the issue of stationarity.

  • Non-stationary vs. Stationary: "Stock price fluctuations are driven by various emotions and institutional factors, making them non-stationary; but the growth of a specific business segment of a company is like a rock—very stable, very stationary." Recce uses an analogy to illustrate: identifying cats on YouTube works because the category "cat" is relatively stable; predicting stock prices fails because the mapping from input variables to output variables does not even exist.
  • Historical Context: Recce's experience in the internet advertising field (Quantcast) serves as a key foundation—"If I know which ad to show you, I know what product you are interested in; if I know the situation across all geographic regions and all products, I know who is winning in the market today." This logic directly transfers to investing: transaction-level data reveals the true health of a company, not market sentiment.
  • Falsification Condition: If a machine learning model attempts to predict short-term stock price movements, Recce believes "the problem itself is misleading—assuming there is a mapping from input to output is already wrong."

Theme 3: Predicting "Failure" Is Easier Than Predicting "Success" — Asymmetric Prediction Space

Recce proposes a counterintuitive judgment: predicting when a company will fail is far easier than predicting when it will achieve great success. This insight stems from his experience helping to build a university admissions essay scoring system.

  • Mechanism Breakdown: In essay scoring, judgments on "poor essays" are highly consistent — there is consensus on error types. However, judgments on "good essays" vary from person to person — "what motivates one person may not motivate another." The same applies to companies: the space of failure (loss of market share, deterioration of customer base, financial issues) is relatively limited and its patterns are identifiable; while the space of success (innovation, market expansion, new business models) is extremely broad and difficult to predict.
  • Data Support: Recce cites research showing that the average retention time of S&P 500 constituents has dropped from 40–50 years to about 12 years, and is still declining. "Trillions of dollars flow into passive investing each year, but the stocks you buy are increasingly less likely to be long-term winners." If data can identify "losers," one can construct an index-like portfolio — overweight winners and underweight losers — earning an extra 50–100 basis points annually, with an operating cost of only 10 basis points.
  • Implication: This could lead to a shift of capital from passive management back to active management — "an AI-driven active management process with low volatility, essentially tracking the index but slightly outperforming, because it excels at identifying negative signals."

Theme 4: The Sustainable Edge in Data Science – Depth of Processing and Multi-Source Cross-Validation

Recce argues that the moat in data science lies not in the data itself (which has become commoditized), but in the "depth of processing from data to information" and the "ability to cross-validate across multiple data sources."

  • Depth of Processing: Most users merely aggregate data into revenue figures ("If you're just aggregating into revenue, why bother with granular data?"). The real advantage lies in building enterprise models—for example, observing changes in the Pareto distribution (80-20 rule): Walmart sees 9% of customers contributing 50% of revenue. If this ratio drops to 8%, it indicates a steeper distribution and higher customer concentration—an unhealthy signal. In contrast, Amazon's distribution is flattening, with a broadening customer base, which is a healthier signal.
  • Multi-Source Cross-Validation: "Digital residue" inherently carries noise and errors, requiring multiple data sources to corroborate each other. Beyond the revenue side, one can also examine the cost side—job postings (a leading indicator for OPEX), raw material prices (e.g., the impact of chicken prices on KFC), background check data, and more.
  • Tech Stack Advantage: Recce emphasizes that the tech stacks of internet companies are far ahead of those in traditional financial institutions. "In traditional institutions, the tech team's first priority is to keep the trains running on time, the second is safety, and a distant third is keeping up with technology. In internet companies, if you don't keep up with technology, you're dead." For example, a Google AI researcher extracted all French house numbers from Google Street View in a single night—something nearly impossible under a traditional IT architecture.

Theme 5: How Data Scientists Are Organized — "Prospectors" Over "Miners"

Recce uses the "miners vs. prospectors" analogy to describe the organizational philosophy of data science teams. He argues that in a completely new domain, letting smart individuals explore on their own (prospector model) is superior to top-down centralized directives (miner model).

  • Mechanism Breakdown: Miner model — the entire company works in unison toward a single goal (e.g., Intel's next-generation processor), offering high efficiency but concentrated risk ("all eggs in one basket"). Prospector model — each scientist has their own lab and does not need to communicate with others; as long as someone strikes gold, the whole company wins.
  • Practical Approach: Short-cycle agile iteration — "Ideas always outnumber resources. Let the team self-organize to select the best ideas, validate them in a short time frame, continue if proven effective, and switch if not." Recce believes this model is not only more effective but also creates a better working environment.
  • Integration with the Investment Process: Recce suggests that traditional fund managers should build a "bridge" with data scientists — "You say you like a company because of strong management. What does 'strong' look like? Is it better cost efficiency? New products? New geographic expansion? Different types of hiring? If you tell me what it looks like, I can go find the data to verify it."

Mentioned Positions

Position Analyst View Key Data
Blue Apron Risk Warning At IPO, customer count and revenue appeared to rise, but after cohort-level decomposition, new customer retention declined and customer acquisition costs increased
Lululemon Bullish (Case Study) By inferring customer gender, the report observed a sustained increase in the proportion of male customers
Walmart Risk Warning (Case Study) 9% of customers contribute 50% of revenue; the Pareto distribution is steepening (rising customer concentration)
Amazon Bullish (Case Study) The Pareto distribution is flattening, indicating a broadening customer base—a healthier signal
Starbucks Neutral (Case Study) Data can measure the conversion rate of loyalty programs and the subsequent changes in spending after conversion
Home Depot Neutral (Case Study) When monthly data is strong, analysts tend to expect mean reversion rather than linear extrapolation
Whole Foods Neutral (Case Study) After Amazon's acquisition, price cuts triggered a "honeymoon period" effect, effectively reopening all stores

Judgments Worth Remembering

1. "Machine learning is useless for predicting stock prices, but extremely useful for predicting business growth." (Recce) — Stock prices are non-stationary, driven by sentiment and institutional factors; business growth is stationary, as solid as a rock. This is the fundamental boundary of machine learning applications.

2. "Predicting when a company will fail is far easier than predicting when it will achieve great success." (Recce) — The scope for failure is limited and patterns are identifiable (market share loss, customer deterioration); the scope for success is infinite and varies by case. This makes "identifying losers" a more viable strategy.

3. "Data has been commoditized, but information has not." (Recce) — Everyone can buy credit card data, but how to extract information from it (e.g., inferring consumer demographics, observing customer cohort changes) is the true moat.

4. "If you only aggregate granular data into revenue figures, why do you even need granular data?" (Recce) — Most users only perform superficial aggregation; the real value lies in reconstructing the company dashboard: breaking it down by product, geography, customer cohort, and timeline.

5. "The direction of change in the Pareto distribution reveals corporate health more than revenue figures do." (Recce) — A steepening distribution at Walmart (rising customer concentration) is an unhealthy signal; a flattening distribution at Amazon (expanding customer base) is a healthy signal.

6. "The prospector model is superior to the miner model — in new domains, letting smart individuals explore independently is more effective than centralized directives." (Recce) — Short-cycle agile iteration, self-organization to select the best ideas, and rapid validation.

7. "If you say you like a company because of excellent management, what does 'excellent' look like? Tell me, and I can go find data to verify it." (Recce) — This is the key bridge connecting fundamental analysis with data science.

8. "The future is already here — it's just not evenly distributed." (Recce, quoting William Gibson) — Data science on Wall Street is still in its early stages; it will completely reshape the industry landscape within 10 years, and sell-side firms may be disrupted before buy-side firms.