The measurement problem when customers never cancel
Non-contractual businesses have a basic accounting problem. A subscriber cancels. A bank account closes. But an occasional buyer at an e-commerce site, marketplace, or specialty retailer can just stop showing up. Silence is ambiguous. Maybe the customer is gone. Maybe she buys once every 20 months.
Buy-Till-You-Die models were built for this setting. The BG/NBD and Pareto/NBD families take each customer's recency, frequency, and time under observation, then estimate future purchasing and a probability of being alive, usually written as P(alive). That number now appears in churn dashboards, customer equity models, CLV systems, and customer-based valuation work.
Ulrich's paper argues that practice has been too casual about what P(alive) is measuring. The issue is not that BTYD models are useless. The issue is narrower and more damaging. A finite-horizon repeat-purchase forecast is auditable. A total alive-customer count is not, at least not from the same data alone.
P(alive) is an infinite-horizon forecast, not a count
The mechanism is simple. In the BG/NBD setup, customers can drop out only after a purchase. So if a customer is classified as alive, that means she will buy again eventually. P(alive) is therefore the limit of a question we can actually test: what is the probability this customer buys again within H months?
Set H to 6, 12, or 18 months, and the forecast can be scored later. Wait long enough and the event either happened or it did not. Set H to infinity, and the event cannot be fully scored. The model has to extrapolate past the observable window.
That turns summed P(alive) into a different object from a repeat-purchase forecast. It is not a count in the normal accounting sense. It is a model-implied total, selected by assumptions about unobserved heterogeneity and dropout. Ulrich calls this dead reckoning. The metaphor fits. The model projects forward from observed transactions, but there is no final fix that tells you the true number of customers still alive.
The identification problem is concrete. A firm with fewer remaining customers buying more often can look a lot like a firm with more remaining customers buying less often. Recency and timing help, but they do not fully break the tie. The likelihood can fit observed purchases while leaving the alive count weakly pinned down.
Similar repeat-purchase forecasts, very different alive-customer totals
The empirical results are the part practitioners should read closely.
On a seven-year panel with 31,683 unique customers, Ulrich compares model specifications that give nearly the same observable forecasts. Their 18-month repeat-purchase forecasts are close. Yet the implied number of alive customers ranges from 3,654 to 27,734. That is a 7.6x spread on the same customer base.
Even a software setting matters. A default weighting parameter in customer-base software moves the alive count by 42 percent while leaving the observable forecast nearly unchanged. On the CDNOW benchmark data, the same pattern appears, with a 2.4x spread in point estimates of the customer base.
This is the key distinction. If the model says 10,000 people will buy within 18 months, we can wait 18 months and check. If the model says 20,000 customers are alive in the infinite-horizon sense, later data can only give a lower bound. Every realized returner proves she was alive at the scoring date. Non-returners do not prove the opposite. In Ulrich's panel, five later years of purchases falsified the maximum-likelihood count from below.
The paper's most useful diagnostic is the category error around calibration. Summed P(alive) overshoots realized 18-month returners by 2.25x. But the model's own 18-month forecast misses by only 1.18x. Calling the first number a miscalibrated 18-month forecast is wrong. It was never an 18-month forecast. It was an infinite-horizon extrapolation being judged against a finite event.
Implications for CLV, churn dashboards, and valuation
This matters because business systems often use P(alive) as if it were a current customer flag.
For CLV, the damage depends on how the model is used. If expected transactions over a stated horizon are used directly, the forecast may be serviceable and testable. If per-customer value is weighted by summed P(alive), the resulting customer equity estimate can move with modeling choices that customers never validated.
For churn dashboards, the problem is labeling. A dashboard that reports "active customers" based on summed P(alive) may be reporting an infinite-horizon model estimate, not active buyers in any operating sense. That can make retention teams think they have a larger recoverable base than the finite-horizon evidence supports.
For customer-based corporate valuation, the issue is sharper. Valuation work often turns customer counts, retention, purchase rates, and margins into enterprise value. If the alive-customer count can swing this much while near-term purchase forecasts barely move, then a valuation based on that count is carrying hidden model risk. The paper does not say customer-based valuation is wrong. It says one popular ingredient is less identified than it looks.
A practitioner audit: use horizons, not mystique
The fix is practical and somewhat unglamorous. Report the forecast at a horizon, then score it.
A useful audit would include:
- Pick operating horizons such as 6, 12, and 18 months, then report expected repeat buyers at each horizon.
- Backtest those forecasts by scoring date, cohort, acquisition channel, and customer age.
- Compare calibration across horizons, not just rank ordering.
- Recalibrate when cohorts drift, especially after pricing, product, or marketing mix changes.
- If management insists on a total alive-customer count, show an interval with realized returners as a lower bound, not a false-precise point estimate.
The point is not to throw away P(alive). It is to stop using an infinite-horizon quantity as if it were an observed customer count. Horizon counts are less grand, but they can be audited.
Why we could not backtest this on our data
We could not build a trading backtest from this paper on our data. The reason is plain. The signal needs customer-level transaction histories: purchase dates, recency, frequency, repeat timing, cohorts, and enough later observation to score forecasts. Our platform has market data, filings, transcripts, press releases, and related public information. It does not have firm-by-firm order-level customer panels.
Public companies also report customer metrics inconsistently. Some disclose active customers, some disclose buyers, some disclose accounts, and definitions change. Without the underlying transactions, we cannot estimate BG/NBD or Pareto/NBD models, audit P(alive), or tell which firms are overstating their recoverable customer base.
That does not weaken the paper's operating point. It only limits direct public-market implementation.
Why this is hard to turn into a public-market signal
A tradable version would need to identify companies where reported customer value depends on inflated or poorly auditable alive-customer estimates. That requires two things investors usually do not have: the company's transaction panel and the model convention used to convert silence into aliveness.
You might look for proxies, such as rising disclosed active customers with flat orders, weakening repeat rates, or changes in KPI definitions. Those are useful red flags, not a clean signal. They mix marketing spend, seasonality, product cadence, price changes, and disclosure choices.
The best use case is inside the firm, or in diligence where the buyer can request raw transactions. Ask for 12-month and 18-month repeat-buyer forecasts by cohort. Ask for calibration plots by scoring date. Ask how much the alive-customer count changes under reasonable model settings. If the forecasted near-term buyers stay put while the claimed customer base moves all over the place, you have found the paper's problem in the wild.