[Paper Review] Distributional Fitting and Tail Analysis of Lead-Time Compositions: Nights vs. Revenue on Airbnb
The paper analyzes daily lead-time compositions for Nights Booked and Gross Booking Value (revenue) on Airbnb as compositional vectors, finding mid-range lead times dominate GBV, Gamma/Weibull fit well, and tail inference is sensitive to truncation.
We analyze daily lead-time distributions for two Airbnb demand metrics, Nights Booked (volume) and Gross Booking Value (revenue), treating each day's allocation across 0-365 days as a compositional vector. The data span 2,557 days from January 2019 through December 2025 in a large North American region. Three findings emerge. First, GBV concentrates more heavily in mid-range horizons: beyond 90 days, GBV tail mass typically exceeds Nights by 20-50%, with ratios reaching 75% at the 180-day threshold during peak seasons. Second, Gamma and Weibull distributions fit comparably well under interval-censored cross-entropy. Gamma wins on 61% of days for Nights and 52% for GBV, with Weibull close behind at 38% and 45%. Lognormal rarely wins (<3%). Nonparametric GAMs achieve 18-80x lower CRPS but sacrifice interpretability. Third, generalized Pareto fits suggest bounded tails for both metrics at thresholds below 150 days, though this may partly reflect right-truncation at 365 days; above 150 days, estimates destabilize. Bai-Perron tests with HAC standard errors identify five structural breaks in the Wasserstein distance series, with early breaks coinciding with COVID-19 disruptions. The results show that volume and revenue lead-time shapes diverge systematically, that simple two-parameter distributions capture daily pmfs adequately, and that tail inference requires care near truncation boundaries.
Motivation & Objective
- Investigate whether volume (Nights) and revenue (GBV) lead-time distributions differ on Airbnb.
- Identify the parametric family that best fits daily lead-time pmfs under interval censoring.
- Assess tail behavior and potential structural changes in the lead-time distributions across time.
- Evaluate nonparametric and parametric approaches for fitting lead-time distributions and compare forecast-relevant scoring.
- Examine implications for revenue forecasting and tail inference under truncation constraints.
Proposed method
- Treat each day’s lead-time allocation as a compositional vector on the 365-simplex.
- Use Wasserstein-1 distance to compare daily Nights vs. GBV pmfs.
- Fit Gamma, Weibull, and Lognormal distributions to interval-censored pmfs via cross-entropy.
- Apply generalized Pareto tail analysis with threshold stability diagnostics (POT approach).
- Employ HAC-robust Bai–Perron structural-break tests on the Wasserstein distance series.
- Compare nonparametric GAM fits using CRPS and KLD as scoring rules.

Experimental results
Research questions
- RQ1Do volume-based and revenue-based lead-time distributions differ systematically across days?
- RQ2Which parametric family (Gamma, Weibull, Lognormal) best describes daily lead-time pmfs under interval censoring?
- RQ3How do tail behaviors differ between Nights and GBV, and are tails stable across thresholds given truncation?
- RQ4Do structural breaks occur in the divergence between Nights and GBV lead-time shapes over the sample?
- RQ5What is the relative performance of parametric vs. nonparametric (GAM) fits for descriptive vs. forecast-prone purposes?
Key findings
- GBV concentrates more in mid-range horizons than Nights beyond 90 days, with ratios up to 75% at 180 days during peak seasons.
- Gamma and Weibull provide comparable fits; Gamma wins on ~55–60% of days for Nights and ~52% for GBV by cross-entropy.
- Lognormal rarely wins; GAMs achieve much lower CRPS but with interpretability trade-offs.
- GPD tail estimates are stable only up to ~150 days due to truncation; beyond that inference destabilizes.
- Five structural breaks in the Wasserstein distance suggest COVID-related shifts and subsequent regime changes in divergence.
- Two-parameter Gamma/Weibull suffice for descriptive lead-time pmfs; GAM offers better in-sample CRPS but less parsimonious.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.