Skip to main content
QUICK REVIEW

[Paper Review] Distributional Fitting and Tail Analysis of Lead-Time Compositions: Nights vs. Revenue on Airbnb

Harrison Katz, Jess Needleman|arXiv (Cornell University)|Jan 17, 2026
Sharing Economy and Platforms0 citations
TL;DR

The paper analyzes daily lead-time compositions for Nights Booked and Gross Booking Value (revenue) on Airbnb as compositional vectors, finding mid-range lead times dominate GBV, Gamma/Weibull fit well, and tail inference is sensitive to truncation.

ABSTRACT

We analyze daily lead-time distributions for two Airbnb demand metrics, Nights Booked (volume) and Gross Booking Value (revenue), treating each day's allocation across 0-365 days as a compositional vector. The data span 2,557 days from January 2019 through December 2025 in a large North American region. Three findings emerge. First, GBV concentrates more heavily in mid-range horizons: beyond 90 days, GBV tail mass typically exceeds Nights by 20-50%, with ratios reaching 75% at the 180-day threshold during peak seasons. Second, Gamma and Weibull distributions fit comparably well under interval-censored cross-entropy. Gamma wins on 61% of days for Nights and 52% for GBV, with Weibull close behind at 38% and 45%. Lognormal rarely wins (<3%). Nonparametric GAMs achieve 18-80x lower CRPS but sacrifice interpretability. Third, generalized Pareto fits suggest bounded tails for both metrics at thresholds below 150 days, though this may partly reflect right-truncation at 365 days; above 150 days, estimates destabilize. Bai-Perron tests with HAC standard errors identify five structural breaks in the Wasserstein distance series, with early breaks coinciding with COVID-19 disruptions. The results show that volume and revenue lead-time shapes diverge systematically, that simple two-parameter distributions capture daily pmfs adequately, and that tail inference requires care near truncation boundaries.

Motivation & Objective

  • Investigate whether volume (Nights) and revenue (GBV) lead-time distributions differ on Airbnb.
  • Identify the parametric family that best fits daily lead-time pmfs under interval censoring.
  • Assess tail behavior and potential structural changes in the lead-time distributions across time.
  • Evaluate nonparametric and parametric approaches for fitting lead-time distributions and compare forecast-relevant scoring.
  • Examine implications for revenue forecasting and tail inference under truncation constraints.

Proposed method

  • Treat each day’s lead-time allocation as a compositional vector on the 365-simplex.
  • Use Wasserstein-1 distance to compare daily Nights vs. GBV pmfs.
  • Fit Gamma, Weibull, and Lognormal distributions to interval-censored pmfs via cross-entropy.
  • Apply generalized Pareto tail analysis with threshold stability diagnostics (POT approach).
  • Employ HAC-robust Bai–Perron structural-break tests on the Wasserstein distance series.
  • Compare nonparametric GAM fits using CRPS and KLD as scoring rules.
Figure 1: Aggregated lead-time distributions for Nights (blue) and GBV (red), 2019–2025. Distributions are day-weighted averages of daily pmfs: $\bar{p}(\ell)=D^{-1}\sum_{d}x_{d,\ell}$ . Both peak near $\ell=0$ and decline rapidly. The curves cross around $\ell=30$ days: below, Nights slightly excee
Figure 1: Aggregated lead-time distributions for Nights (blue) and GBV (red), 2019–2025. Distributions are day-weighted averages of daily pmfs: $\bar{p}(\ell)=D^{-1}\sum_{d}x_{d,\ell}$ . Both peak near $\ell=0$ and decline rapidly. The curves cross around $\ell=30$ days: below, Nights slightly excee

Experimental results

Research questions

  • RQ1Do volume-based and revenue-based lead-time distributions differ systematically across days?
  • RQ2Which parametric family (Gamma, Weibull, Lognormal) best describes daily lead-time pmfs under interval censoring?
  • RQ3How do tail behaviors differ between Nights and GBV, and are tails stable across thresholds given truncation?
  • RQ4Do structural breaks occur in the divergence between Nights and GBV lead-time shapes over the sample?
  • RQ5What is the relative performance of parametric vs. nonparametric (GAM) fits for descriptive vs. forecast-prone purposes?

Key findings

  • GBV concentrates more in mid-range horizons than Nights beyond 90 days, with ratios up to 75% at 180 days during peak seasons.
  • Gamma and Weibull provide comparable fits; Gamma wins on ~55–60% of days for Nights and ~52% for GBV by cross-entropy.
  • Lognormal rarely wins; GAMs achieve much lower CRPS but with interpretability trade-offs.
  • GPD tail estimates are stable only up to ~150 days due to truncation; beyond that inference destabilizes.
  • Five structural breaks in the Wasserstein distance suggest COVID-related shifts and subsequent regime changes in divergence.
  • Two-parameter Gamma/Weibull suffice for descriptive lead-time pmfs; GAM offers better in-sample CRPS but less parsimonious.
Figure 2: Daily Wasserstein-1 distance between Nights and GBV, with structural breakpoints (dashed vertical lines) identified via Bai–Perron with HAC standard errors. The series averages 8.67 (95% CI: 8.42–8.92). Seasonality peaks in summer. Early breaks (2020–2021) align with COVID disruptions; lat
Figure 2: Daily Wasserstein-1 distance between Nights and GBV, with structural breakpoints (dashed vertical lines) identified via Bai–Perron with HAC standard errors. The series averages 8.67 (95% CI: 8.42–8.92). Seasonality peaks in summer. Early breaks (2020–2021) align with COVID disruptions; lat

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.