[论文解读] Picking Winners: A Data Driven Approach to Evaluating the Quality of Startup Companies
本文提出一种基于布朗运动首达时的、数据驱动的随机模型,利用创始人、行业和投资者特征来评估初创企业质量。通过贝叶斯推断与贪心组合优化,该模型构建的组合退出率最高可达60%,接近顶级风投机构的两倍,展现出在预测高回报初创企业结果方面的强劲表现。
We consider the problem of evaluating the quality of startup companies. This can be quite challenging due to the rarity of successful startup companies and the complexity of factors which impact such success. In this work we collect data on tens of thousands of startup companies, their performance, the backgrounds of their founders, and their investors. We develop a novel model for the success of a startup company based on the first passage time of a Brownian motion. The drift and diffusion of the Brownian motion associated with a startup company are a function of features based its sector, founders, and initial investors. All features are calculated using our massive dataset. Using a Bayesian approach, we are able to obtain quantitative insights about the features of successful startup companies from our model. To test the performance of our model, we use it to build a portfolio of companies where the goal is to maximize the probability of having at least one company achieve an exit (IPO or acquisition), which we refer to as winning. This $\ extit{picking winners}$ framework is very general and can be used to model many problems with low probability, high reward outcomes, such as pharmaceutical companies choosing drugs to develop or studios selecting movies to produce. We frame the construction of a picking winners portfolio as a combinatorial optimization problem and show that a greedy solution has strong performance guarantees. We apply the picking winners framework to the problem of choosing a portfolio of startup companies. Using our model for the exit probabilities, we are able to construct out of sample portfolios which achieve exit rates as high as 60%, which is nearly double that of top venture capital firms.
研究动机与目标
- 开发一种定量、数据驱动的框架,以在早期数据有限的情况下评估初创企业质量。
- 将初创企业的发展建模为首达时问题,使用具有漂移和扩散参数的时齐布朗运动,其参数由创始人、行业和投资者特征推导得出。
- 制定一种投资组合优化策略,以最大化至少一次成功退出(首次公开募股或并购)的概率,即‘选中赢家’。
- 通过样本外投资组合验证模型,并与顶级风投表现进行对比。
- 为机构投资者和零售投资者提供一种可扩展、有理论依据的方法,仅使用公开可获取的数据进行早期初创企业筛选。
提出的方法
- 将初创企业成功建模为首达时问题,即时间齐次布朗运动的首达时间,其中漂移和扩散参数为初创企业特征的函数。
- 采用贝叶斯分层模型估计布朗运动的参数,纳入对退出概率预测不确定性的建模。
- 基于包含数万家初创企业数据的大规模数据集,构建退出概率估计,数据包括创始人背景、行业和初始投资者信息。
- 将投资组合构建问题形式化为组合优化问题,以最大化所选初创企业中至少一次退出的概率。
- 应用具有强理论性能保证的贪心算法,选择高影响力初创企业投资组合。
- 在多个年份(2011–2012)测试模型,涵盖异方差独立与相关模型,以及鲁棒变体。
实验结果
研究问题
- RQ1基于布朗运动首达时的随机模型,仅使用创始人、行业和投资者数据,能否有效预测初创企业退出结果?
- RQ2不同的建模假设(同方差 vs. 异方差,独立 vs. 相关)如何影响退出概率预测的准确性与鲁棒性?
- RQ3贪心组合优化方法是否能在选择初创企业投资组合以最大化至少一次退出概率方面,实现强性能保证?
- RQ4在仅使用公开可获取数据的前提下,该模型在退出率方面能在多大程度上超越顶级风投机构?
- RQ5该模型能否推广至其他低概率、高回报的选择问题,如新药研发或影视制作?
主要发现
- 该模型构建的样本外初创企业投资组合,退出率最高达60%,显著优于顶级风投机构通常实现的约30%的退出率。
- 采用鲁棒的异方差相关模型,2012年投资组合的累积目标值达到1.0,表明在所选头部项目中至少一次退出的概率为100%。
- 贪心优化算法实现了强性能保证,且持续选择出的退出概率高于随机或启发式方法的替代方案。
- 该模型的退出概率估计对创始人和投资者质量高度敏感,顶级创始人和投资者显著提升了漂移参数,从而提高了早期退出的可能性。
- 异方差相关模型在捕捉初创企业成功的真实方差与相关性结构方面优于简单模型,尤其在高不确定性情境下表现更优。
- 该框架成功识别出如Struq、Funzio和Metaresolver等高潜力初创企业,这些企业后来均实现了并购或获得重要融资轮次,验证了模型的预测能力。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。