[論文レビュー] Picking Winners: A Data Driven Approach to Evaluating the Quality of Startup Companies
本論文は、創設者、業界、投資家特徴を用いて、ブラウン運動の最初の通過時間に基づくデータ駆動型確率的モデルを提案し、スタートアップの質を評価する。ベイズ推論とグリーディーな組み合わせ最適化を用いて、最大60%の排出率を達成するポートフォリオを構築し、トップVCファンドのほぼ2倍の性能を示し、高リターンのスタートアップ成果を予測する上で顕著な効果を発揮した。
We consider the problem of evaluating the quality of startup companies. This can be quite challenging due to the rarity of successful startup companies and the complexity of factors which impact such success. In this work we collect data on tens of thousands of startup companies, their performance, the backgrounds of their founders, and their investors. We develop a novel model for the success of a startup company based on the first passage time of a Brownian motion. The drift and diffusion of the Brownian motion associated with a startup company are a function of features based its sector, founders, and initial investors. All features are calculated using our massive dataset. Using a Bayesian approach, we are able to obtain quantitative insights about the features of successful startup companies from our model. To test the performance of our model, we use it to build a portfolio of companies where the goal is to maximize the probability of having at least one company achieve an exit (IPO or acquisition), which we refer to as winning. This $\ extit{picking winners}$ framework is very general and can be used to model many problems with low probability, high reward outcomes, such as pharmaceutical companies choosing drugs to develop or studios selecting movies to produce. We frame the construction of a picking winners portfolio as a combinatorial optimization problem and show that a greedy solution has strong performance guarantees. We apply the picking winners framework to the problem of choosing a portfolio of startup companies. Using our model for the exit probabilities, we are able to construct out of sample portfolios which achieve exit rates as high as 60%, which is nearly double that of top venture capital firms.
研究の動機と目的
- 初期段階のデータが限られる中で、定量的かつデータ駆動型のフレームワークを構築すること。
- 創設者、業界、投資家特徴から導かれるドリフトおよび拡散パラメータを有するブラウン運動を用いて、スタートアップの進化を最初の通過時間問題としてモデル化すること。
- 少なくとも1つの成功した排出(IPOまたは買収)の確率を最大化するポートフォリオ最適化戦略を構築すること(「勝者を選ぶ」として定義)。
- アウトオブサンプルのポートフォリオを用いてモデルを検証し、トップベンチャーキャピタルのパフォーマンスと比較すること。
- 公に入手可能なデータを用いて、機関投資家および個人投資家がスケーラブルで原則に基づいた方法で初期段階のスタートアップを選定できるようにすること。
提案手法
- スタートアップの成功を、時間に依存しないブラウン運動の最初の通過時間としてモデル化し、ドリフトおよび拡散がスタートアップ特徴の関数として定義する。
- ベイズ階層モデルを用いてブラウン運動のパラメータを推定し、排出確率予測における不確実性を組み込む。
- 創設者背景、業界、初期投資家情報を含む数万件のスタートアップの大きなデータセットから排出確率推定値を構築する。
- 選択されたスタートアップのうち少なくとも1つが排出する確率を最大化するという目的で、ポートフォリオ構築を組み合わせ最適化問題として定式化する。
- 理論的性能保証が強く、高インパクトのスタートアップポートフォリオを選択できるグリーディーなアルゴリズムを適用する。
- 複数年の期間(2011–2012年)にわたり、異分散性を持つ独立および相関モデル、およびロバストなバージョンを用いてモデルをテストする。
実験結果
リサーチクエスチョン
- RQ1最初の通過時間に基づくブラウン運動の確率的モデルは、創設者、業界、投資家データのみを用いて、スタートアップの排出結果を効果的に予測できるか?
- RQ2モデルの仮定(等分散性対非等分散性、独立対相関)が、排出確率予測の正確性およびロバスト性に与える影響は何か?
- RQ3グリーディーな組み合わせ最適化アプローチは、少なくとも1つの排出の確率を最大化するスタートアップポートフォリオ選択において、強力な性能保証を達成できるか?
- RQ4公に入手可能なデータのみを用いて、モデルはトップクラスのベンチャーキャピタルファンドをどの程度上回る排出率を達成できるか?
- RQ5本モデルは、スタートアップに限らず、ドラッグ開発や映画製作など、低確率・高リターンの選択問題へ一般化可能か?
主な発見
- モデルは、アウトオブサンプルのスタートアップポートフォリオを構築し、最大60%の排出率を達成し、通常30%程度の排出率を達成するトップVCファンドを著しく上回った。
- ロバストな非等分散相関モデルを用いた2012年のポートフォリオは、累積目的値が1.0に達し、上位選択された企業のうち少なくとも1つが排出する確率が100%であることを示した。
- グリーディー最適化アルゴリズムは強力な性能保証を達成し、ランダムまたはヒューリスティックな代替手法よりも常に高い排出確率のポートフォリオを一貫して選択した。
- モデルの排出確率推定値は、創設者および投資家の質に非常に感受的であり、トップクラスの創設者および投資家はドリフトパラメータを上昇させ、早期の排出確率を高める要因となった。
- 非等分散相関モデルは、特に高い不確実性下でのスタートアップ成功の真の分散および相関構造を捉える点で、単純なモデルを上回った。
- フレームワークは、後に買収または大規模な資金調達ラウンドを達成した、Struq、Funzio、Metaresolverなどの高潜在力スタートアップを効果的に同定し、モデルの予測力の妥当性を検証した。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。