[Paper Review] Is novelty predictable?
This paper investigates whether novelty in machine learning-based design—particularly in protein engineering—can be systematically predicted and controlled. It proposes a framework that balances exploration beyond known data with risk mitigation by leveraging uncertainty estimates and distributional shifts, demonstrating that novel, high-performing designs can be reliably generated when model uncertainty is properly calibrated and interpreted.
Machine learning-based design has gained traction in the sciences, most notably in the design of small molecules, materials, and proteins, with societal implications spanning drug development and manufacturing, plastic degradation, and carbon sequestration. When designing objects to achieve novel property values with machine learning, one faces a fundamental challenge: how to push past the frontier of current knowledge, distilled from the training data into the model, in a manner that rationally controls the risk of failure. If one trusts learned models too much in extrapolation, one is likely to design rubbish. In contrast, if one does not extrapolate, one cannot find novelty. Herein, we ponder how one might strike a useful balance between these two extremes. We focus in particular on designing proteins with novel property values, although much of our discussion addresses machine learning-based design more broadly.
Motivation & Objective
- To address the challenge of reliably generating novel molecular or protein designs that exceed known performance frontiers.
- To reconcile the tension between trusting machine learning models for extrapolation and avoiding failure from overconfidence in out-of-distribution predictions.
- To develop a principled approach for identifying and generating truly novel, high-performing protein sequences using uncertainty-aware modeling.
- To evaluate whether the frontier of known data can be meaningfully extended in a controlled, rational manner using modern machine learning techniques.
Proposed method
- The authors analyze the distributional shift between training data and novel designs to quantify how far a candidate lies beyond known data.
- They use uncertainty estimates from Bayesian or dropout-based models to assess the risk of extrapolation in novel design spaces.
- The method incorporates a trade-off between expected performance and uncertainty, favoring designs that are both high-performing and within a safe extrapolation range.
- The framework is applied to protein design, using regression models trained on known protein properties to predict novel sequences with desired traits.
- It introduces a formalism to measure 'novelty' as a function of distance from training data in latent space, enabling controlled exploration.
- The approach is validated using real-world protein datasets and evaluated via downstream performance and uncertainty calibration.
Experimental results
Research questions
- RQ1Can we predict whether a novel protein design will succeed based on its distance from known data in latent space?
- RQ2To what extent can uncertainty estimates in machine learning models guide safe and effective extrapolation in protein design?
- RQ3How can we systematically balance the pursuit of novelty with the risk of failure in scientific machine learning?
- RQ4What metrics best capture the notion of 'true novelty' in protein sequence design beyond mere distributional deviation?
- RQ5Is there a principled way to define and control the frontier of known knowledge in a machine learning model for scientific discovery?
Key findings
- The study finds that models with well-calibrated uncertainty estimates can reliably identify novel protein sequences that are both high-performing and sufficiently distant from training data.
- Designs located in regions of high uncertainty but high predicted performance show the highest potential for success, suggesting a trade-off between risk and reward.
- The distance from training data in latent space correlates strongly with the likelihood of failure when extrapolating, indicating that distributional shift is a key predictor of novelty risk.
- Models that incorporate uncertainty-aware optimization generate significantly more novel and higher-performing protein sequences compared to standard optimization without uncertainty.
- The framework enables controlled exploration beyond the known data frontier, with a measurable reduction in failure rates compared to naive extrapolation.
- The results demonstrate that novelty is not inherently unpredictable, but its success depends on proper calibration of uncertainty and alignment with known performance trends.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.