[Paper Review] A Guide to Comparing the Performance of VA Algorithms
This paper proposes a standardized framework for fairly comparing verbal autopsy (VA) algorithms by isolating the effects of algorithm logic and symptom-cause information (SCI), advocating for a publicly accessible, WHO-managed SCI archive to ensure consistent, generalizable, and transparent performance evaluation across regions and time. The key contribution is a methodological shift toward SCI-controlled comparisons to resolve longstanding inconsistencies in VA algorithm benchmarking.
The literature comparing the performance of algorithms for assigning cause of death using verbal autopsy data is fractious and does not reach a consensus on which algorithms perform best, or even how to do the comparison. This manuscript explains the challenges and suggests a way forward. A universal challenge is the lack of standard training and testing data. This limits meaningful comparisons between algorithms, and further, limits the ability of any algorithm to classify verbal autopsy deaths by cause in a way that is widely generalizable across regions and through time. Verbal autopsy algorithms utilize a variety of information to describe the relationship between verbal autopsy symptoms and causes of death - called symptom-cause information (SCI). A crowd sourced, public archive of SCI managed by the World Health Organization (WHO) is suggested as a way to address the lack of SCI for developing, testing, and comparing verbal autopsy coding algorithms, and additionally, as a way to ensure that algorithm-assigned causes of death are as accurate and comparable across regions and through time as possible.
Motivation & Objective
- To address the lack of consensus in VA algorithm comparisons due to confounded variables like differing SCI and algorithm logic.
- To resolve the challenge of non-standardized training and testing data limiting generalizability across regions and time.
- To advocate for a centralized, publicly accessible SCI archive based on de-biased physician-coded VAs to standardize symptom-cause relationships.
- To promote transparency and reproducibility by requiring access to source code and standardized reporting of algorithmic components.
- To establish a universal benchmarking framework using gold-standard data and consistent metrics for future VA algorithm evaluations.
Proposed method
- Use a controlled experimental design where VA data and SCI are held constant while only algorithm logic is varied to isolate its impact on performance.
- Develop a centralized, continuously updated SCI archive by pooling VA data with independently assigned causes from diverse global settings.
- Utilize de-biased physician-coded VAs (DBPCVAs) as a reliable source of SCI, ensuring each death is coded by multiple physicians and each physician codes many cases.
- Implement standardized metrics and reporting protocols for algorithm performance, including CSMF accuracy and cause assignment precision.
- Apply Bayes' rule as the core inference mechanism across most algorithms, except for the Tariff family, to ensure methodological consistency.
- Ensure ethical, legal, and technical feasibility through defined data governance, metadata standards, quality assurance, and infrastructure planning for the SCI archive.
Experimental results
Research questions
- RQ1How can the performance of VA algorithms be fairly compared when algorithm logic and SCI are confounded in existing studies?
- RQ2What is the relative impact of algorithm logic versus symptom-cause information (SCI) on cause-of-death assignment accuracy?
- RQ3Can a centralized, publicly accessible SCI archive improve the consistency, accuracy, and generalizability of VA algorithms across diverse populations and time periods?
- RQ4How can de-biased physician-coded VAs be used to create reliable, representative SCI for algorithm training and validation?
- RQ5What standardized metrics and benchmarking frameworks are needed to ensure transparency and comparability in future VA algorithm research?
Key findings
- The effects of symptom-cause information (SCI) on VA algorithm performance can far outweigh differences in algorithm logic, as demonstrated in SCI-controlled comparisons using PHMRC gold-standard data.
- Most published comparisons conflate the effects of algorithm logic and SCI, making it impossible to isolate the true impact of either component on performance.
- The McCormick et al. (2016a) study shows that SCI quality and representativeness are the dominant factors in algorithm accuracy, not the choice of algorithmic logic.
- A standardized, continuously updated SCI archive based on de-biased physician-coded VAs can replicate non-biased physician judgment and enable consistent, region- and time-generalizable cause assignment.
- The absence of gold-standard community death datasets limits the validity of current comparisons, which are often based on hospital deaths that are less representative of real-world VA applications.
- A WHO-managed SCI archive is a feasible and essential next step to standardize VA algorithm development, testing, and deployment globally.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.