Skip to main content
QUICK REVIEW

[Paper Review] Accurate Chemistry Collection: Coupled cluster atomization energies for broad chemical space

Sebastian Ehlert, Jan Hermann|ArXiv.org|Jun 17, 2025
Machine Learning in Materials Science3 citations
TL;DR

The paper presents MSR-ACC/TAE25, a large CCSD(T)/CBS-based dataset of 76,879 total atomization energies covering broad chemical space up to argon, created to enable data-driven thermochemistry methods with sub-chemical accuracy.

ABSTRACT

Accurate thermochemical data with sub-chemical accuracy (within 1 kcal mol$^{-1}$ of the empirical ground truth) are essential for advancing computational chemistry methods. However, existing datasets that reach this level of accuracy remain limited in size or scope. This hinders the development of data-driven methods with predictive accuracy across the broad chemical space of closed-shell, neutral molecules. Here we present Microsoft Research Accurate Chemistry Collection (MSR-ACC) and its first release, MSR-ACC/TAE25, comprising 73,040 total atomization energies at the CCSD(T)/CBS level obtained with the W1-F12 thermochemical protocol. The dataset is constructed to exhaustively cover the chemical space of closed-shell, charge-neutral, covalently bound equilibrium molecular structures containing up to 5 non-hydrogen atoms drawn from elements up to argon and lacking significant multireference character. The dataset and its canonical train and validation splits are openly available on Zenodo in the QCSchema format under the CDLA Permissive 2.0 license. This first release of MSR-ACC enables data-driven approaches for developing predictive computational chemistry methods with unprecedented accuracy and scope.

Motivation & Objective

  • Provide sub-chemical-accuracy TAE data to benchmark and train computational methods.
  • Exhaustively cover chemical space for elements up to argon without bias toward common subspaces.
  • Enable data-driven approaches (ML, DFT, semi-empirical) with unprecedented scope and accuracy.
  • Filter out systems with significant multireference character or triplet ground states to ensure CCSD(T)-based labeling.

Proposed method

  • Generate exhaustive molecular graphs for up to five non-hydrogen atoms using three graph-generation strategies (combinatorial enumeration, degree-sequence sampling, and an autoregressive GPT-2 based model).
  • Optimize structures through a multi-step protocol: UFF → GFN2-xTB sampling → r2SCAN-3c → B3LYP-D3(BJ)/def2-TZVPP.
  • Label TAEs at the W1-F12 CCSD(T)/CBS level with Hartree–Fock extrapolated to CBS, CCSD-F12 energies, and (T) corrections.
  • Apply filtering criteria: exclude %TAE[(T)] > 6% and S0–T1 gaps from positive values to ensure single-reference character.
  • Provide data records in Zenodo with QCSchema formatting and extras including W1-F12 TAE components.

Experimental results

Research questions

  • RQ1How can one achieve a broad, bias-free coverage of chemical space for TAEs up to argon with CCSD(T)-level accuracy?
  • RQ2What proportion and characteristics of molecules exhibit significant post-CCSD(T) contributions, and how can reliable labeling be ensured?
  • RQ3Can a large, openly accessible TAE dataset enable robust development of ML and DFA methods with sub-chemical accuracy across diverse chemistries?
  • RQ4What quality controls (e.g., singlet-triplet gaps, multireference diagnostics) effectively filter problematic species without excluding valid single-reference systems?

Key findings

  • MSR-ACC/TAE25 contains 76,879 charge-neutral closed-shell TAEs labeled at CCSD(T)/CBS via the W1-F12 protocol.
  • The dataset spans elements up to argon with up to five non-hydrogen atoms and is not dominated by nondynamical correlation.
  • Filtering using %TAE[(T)]>6% and positive S0–T1 gaps removes multireference/ triplet-containing species to ensure single-reference labeling.
  • W1-F12 TAEs show expected component distributions across HF, CCSD, (T), and CV contributions, with TAE values ranging across a broad spectrum.
  • Data records are released with training/validation splits and supplementary W1-F12 energy components for machine-learning applications.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.