Life sciences · Preprint
arXiv · September 4, 2026
Early or partial results. Treat as a signal, not a conclusion.
This preprint proposes two Hessian-derived data augmentation schemes (UniAug and ModeAug) to improve machine-learning interatomic potentials without modifying training architectures. The work is computational and has not undergone peer review; it addresses a technical problem in molecular simulation rather than a clinical or biological outcome.
Preprint. Intervention: Hessian-derived data augmentation via isotropic Gaussian displacement (UniAug) and normal mode-weighted displacement (ModeAug) applied to machine-learning interatomic potential training..
Both UniAug and ModeAug achieve effective Hessian augmentation using simple Taylor expansions without computational overhead from higher-order backpropagation. Methods integrate seamlessly into existing MLIP architectures as plug-and-play augmentation schemes. Comprehensive evaluation across non-equilibrium and equilibrium datasets demonstrates enhanced model accuracy.
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
The source did not state who this applies to in practice.
This is a methodological proposal for improving machine-learning interatomic potentials through data augmentation, demonstrated on benchmark datasets but lacking clinical or real-world validation relevant to life sciences.
As stated by the source record.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
What is missing. This record has no reported figures. That is a gap in the analysis, not a judgement about the study.
While machine-learning interatomic potentials (MLIPs) have successfully learned potential energy surfaces (PES) and atomic forces, many practical applications, such as vibrational analysis and transition state search, rely heavily on the PES Hessian. Yet, standard MLIPs tend to be trained on energy and forces alone, leaving Hessian information largely unexploited. Meanwhile, existing methods that explicitly incorporate the Hessian into training objectives require architectural modifications and introduce significant computational and memory overheads due to higher-order backpropagation. To address these limitations, we propose two Hessian-derived data augmentation schemes: isotropic Gaussian displacement (\textbf{UniAug}) and normal mode-weighted displacement (\textbf{ModeAug}). Both methods utilize simple Taylor expansions, achieving effective augmentation without altering training objectives or extending the autograd graph. This allows seamless, plug-and-play integration with existing architectures and training pipelines. Comprehensive evaluations across non-equilibrium and equilibrium datasets demonstrate that our approach enhances model accuracy while providing practical, task-specific guidelines.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.