Life sciences · Preprint
arXiv · September 9, 2026
Early or partial results. Treat as a signal, not a conclusion.
This preprint systematically documents previously unreported biases in major protein-protein interaction databases (HIPPIE, IntAct, STRING, PDB-derived) and proposes an optimization-based pipeline to detect and mitigate them. The work identifies multiple classes of shortcuts (topological, self-interaction, taxonomic, functional) that machine learning models may exploit instead of learning biological signal, but the proposed pipeline's impact on actual ML prediction accuracy is not empirically demonstrated.
Methodological audit with computational tool development. Protein-protein interaction data from HIPPIE, IntAct, STRING databases and 3D-structural data from the Protein Data Bank (PDB). Intervention: Optimization-based dataset splitting and negative sampling pipeline formulated as integer linear programs. Compared with: Standard random data splitting and existing negative sampling approaches.
Random data splitting introduces strong topological shortcuts that advantage models trained and tested on overlapping protein sets Removing train-test protein overlap eliminates topological shortcuts but retains usable shortcuts from self-interactions, taxonomic identity, and functional relatedness The prevalence of retained shortcuts (self-interaction, taxonomic, functional) depends on the data source (HIPPIE, IntAct, STRING, PDB)
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
This work is primarily relevant to machine learning researchers and bioinformaticians developing PPI prediction models. Clinicians and bench researchers using PPI databases should be aware that these resources contain systematic biases that may distort which interactions are well-characterized and that models trained on them may not reflect true biological mechanisms.
A methodological audit and tool development study identifying and proposing mitigation strategies for biases in PPI datasets; important for ML validation but lacks empirical validation of the proposed pipeline's effectiveness on downstream prediction tasks.
As stated by the source record.
This work is primarily relevant to machine learning researchers and bioinformaticians developing PPI prediction models. Clinicians and bench researchers using PPI databases should be aware that these resources contain systematic biases that may distort which interactions are well-characterized and that models trained on them may not reflect true biological mechanisms.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
What is missing. This record has no reported figures. That is a gap in the analysis, not a judgement about the study.
Protein-protein interaction (PPI) databases do not faithfully reflect biological realities. Instead, they are influenced by study and technical biases that distort certain protein and interaction attributes. Machine learning models can exploit these as learning shortcuts if the negative dataset is not constructed with care. So far, the shortcuts introduced during PPI dataset construction have only been examined in isolation. Here, we systematically characterize both reported and, to our knowledge, previously unreported biases in PPI datasets that lead machine learning models to learn shortcuts instead of biological signal. We analyze HIPPIE, IntAct, and STRING, dedicated PPI databases, as well as two datasets derived from 3D-structural information in the Protein Data Bank (PDB). We show that random data splitting introduces strong topological shortcuts. When train-test protein overlap is removed, the resulting datasets still retain usable shortcuts stemming from self-interactions, taxonomic identity, and functional relatedness, whose prevalence interestingly depends on the data source. We further show that sampling negatives from a set of high-confidence non-interactors, an intuitively appealing choice, can amplify the shortcut stemming from functional relatedness. To detect and mitigate these biases, we provide an open Nextflow pipeline that combines similarity-aware, data-loss-minimizing dataset splitting with bias-minimizing negative sampling, both formulated as integer linear programs. Its key concept of quantifying biases to minimize them through optimization-based negative sampling can, in principle, be extended to any machine learning problem where the pool of negative candidates is much larger than the positives and is thus of interest also beyond PPI prediction.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.