Life sciences · Preprint
arXiv · September 5, 2026
Early or partial results. Treat as a signal, not a conclusion.
PhenoBench is a newly constructed, executable benchmark framework built on a deeply phenotyped cohort of over 13,000 participants, defining 90 clinically grounded prediction tasks across 15 domains and 26 input modalities. Comparative evaluation of tabular and language models showed modest improvements of foundation models over standard baselines (median R² gain 0.004), with language models showing task-specific capability gaps and inconsistent performance without cohort-specific fitting.
Benchmark framework study with comparative model evaluation on a deeply phenotyped observational cohort. Over 13,000 participants from the Human Phenotype Project who completed the initial visit. No eligibility criteria, demographic, or clinical characteristics reported.. Intervention: Evaluation of tabular foundation models and language models on prediction tasks derived from multimodal longitudinal phenotypic data. Compared with: Ridge regression baseline (for tabular models); standard task-specific models (for language models).
Benchmark comprises 90 clinically grounded tasks across 15 domains and 26 input modalities from over 13,000 participants Tabular foundation models improved on ridge baseline by a median of 0.004 R² (95% CI, 0.002–0.006) across 160 regression comparisons spanning 52 tasks Language models showed task-specific capability gaps and shared failures of scale, rarely surpassing models fitted on the same fields without cohort-specific fitting
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
This benchmark may enable standardized comparison of predictive models on multimodal phenotypic data, but does not yet establish which models or measurements should be used in clinical practice. The modest gains of foundation models and inconsistency of language models suggest that current pretrained approaches offer limited advantage over task-specific fitting in this setting.
A descriptive benchmark study presenting a novel evaluation framework and initial comparative results across multiple model classes, without reporting definitive clinical utility or practice-relevant outcomes.
As stated by the source record.
Quoted from the source exactly as published.
This benchmark may enable standardized comparison of predictive models on multimodal phenotypic data, but does not yet establish which models or measurements should be used in clinical practice. The modest gains of foundation models and inconsistency of language models suggest that current pretrained approaches offer limited advantage over task-specific fitting in this setting.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
Deeply phenotyped cohorts combine clinical, imaging, molecular, and wearable observations across timescales from seconds to years. This breadth can reveal which measurements inform which health-related questions, but heterogeneous analyses are not directly comparable. We present PhenoBench, an executable benchmark built around the Human Phenotype Project, in which more than 13,000 participants have completed the initial visit. Each question fixes the target, eligible population, timing, and allowed information; its evaluation contract specifies the split, metric, baseline, and claim boundary. The benchmark defines 90 clinically grounded tasks across 15 domains and 26 input modalities. Measurements showed question- and representation-dependent predictive value, including positive, near-zero, and negative changes in held-out performance relative to matched baselines. We used PhenoBench to evaluate emerging tabular foundation models across 160 matched regression comparisons spanning 52 tasks. These models ranked above standard task-specific models in aggregate but, averaged across the three pretrained models within each cell, improved on ridge by a median of only 0.004 $R^2$ (95% CI, 0.002--0.006). We then used the same cohort data and evaluation contracts to evaluate 14 language models, collectively covering 40 tasks spanning phenotype recovery, classification, follow-up forecasting, and participant ordering. Without cohort-specific fitting, language models made informative predictions on some tasks, but showed task-specific capability gaps, shared failures of scale, and rarely surpassed models fitted on the same fields. PhenoBench turns a multimodal longitudinal cohort into a versioned, auditable evaluation system where new questions, measurements, and models can be added without redefining existing comparisons.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.