Life sciences · Preprint
arXiv · September 9, 2026
Early or partial results. Treat as a signal, not a conclusion.
OmniMed-FL is a controlled systems benchmark comparing eight multimodal fusion strategies and federated averaging methods on synthetic chest radiograph–note pairs across simulated non-IID hospital clients. The authors explicitly note these are descriptive proxy comparisons, not estimates of diagnostic performance or deployment readiness; multimodal fusion outperformed unimodal approaches (macro-F1 0.956 vs 0.934 text, 0.664 images on synthetic data), but findings are not generalizable to real clinical deployment.
Controlled systems study; benchmark of federated learning architectures and multimodal fusion strategies. Proxy corpus: 3,000 public chest radiographs and 3,000 class-conditioned synthetic notes, class-matched (not patient-matched). Five-class classification task: Normal, Pneumonia, COVID-19, Pleural Effusion, Cardiomegaly.. Intervention: OmniMed-FL framework with eight multimodal fusion strategies and federated learning algorithms (FedAvg, FedProx, FedMME, SCAFFOLD-AdamW) across federated clients.. Compared with: Local-only training (baseline), unimodal (text-only and image-only) learning, and non-federated (centralized) learning on the same corpus.. n = 6,000.
FedProx achieved macro-F1 0.737±0.085 vs FedAvg 0.662±0.074 over 3–5 clients with severe skew (α=0.1) SCAFFOLD-AdamW adaptation scored 0.070±0.015, substantially lower than FedProx Multimodal fusion scored 0.956 on synthetic corpus vs 0.934 text-only and 0.664 image-only
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
This is a systems benchmark, not a clinical validation study. The findings on federated multimodal fusion architecture are not directly applicable to clinical deployment because notes are synthetic and not patient-linked. Clinicians and implementers should view this as evidence that federated learning and multimodal fusion are technically feasible within communication and privacy constraints, but not as proof of diagnostic accuracy or readiness for production use.
A controlled systems study using synthetic data and non-patient-matched pairing to benchmark federated learning architectures; descriptive proxy comparisons only, not diagnostic validation or deployment readiness.
As stated by the source record.
Quoted from the source exactly as published.
This is a systems benchmark, not a clinical validation study. The findings on federated multimodal fusion architecture are not directly applicable to clinical deployment because notes are synthetic and not patient-linked. Clinicians and implementers should view this as evidence that federated learning and multimodal fusion are technically feasible within communication and privacy constraints, but not as proof of diagnostic accuracy or readiness for production use.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
Simultaneous assessment of medical imaging and patient records is often required in clinical diagnosis. However, standard machine learning algorithms cannot analyze these data types together. Meanwhile, compliance with HIPAA and GDPR can constrain centralized aggregation of sensitive patient data. This leaves a crucial void of secure fusion of visual and textual context across distant networks. Thus, we present OmniMed-FL, a controlled systems study of multimodal federated learning for five-class clinical condition classification (Normal, Pneumonia, COVID-19, Pleural Effusion, Cardiomegaly). Our proxy corpus pairs 3,000 public chest radiographs with 3,000 class-conditioned synthetic notes, matched by class, not by patient. The framework benchmarks eight fusion strategies, three initializations, four missing-text imputation rules, and matched federated baselines under non-IID Dirichlet partitioning across 3 to 20 hospital clients. As all notes are synthetic and pairing is not patient-level, these are descriptive proxy comparisons, not estimates of diagnostic performance or deployment readiness. Within those limits with clients ($K=5$) and severe skew ($α=0.1$), local-only training achieves a macro-F1 score of 0.297, FedAvg achieves $0.662\pm0.074$, FedProx $0.737\pm0.085$, a matched FedMME-style one-shot ensemble $0.647\pm0.080$, and our SCAFFOLD-AdamW adaptation $0.070\pm0.015$, the 0.075 FedProx-FedAvg gap falling inside the wider of the two two-seed standard deviations. Over a $4\times3$ grid, label skew costs up to 0.27 F1 whereas a near-sevenfold client increase costs at most 0.10, while bidirectional volume grows linearly to 183.5 GiB at $K=20$. Multimodal fusion leads on both corpora, scoring 0.956 against 0.934 for text and 0.664 for images on the synthetic corpus and 0.906 against 0.880 and 0.737 on the radiograph corpus, for $2.3\times$ the model state of text alone.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.