Life sciences · Preprint
arXiv · September 3, 2026
Raises a question worth testing. It does not answer one.
FLY-EVAL++ is a proposed evaluation framework that applies deterministic verification and rubric-guided scoring to measure safety compliance, physical feasibility, and constraint satisfaction in LLM flight trajectory predictions. Applied to 66 models, it shows substantial variation in safety performance independent of predictive accuracy, indicating that standard accuracy metrics alone are insufficient for safety-critical applications.
Systematic benchmarking study with fixed rubric evaluation. Large language models (specific model names and versions not listed in abstract); no patient, clinical, or human subject involvement.. Intervention: FLY-EVAL++ evaluation framework applied to flight trajectory and attitude prediction tasks. n = 66.
Models with comparable predictive performance differ by more than 28 points in safety score Recurrent failures include safety violations under physically plausible predictions and instability in multi-step rollouts Safety compliance is the most discriminative dimension of model behavior across the 66 LLMs evaluated
No external validation against independent safety assessments or real-world flight data described. Models with comparable predictive performance differ by more than 28 points in safety score
The source did not state who this applies to in practice.
This is a methodological paper proposing and demonstrating a new evaluation framework for LLMs in safety-critical domains, not a clinical trial or empirical validation of a therapeutic or diagnostic intervention; it raises questions about LLM safety assessment rather than answering clinical questions.
As stated by the source record.
Quoted from the source exactly as published.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
Evaluating large language models (LLMs) in safety-critical, physics-governed environments requires more than accuracy-based metrics, because predictions that are numerically close to the ground truth can still violate operational constraints, combine fields in physically inconsistent ways, or fail to produce usable structured outputs. Existing evaluation protocols do not measure these failure modes reliably. We propose FLY-EVAL++, an evidence-driven evaluation protocol that combines deterministic verification of protocol compliance, physical feasibility, and safety constraints with fixed rubric-guided aggregation into interpretable multi-dimensional scores. We instantiate FLY-EVAL++ for Flight Trajectory and Attitude Prediction (FTAP) by extending the PilotBench setting with history-conditioned and multi-step prediction tasks. Across 66 LLMs, safety compliance is the most discriminative dimension of model behavior: models with comparable predictive performance differ by more than 28 points in safety score, and we observe recurrent failures including safety violations under physically plausible predictions and instability in multi-step rollouts. These results show that evaluation in safety-critical domains should measure constraint satisfaction and structured validity explicitly rather than rely on accuracy-centric reporting alone.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.