Life sciences · Preprint
arXiv · September 9, 2026
Posted before peer review. The findings may change or fail to hold.
This preprint describes IBIB (IB2), a protocol for measuring enterprise AI system capability by serving route rather than model identifier alone, acknowledging that deployment performance depends on weights, serving configuration, precision, output contract, and harness. Empirical testing across eleven systems demonstrates that capability is not reliably indicated by advertised model identifiers and that serving-arm choice and reliability inclusion materially affect reported performance metrics.
Methodological protocol with empirical validation across multiple systems. Eleven enterprise AI systems; specific model identifiers, vendors, and deployment contexts not disclosed in abstract.. Intervention: IBIB (IB2) measurement protocol, including capability-binding preflight, reliability-inclusive scoring, and score-blind adjudication.. Compared with: Conventional model-identifier-based benchmarking (18 audited benchmarks described as scoring advertised model identifiers only)..
Two complete single-route runs on identical weights later failed distinct predicates of the finalized binding gate, while a third passed that gate before a fresh run, showing capability availability is measurable but not stable across runs. Serving-arm choice moved one declared revision and precision from 77.38 to 82.54, paired interval [0.11, 10.60], indicating that serving configuration substantively affects performance. Four of seven suites saturate under a six-system band, with spread almost entirely from governed database work and multi-tab joins, suggesting non-uniform discrimination across task categories.
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
The source did not state who this applies to in practice.
This is a methodological protocol paper on arXiv describing a measurement framework for enterprise AI systems; it reports empirical results from eleven systems but is not peer-reviewed and does not present a clinical or comparative efficacy claim.
As stated by the source record.
Quoted from the source exactly as published.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
Enterprises deploy systems, not checkpoints. Usable capability depends jointly on weights, serving route, precision, output contract, and harness, yet all 18 audited benchmarks score advertised model identifiers. We treat this as measurement error and give a protocol that makes it reportable. It has three parts. A gold-blind capability-binding preflight verifies that a route can execute the evaluation contract before any task reaches it; a reliability-inclusive first-pass scoring rule keeps failure in the score while keeping unsupported capability out; and adjudication is structurally score-blind. We call the protocol IB2 and release its algorithms, classification tables, request contract, and manifest schemas. Its reference instantiation, 128 locked tasks and 987 assertions over document, spreadsheet, chart, tool and database work, stays sealed: the procedure is the artifact, not the corpus. Across eleven systems, four results. Capability availability is measurable: two complete single-route runs on identical weights later failed distinct predicates of the finalized binding gate, while a third passed that gate before a fresh run. The advertised identifier exposed neither limit. Discrimination is not uniform: four of seven suites saturate under a six-system band, with the spread almost entirely from governed database work and multi-tab joins, so we report interval-backed resolution groups, not ranks; two of the nominal five-label output's four cuts fail multiplicity adjustment. Serving-arm choice moved one declared revision and precision from 77.38 to 82.54, paired interval [0.11,10.60], though the arms differ in access mode, harness generation, and the serving tool-call parser, and harness generation is a property of our evaluator, not any endpoint. Excluding failed responses from denominators changes the point ordering, so reliability inclusion changes a conclusion, not its wording.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.