Vaccine Coverage and Hesitancy · Journal article
Journal of Medical Systems · August 8, 2026
Early or partial results. Treat as a signal, not a conclusion.
This technical audit demonstrates that reasoning-enabled LLMs (o4-mini and Gemini 2.5 Flash) can classify COVID-19 vaccine stance on social media with macro-F1 scores around 0.8, but identifies systematic reasoning failures in chain-of-thought summaries (target confusion, sarcasm misinterpretation, label–rationale incoherence) that may limit trustworthiness for public health surveillance applications. The study proposes explainability-readiness metrics beyond accuracy but provides no evidence of clinical or epidemiological utility.
Mixed-methods technical audit with quantitative performance evaluation and qualitative content analysis. 3,060 rehydrated COVID-19 vaccination tweets from X (formerly Twitter), pre-labelled by human annotators as positive, negative, or neutral stance; no human subjects or clinical population. Intervention: Zero-shot stance classification using o4-mini and Gemini 2.5 Flash LLMs with reasoning-intensive settings enabled. Compared with: Human-annotated stance labels (positive, negative, neutral) and comparison between o4-mini and Gemini model outputs. n = 3,060. Not specified; social media data source (X/Twitter) is global but study location of analysis and annotation not stated.
Both models reached macro-F1 around 0.8 at best-performing settings, with o4-mini accuracy 0.819 vs Gemini 0.799 (McNemar p = 0.0015) Δmacro-F1 between models = 0.020 (95% CI 0.008–0.032) Gemini returned reasoning summary for all tweets; o4-mini did so for 64.7%
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
This work does not assess clinical outcomes or decision-making impact. It is a technical audit of LLM reasoning transparency; clinicians and public health professionals should interpret findings as raising caution flags about the readiness of current reasoning LLMs for deployment in epidemiological surveillance without additional safeguards against systematic reasoning failure modes.
This is a single mixed-methods technical audit of LLM performance on a specific task (stance classification) with no clinical trial data, no patient outcomes, and no direct evidence of impact on clinical practice or public health decision-making.
As stated by the source record.
Quoted from the source exactly as published.
This work does not assess clinical outcomes or decision-making impact. It is a technical audit of LLM reasoning transparency; clinicians and public health professionals should interpret findings as raising caution flags about the readiness of current reasoning LLMs for deployment in epidemiological surveillance without additional safeguards against systematic reasoning failure modes.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
Abstract This mixed-methods study assessed whether reasoning-enabled large language models (LLMs) can classify stances towards COVID-19 vaccination on X (formerly Twitter) and whether model-generated chain-of-thought (CoT) summaries contain reasoning failures relevant to transparent and auditable public health applications. Zero-shot stance classification by o4-mini and Gemini 2.5 Flash (Gemini) was evaluated on 3,060 rehydrated COVID-19 vaccination tweets against human-annotated labels (positive, negative, neutral). We reported accuracy and macro-F1, measured CoT availability, and qualitatively analysed dual-error cases (tweets misclassified by both models) using Mayring’s content analysis guided by the FUTURE-AI framework. At each model’s best-performing setting, both models reached macro-F1 around 0.8, with o4-mini outperforming Gemini (accuracy 0.819 vs. 0.799, McNemar p = 0.0015; Δmacro-F1 = 0.020, 95% CI 0.008–0.032). Under the reasoning-intensive settings, CoT availability differed: Gemini returned a reasoning summary for all tweets, whereas o4-mini did so for 64.7%. Among 1,981 tweets with CoTs from both models, 295 (14.9%) were dual-errors; in 88.8%, both models produced the same wrong label, suggesting shared failure modes. Qualitatively, both models showed the same errors: target confusion (policy vs. vaccine), literal readings of sarcasm, and label–rationale mismatches, recurring across models despite their markedly different CoT lengths. Reasoning LLMs can therefore classify stance accurately, but their readiness for transparent public health applications depends on whether a CoT is available at all and whether it is coherent with the label it accompanies (label–rationale coherence). CoT availability, label–rationale coherence, and safeguards against systematic reasoning failures offer candidate explainability-readiness metrics, alongside accuracy, for trustworthy digital epidemiology.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.