Life sciences · Preprint
arXiv · September 10, 2026
Early or partial results. Treat as a signal, not a conclusion.
This preprint describes E-CONAN, a new multilingual benchmark for Arabic natural language inference composed of sentence pairs from four heterogeneous sources. Nine multilingual pretrained models and five LLMs were evaluated using zero-shot classification and fine-tuning; results claim E-CONAN offers broader assessment than existing datasets (XNLI, ArNLI) but without peer review or formal validation of this claim.
Benchmark dataset development with descriptive post-hoc model evaluation. Arabic sentence pairs from multiple sources: automatic translation, human-validated machine translation, hand-crafted materials from Arabic teaching resources, and news headlines containing rumors. Intervention: E-CONAN benchmark datasets (E-CONAN-2 for 2-way RTE; E-CONAN-3 for 3-way NLI). Compared with: Comparison datasets: ArNLI and XNLI; baseline model: MARBERT (Arabic-specific); comparison groups: multilingual pretrained models (9 models, zero-shot), LLMs (5 models).
E-CONAN comprises two datasets: E-CONAN-2 (2-way RTE) and E-CONAN-3 (3-way NLI) Four sentence pair sources: automatically-translated, human-validated machine-translated, hand-crafted from Arabic teaching materials, and news headlines with rumors Nine multilingual pretrained models evaluated on E-CONAN, ArNLI, and XNLI using zero-shot classification
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
The source did not state who this applies to in practice.
A new benchmark dataset paper without peer review, presenting resource development and descriptive evaluation across multiple models without controlled comparisons or definitive performance claims.
As stated by the source record.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
What is missing. This record has no reported figures. That is a gap in the analysis, not a judgement about the study.
Natural Language Inference processes pairs of sentences to extract their semantic relations. NLI has been a hot research topic, integrated as a main component in other NLP applications. Despite significant advancements in textual inference across various languages all around the world, Arabic language still suffers from limited resources in this domain. To address this gap, this paper introduces E-CONAN benchmarks that are composed of sentences pairs from various sources: (1) automatically-translated pairs, (2) human-validated machine-translated pairs, (3) hand-crafted pairs from teaching Arabic as foreign language books, and (4) headlines pairs from different news channels containing rumors. E-CONAN contains two benchmark datasets, E-CONAN-2, a 2-way dataset (RTE) and E-CONAN-3, a 3-way dataset (NLI). Additionally, we have used E-CONAN benchmarks to evaluate 9 state-of-the-art multilingual pretrained models using zero-shot classification. Models were evaluated across the ArNLI, XNLI, and E-CONAN datasets. Results show that E-CONAN is a potentially valuable resource for evaluating model generalization and even for fine-tuning pre-trained models. Its diverse composition, derived from a combination of sources, offers a broader and more robust assessment compared to XNLI and ArNLI. In addition, we have evaluated 5 LLMs on E-CONAN-3 dataset. Moreover, we incorporated MARBERT as a representative Arabic-specific baseline and conducted performance evaluation comparison to demonstrate how Arabic-specific models scale against cross-lingual and LLM-based approaches on the E-CONAN benchmarks. Furthermore, we conducted detailed qualitative and quantitative error analysis to analyze frequent error patterns. E-CONAN benchmarks will be publicly available, we hope that it will enrich research community in Arabic textual entailment and natural language inference.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.