Life sciences · Preprint
arXiv · September 10, 2026
Early or partial results. Treat as a signal, not a conclusion.
This preprint introduces DeFiFlowBench, a 207-prompt benchmark for evaluating natural-language DeFi workflow synthesis, and proposes Koan-Safe, a hybrid method combining intent parsing, generation, and structural repair. On 75 held-out prompts, Koan-Safe achieved a 0.67 safety proxy score compared to 0.33 for baselines, with no unsafe executions in saved outputs; however, ablation and diagnostic testing showed that disabled enforcement or permissive thresholds can still authorize unsafe trades.
Benchmark development with controlled held-out evaluation and ablation testing. Natural-language prompts for DeFi workflow synthesis, authored by a team; configurations tested on simulated EVM. Intervention: Koan-Safe: a hybrid method combining intent parser, replaceable generator, and structural repair with default safety parameters. Compared with: Direct prompting, constrained prompting, and few-shot prompting baselines. n = 207.
Direct, constrained, and few-shot prompting produced 14–19 unsafe held-out executions per configuration under fixed 5% price-impact cap Koan-Safe hybrid variant scored 0.67 on static safety proxy versus 0.33 for best baseline on 75 held-out prompts Koan-Safe recorded no unsafe executions on saved benchmark outputs
Safety evaluation relies on a static proxy metric and local EVM execution; real-world DeFi market conditions not modelled Ablation shows that default safety parameters and thresholds can be circumvented; generalizability of proposed fixes uncertain
The source did not state who this applies to in practice.
This is a preprint describing a new benchmark and mitigation method for DeFi workflow safety, with controlled testing on held-out prompts but no peer review, external validation, or comparison against published baselines in a clinical or regulatory context.
As stated by the source record.
Quoted from the source exactly as published.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
A structurally valid DeFi workflow can still authorize a costly trade. We introduce DeFiFlowBench, a benchmark of 207 team-authored prompts for natural-language DeFi workflow synthesis. It measures graph coverage, configuration completeness, and declared safety predicates, then tests supported trade configurations on a local EVM. Direct, constrained, and few-shot prompting produce 14-19 unsafe held-out executions per configuration under a fixed 5% price-impact cap. A slippage bound derived from a quote does not prevent the price impact of the order itself. We propose Koan-Safe, which combines a prompt-only intent parser, a replaceable generator, and structural repair with default safety parameters. On 75 held-out workflow prompts, its hybrid variant scores 0.67 on the static safety proxy, compared with 0.33 for the best baseline. Koan-Safe records no unsafe executions on the saved benchmark outputs. A matched-candidate ablation produces 14-17 unsafe executions when enforcement is disabled. Additional tests expose the limits of default injection: permissive existing thresholds can still authorize unsafe trades. A separately evaluated policy cap addresses this failure on a 36-case diagnostic grid. These results support explicit trade protections and execution-based evaluation, while distinguishing declared safety from a general guarantee.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.