Life sciences · Preprint
arXiv · August 10, 2026
Posted before peer review. The findings may change or fail to hold.
This is an unpublished technical preprint proposing a zeroth-order optimization framework to allow LLM agents to improve beyond their inherent capability boundaries. The authors report that their method obtains more successful trajectories and outperforms baselines, especially on difficult examples, but the work has not undergone peer review and lacks detailed quantitative reporting of results and comparisons.
Preprint. LLM agents tested on multiple deep research benchmarks. Intervention: Zeroth-order self-evolution framework with LoRA perturbation, gradient estimation, and supervised fine-tuning. Compared with: Strong baselines (names not specified).
Method obtains substantially more successful trajectories compared to strong baselines Consistently outperforms baselines especially on difficult examples
Safety was not reported in the material analysed. Check the source before drawing any conclusion about harm.
The source did not state who this applies to in practice.
This is an unrefereed arXiv preprint describing a novel optimization framework for LLM agents; it reports experimental results but lacks peer review and independent validation.
As stated by the source record.
Graded across the dimensions that decide whether you should act, each from what the source actually supports. There is no single score, and where a dimension was not assessed it says so.
What is missing. This record has no reported figures. That is a gap in the analysis, not a judgement about the study.
Self-evolving methods improve the capabilities of LLM agents by sampling trajectories from the underlying LLMs and learning from these trajectories. However, these methods struggle to learn beyond the inherent capability boundary of the agents, since the agents cannot sample correct trajectories on difficult examples for further improvements. In this paper, we propose a zeroth-order self-evolution framework that enables agents to learn beyond their capability boundary by perturbing LLM parameters to adapt to difficult examples without any trajectory annotations. Specifically, we perturb LoRA parameters of LLMs, run the agent, compute the losses under the perturbed and original parameters, and use the loss difference to estimate gradients and further update the LoRA parameters. We sample trajectories using the updated LLMs for supervised fine-tuning to break through the capability boundary of the agents, forming a closed self-evolution loop. We introduce a parallel perturbation inference mechanism and an adaptive lookup mechanism to reduce time consumption in zeroth-order optimization, with an answer perplexity loss that provides smooth and stable zeroth-order loss values. Experiments on multiple deep research benchmarks show that our method obtains substantially more successful trajectories and consistently outperforms strong baselines, especially on difficult examples. The code and released artifacts are available at https://github.com/hidk1911/ZOForLLMAgents.
Taken from the source record, never inferred. Follow any of these and new work involving them reaches your briefing.