Science Explained/Brief
Study probes when LLM draft-verify-revise pipelines disagree on 'previous'
A preprint abstract reports that in draft-verify-revise LLM pipelines, a context-dependent expression such as 'previous' can be resolved differently by different stages, producing a deictic shift. The study used a synthetic dataset of 10 base examples in three conditions and tested six models across 21 reasoning effort configurations.
BriefPublished 14 September 20261 min read1 linked source · 5 checked facts
Balanced accuracy ranged from 0.156, below chance, to near-perfect. GPT-5.2 rose from 0.156 without reasoning to 0.942 at its highest reasoning effort level, while Gemini 3 Pro stayed above 0.94 at every level. Gemini 3 Pro at low reasoning effort outscored GPT-5.2 at xhigh reasoning effort for roughly 5% of the cost per trial. When the meta-evaluator erred, it tended to rely on surface cues rather than operational reasoning. The abstract advises context engineers to make the intended referent explicit at each stage.
Our view
This is a controlled synthetic demonstration, so it shows a possible failure mode rather than a settled rate for real-world pipelines.
What the reporting says: Balanced accuracy ranged from 0.156, below chance, to near-perfect, and that GPT-5.2 rose from 0.156 without reasoning to 0.942 at its highest reasoning effort level.