The first complete replay gave +0.163 under direct scoring and −0.160 with an externally generated view. Both channels saw the same frozen before and after agent states. I had assumed a step score's sign would survive changes to prompts and evaluators, and these results broke that assumption.
The actor had already finished. Its stored state and candidate answers stayed fixed; the measurement setup changed. I find it more useful to describe this as measurement than judgment. The frozen transition is the signal, while the prompt, representation, readout, and hosted model form the channel through which I observe it.
We call the within-channel update Agent Step Value, or ASV. Comparing matched ASVs with Ω on a complete state-by-channel face audits whether that update keeps its direction across channels. The audit leaves universal evaluator correctness and causal credit untouched. I now want transport checked before a step score is pooled, substituted, or reused as supervision.
The same frozen transition can receive a positive ASV in one declared channel and a negative ASV in another. A channel-free step score omits part of the measurement.
transitions complete across all four cyclic candidate layouts, from 1,100 replayed
mean ASV, with a 95% interval from +0.102 to +0.218
mean ASV, with a 95% interval from −0.244 to −0.079
matched channel difference, with a 95% interval from −0.418 to −0.232
The result that made me distrust a naked step score
Step-level scores are attractive because they turn a messy trajectory into a sequence of numbers. A positive number says the agent made progress. A negative one says it moved away from the target. Once that number exists, it is easy to pool it with other scores, compare it across evaluators, or feed it into another system as supervision.
I had been treating the sign as the sturdy part. Different evaluators might disagree about magnitude, but surely a useful retrieval step would still point in the same direction. The replay did not support that intuition. The direct channel produced a positive mean update of +0.163. The generated-view channel produced a negative mean update of −0.160. The matched interaction between them was −0.323, and its 95% interval stayed below zero.
The individual transitions made the result harder to dismiss as one aggregate quirk. Direct and generated-view gains landed on opposite sides of zero for 507 of the 1,004 complete transitions. The flips went both ways: 240 moved from positive under direct scoring to negative under the generated view, while 267 moved from negative to positive. The population means gave a strict reversal, but the paired story was bidirectional.
Negative interactions were not more common. The median Ω was +0.014, and 496 of 1,004 interactions were negative. Their larger magnitudes pulled the mean below zero. The 10% trimmed mean was still −0.033, and every leave-one-trajectory-out mean stayed between −0.335 and −0.306.
The plot supports a narrower claim than "one channel made every step worse." Channel choice often changed the sign assigned to a frozen transition, and the population mean reversed. A score shown without its channel leaves out a condition that made the measurement possible.
A step value belongs to its evaluation channel
By evaluation channel, I mean the full operational scoring configuration. That includes the evaluator prompt, the readout used to turn its response into candidate energies, the hosted model stack, any derived view supplied to it, and the declared presentation settings. Change any of those pieces and the measurement channel has changed, even if the actor trace and candidate set remain fixed.
Within one channel, ASV measures the change in the reviewed target's margin from the state
before a step to the state after it. The target margin is the target candidate's energy minus
the log-sum-exp of the competing candidates. In one channel, ASV is therefore
Gρ = Fρ(x+) − Fρ(x−). In signal-processing terms, a channel-specific additive
level behaves like a DC offset. The before/after difference cancels that offset, but it does
not remove a channel response that depends on the state.
Ω = ΔxΔρ F First compare before with after inside each channel. Then compare those two updates.
A nonzero Ω is what remains after those within-channel differences are compared. It says the channel responded differently to the state change itself. In communications language, the response depends on the signal state. The measurement identifies a state-by-channel interaction. It does not identify a nonlinear mechanism inside the evaluator.
Independent recalibration sets the boundary. A channel-specific constant still cancels, and a positive gain can stretch or shrink ASV without reversing its sign. Ω is not generally invariant to separate gains, but a strict reversal is. One channel stays positive and the other negative under every admissible positive recalibration.
Averaging over prompts does not prove stability. A weighted average of channel gains defines another channel mixture. If the component gains straddle zero, changing the weights can change the pooled sign. The resulting number may be less noisy even though its direction still depends on an undeclared mixture.
Why the four corners matter
Auditing transport requires four scores. We have two frozen agent states, before and after, and two evaluation channels, reference and alternative. Each state must pass through each channel. The resulting two-by-two face gives a before/after gain inside each row, which we can then compare.
Same frozen transition, reference instrument, first endpoint.
Same instrument, second endpoint. This row gives the reference ASV.
Same first endpoint replayed through the alternative instrument.
This row gives the alternative ASV. Comparing rows isolates the interaction.
A common shortcut uses only the diagonal: score the before state in one channel and the after state in another. That comparison looks efficient, but it mixes three things. It contains the reference step gain, the baseline shift between channels, and the state-by-channel interaction. With one or more corners missing, those pieces cannot be separated unless we add assumptions about how the instrument behaves.
Among unrestricted two-by-two corner designs, the complete face is both necessary and sufficient. If any corner is missing, its unknown value can change Ω while every observed score stays fixed.
Repeated evaluator calls do not fill the missing corners either. Repeats tell us whether a declared channel is stable across acquisitions. The complete face tells us whether an update transports across channels. Repeatability and transport are different questions, and I no longer want one used as a substitute for the other.
In practice, I freeze the actor transition and candidate set, then replay all four cells under explicit layout and acquisition rules. Resampling stays at the trajectory level so repeated steps from one task remain together. Compared with a diagonal shortcut, the two extra observations are what separate baseline channel shift from the interaction.
What happened across 1,100 PubMed transitions
The empirical study used 100 open question-answering tasks from a PubMed evidence workflow. Each task had three concrete answer candidates plus a none-of-the-above option. We replayed 1,100 frozen agent transitions through direct and generated-view channels, rotating the four candidates through four cyclic label layouts so every candidate occupied every physical position once.
A transition entered the primary analysis only when both endpoints were complete in both channels across all four layouts. That left 1,004 transitions. On this common cohort, the direct mean was +0.163 with a 95% interval of +0.102 to +0.218. The generated-view mean was −0.160 with an interval of −0.244 to −0.079. Their interaction was −0.323 with an interval of −0.418 to −0.232. All four layout-specific interactions were negative as well.
The main estimate follows the empirical mixture of reviewed targets in the complete cohort. A post hoc sensitivity gave equal weight to all four target types. Under that weighting, the interaction remained negative at −0.213, with an interval from −0.343 to −0.090. The direct mean moved to +0.004, with an interval from −0.044 to +0.057, so the strict mean reversal was no longer supported. The global interaction survives this reweighting, but the claim that both population means sit on opposite sides of zero depends on the empirical target mixture. The estimand, completeness rule, and population weighting therefore need to travel together when the result is reported.
The generated view carried the reversal
After detecting the interaction, we asked where the negative component entered. A symmetric retrieval face crossed the agent state with a frozen view and averaged the two possible paths through that face. It assigned +0.224 to the state coordinate, with an interval that crossed zero, and −0.915 to the generated-view coordinate, with an interval from −1.321 to −0.522. For two factors, this symmetric split is the Shapley allocation. We use it only to describe where the diagonal change enters.
That allocation points to the view coordinate, but it does not yet tell us why. To reduce the prompt-structure alternative, we compared quote and generated views under the same template and neutral authority. The mean generated-minus-quote interaction was −1.306, with an interval from −1.822 to −0.797. Quote minus empty was −0.411, with an interval from −0.632 to −0.184. Both intervals excluded zero. Generated minus quote was the larger negative contrast, and quote also differed from empty.
The matched cube also varied the authority instruction. Empty, quote, and generated gains moved from +1.314, +0.903, and −0.403 under neutral authority to +0.579, +0.565, and −0.646 under evidence-primary authority. The authority change in the generated-minus-quote interaction was +0.095, with a 95% interval from −0.154 to +0.349. Its 90% interval, −0.114 to +0.310, extended beyond the prespecified ±0.25 equivalence band. This run could not establish that authority combined additively with view.
Every scorer already received the same projected state, question, and candidate contract. Quote and generated views were derived from those retained inputs; the empty condition added nothing. With every other part of the channel fixed, the paper's benchmark predicts equal response distributions and stored-success AUCs across views. The observed channels departed from both predictions. That still does not reveal the evaluator's internal mechanism.
We then followed the sign along a three-vertex bridge. Vertex A used a joint four-label readout on the DeepSeek stack. Vertex B changed to a candidate-wise binary readout on the same stack. Vertex C kept that binary contract and moved to GPT-4o. All three vertices reused the same actor states, candidates, quotes, and generated views. The bridge retained all 100 trajectories across four repeated acquisitions and every vertex.
generated-minus-quote interaction; quote-minus-generated AUC +0.169
generated-minus-quote interaction; quote-minus-generated AUC +0.165
generated-minus-quote interaction; quote-minus-generated AUC +0.267
Every interaction interval stayed below zero. Quote gains were positive and generated-view gains were negative at every vertex. Quote also ranked stored task success better at all three vertices. The paired AUC intervals were +0.085 to +0.261 at A, +0.081 to +0.250 at B, and +0.143 to +0.393 at C.
The bridge supports only a directional comparison. Each vertex lives on its own operational scale, so the three magnitudes are not comparable. Their signs remained negative after the readout change and after the hosted evaluator stack changed.
On 96 jointly complete retrieval trajectories, the own-minus-donor gain contrast was −4.080, with an interval from −5.060 to −3.104. Reversing the state/view presentation order produced an imprecise −0.403 shift, while evidence-primary authority shifted the gain by +2.272. A length-only account does not fit these controls, and the evaluator's use of the redundant view remains unresolved.
What ASV can and cannot tell us
ASV gives the before/after margin change inside one declared channel. Matched ASVs on a complete face show whether that change keeps its direction across channels, and Ω separates an additive offset from a state-dependent response without assuming that the channels differ only by a constant. Together, those measurements support a pre-use transport audit for evaluator-derived reward labels, search heuristics, and process measurements.
ASV cannot assign causal credit to the actor. A retrieval step and the state around it are part of an observed trace. The interaction tells us how an instrument measured the frozen change, not what would have happened under a counterfactual action. Calling ASV a causal reward would claim more than the design identifies.
The study also does not establish provider-wide invariance or explain evaluator internals. It covers 100 tasks from one actor workflow, four layouts from one serialization orbit, and one bridge path across two readouts and two hosted stacks. The quote/generated contrast changes several representational details at once. Stored success is a criterion from the same corpus, not an external judgment of step quality. The bridge leaves one joint-readout GPT-4o vertex unobserved, so the compound endpoint cannot be fully decomposed.
The equal-target sensitivity adds another boundary. It preserved the negative global interaction but left the matched-cube localization and the all-vertex strict-reversal conjunction unresolved. It also did not support a strict mean reversal. New tasks, actors, view constructions, candidate contracts, and external criteria still need their own complete faces.
The PubMed questions are research artifacts from public literature. They contain no patient records and support no diagnosis or treatment claim. Consequential use still needs source evidence, evaluator disclosure, channel-sensitivity results, and human review.
I now treat an evaluator-produced step value as a signal read through a declared measurement channel, and I store that channel with the score. Before pooling the value, comparing it with another instrument, or using it as supervision, I want the complete face replayed. This study showed why: moving the same frozen steps between channels reversed the population mean.
Keep the channel with the score
The same actor steps produced +0.163 and −0.160 after only the measurement setup changed. I do not want either number pooled or reused until that transport has been tested.