Independent agent loops · training feedback → frozen-policy evaluation

Loops

LoopTrace / SLOBackend · model / effortStatusTurnsBest trainSelected eval

Compare backends on the same episode

All backends on matching training conditions. Each point is an individual turn's reward; dashed lines show best so far. Failed or unscored loops remain listed below. Eval is held out; policy selection follows each loop's training reward.

Training

Evaluation

BackendStatusTurnsBest trainSelected eval

Reward per turn & best so far

Every scored attempt has a reward point. Dynamics, unchanged files and failed calls have no reward. Evaluation is computed after generation and is never sent to the agent. The selected policy follows training reward; eval's best curve is descriptive.

Training feedback

Evaluation

Default-policy baseline scores

Training baseline · shown before turn 1

Eval baseline · held out from agent

Turn trajectory

Select a turn to inspect it. Page turn numbers start at 1; raw record indices start at 0. Elapsed time includes the initial baseline.

TurnKindTrainTrain SLOEvalEval SLOElapsedInput tokensOutput tokensTools

Inspect a turn

Records & configuration

Public snapshot: prompts, replies, tool events, policies and score summaries are included. Local paths are anonymized. Full per-request replay logs remain in the original local run.

Loop configuration