Independent agent loops · training feedback → frozen-policy evaluation
Loops
| Loop | Trace / SLO | Backend · model / effort | Status | Turns | Best train | Selected eval |
|---|
Compare backends on the same episode
All backends on matching training conditions. Each point is an individual turn's reward; dashed lines show best so far. Failed or unscored loops remain listed below. Eval is held out; policy selection follows each loop's training reward.
Training
Evaluation
| Backend | Status | Turns | Best train | Selected eval |
|---|
Reward per turn & best so far
Every scored attempt has a reward point. Dynamics, unchanged files and failed calls have no reward. Evaluation is computed after generation and is never sent to the agent. The selected policy follows training reward; eval's best curve is descriptive.
Training feedback
Evaluation
Default-policy baseline scores
Training baseline · shown before turn 1
Eval baseline · held out from agent
Turn trajectory
Select a turn to inspect it. Page turn numbers start at 1; raw record indices start at 0. Elapsed time includes the initial baseline.
| Turn | Kind | Train | Train SLO | Eval | Eval SLO | Elapsed | Input tokens | Output tokens | Tools |
|---|
Inspect a turn
Records & configuration
Public snapshot: prompts, replies, tool events, policies and score summaries are included. Local paths are anonymized. Full per-request replay logs remain in the original local run.