Training a weak SWE agent is a familiar problem: collect correct trajectories, remove bad actions, and run supervised fine-tuning. Training an agent that is already strong is a different problem.
Traditional filtering asks: Is this a good trajectory?
We also need to ask: Does this good trajectory still contain anything new for this student?
The argument of this note is that data selection for heavily post-trained, behaviorally mature SWE agents should move from teacher-centric quality filtering toward student-relative residual learning. Data quality is no longer only an intrinsic property of a trajectory. It is a property of the interaction between the data, the current student, and the states that student will actually visit.
SWE-Lego- and SWE-Master-style filtering remains a necessary first stage. After it, teacher-forced NLL is a useful and inexpensive second-stage signal for estimating how much information a candidate trajectory still contains for the student. The next step is to select supervision from student rollouts, student failure states, and student-relative issue difficulty, bringing the training distribution closer to deployment.
Quality filtering asks whether a trajectory is worth teaching to any model. Residual learning asks which parts are still worth teaching to this student.
#1. Background: why v14.d should have worked
Recent SWE-agent research offers a remarkably consistent lesson: SFT is often the most cost-effective stage of capability growth for a small model. SWE-Lego reaches 42.2% and 52.6% on SWE-bench Verified with Qwen3-8B and Qwen3-32B using high-quality data, erroneous-action masking, and a difficulty curriculum. Orchard's Qwen3-30B-A3B-Thinking reaches 64.3% after SFT and 67.5% after RL. SWE-Master similarly combines teacher rollout, data cleaning, long-horizon SFT, execution-feedback RL, and inference scaling in one pipeline.13
These results established the standard teacher-centric recipe: sample recent issues, run a strong teacher in a real container, retain trajectories that pass tests and protocol checks, mask incorrect or inefficient actions, and organize the remainder by difficulty, length, or reward.
v14.d followed that recipe. It used successful MiniMax-M2.5 trajectories from Orchard-SWE. The final set contains 16,124 trajectories covering 10,551 manifest issues: 72.8% from Scale-SWE and 27.2% from SWE-rebench. Filtering required external resolution, a valid protocol, no Git-history leakage, no test contamination, and no explicit unresolved self-test failure. Training retained full thinking and multi-turn observations.2
By conventional standards, this was ideal data: real issues, real execution, a strong teacher, successful trajectories, and full long-horizon context. We expected it to make Qwen3.5-4B a better SWE agent.
#2. v14.d: healthy training, no better policy
We ran full-parameter SFT on Qwen3.5-4B with 64K context, two epochs, global batch size 32, and a peak learning rate of 1e-5. The complete 1,008-step W&B history looks entirely normal: mean update loss is 0.243, the final 50-step mean is 0.214, held-out loss falls from 0.2646 to 0.2541, and every recorded loss and gradient norm is finite.5
SFT v14.d training log
Independent teacher forcing confirms that the model learned the training distribution. Base-model token NLL on v14.d train-300 is 0.307 and falls to 0.211 after SFT; validation-256 falls from 0.300 to 0.252.15
Closed-loop behavior did not deliver the expected gain. After repairing the old evaluation, the final merged 500-problem result passes 34.2%. Only 53.6% of trajectories actively submit, and 40.6% consume the full 250-turn budget. The model becomes better at predicting MiniMax-M2.5 while continuing to wander and fail to converge.6
Evaluation scope. The 34.2% is the project's final merged result. The first 200 problems include selective reruns, so it is not a single unbiased pass@1 run. The uniformly configured problems 301–500 pass at 30.0%. This limits leaderboard comparability, but it does not remove the behavioral evidence in the low submit rate and high turn-cap rate.6
Our first suspicion was not that the data was too familiar to a strong student. It was more mundane: was there a hidden bug in the data, prompt, chat template, loss mask, or training recipe?
#3. A forensic audit found no bug that explained the result
We rechecked the full chain.26
| Audit layer | What we checked | Result |
|---|---|---|
| Data selection | resolution, leakage, test contamination, split isolation, frozen hashes | passed |
| Action protocol | whether training prompts and the deployed harness both required one action per turn | passed |
| Chat template | full thinking, multi-turn observations, and historical assistant reasoning | passed |
| Loss mask | which tokens were supervised and whether assistant terminators were omitted | passed |
| Optimization | loss, held-out loss, gradient norms, checkpoints, and the actual run config | passed |
Data protocol: v14.c to v14.d changed one contract
v14.c had already completed strict selection and full-corpus structural audits, but the original Orchard first-user prompt still said “at least one command,” while the exported mini-swe harness accepted exactly one action per turn. v14.d did not reselect more favorable data or rewrite teacher behavior. On the frozen v14.c split, it changed each first user message to require exactly one <think>...</think> block and exactly one shell action inside one bash block.
Issue selection, all 745,175 training assistant turns, all 11,624 validation assistant turns, and every assistant/observation body remained unchanged. The builder pinned source hashes, failed closed on unrecognized prompt shapes, rewrote all 16,124/256 rows, and verified one balanced native think block per assistant turn.26
Chat template and loss mask: training retained historical reasoning
Training used the installed MS-SWIFT template with preserve_thinking=true and loss_scale=default. We re-encoded early examples and the largest serialized examples, comparing frozen ChatML against trainer output token by token. Across 32 audited rows, input IDs and label masks matched exactly. System, user, and tool observations were context only; each assistant turn's reasoning, action body, and <|im_end|>\n were supervised.26
There was a real chat-template bug, but it belonged to the old evaluation path: later turns omitted historical assistant reasoning from context. The repaired evaluation preserves that reasoning. Training data and the training loss mask never used the faulty path, and low submission plus high turn-cap behavior persisted after the repair.6
Training recipe: the likelihood objective was optimized stably
The completed run used 8×B300, BF16, Flash Attention 2, ZeRO-3, gradient checkpointing, 64K without packing, global batch 32, cosine decay, and peak LR 1e-5. These facts come from the actual W&B run config, not from later launcher defaults.5 Training loss, held-out loss, and independent NLL all point in the same direction: optimization was stable and the student moved closer to the teacher distribution.
This does not prove that every design choice in v14.d was optimal. It rules out the simpler story: the result was not caused by silently changing the dataset, leaving train and deployment protocols misaligned, dropping historical reasoning during training, supervising the wrong token classes, or suffering numerical optimizer failure. This was an implementation-correct SFT that succeeded at its objective and failed to improve closed-loop behavior.
#4. TMax turns one failure into a repeatable pattern
At this point, v14.d could still have been an isolated case. TMax provides a cleaner independent contrast: with the same TMax-SFT data and a similar recipe, an older Qwen3-8B improves while the heavily post-trained Qwen3.5-9B degrades. On Terminal Bench Lite, Qwen3-8B moves from 7.3% to 11.5%; Qwen3.5-9B moves from 41.9% down to 35.5%.4
We computed assistant-token teacher-forced NLL on the same 1,000 TMax-SFT trajectories using each model's official chat template:
| Student | Token NLL | Bits per byte | SFT result on the same data |
|---|---|---|---|
| Qwen3.5-9B | 0.274 | 0.114 | 41.9 → 35.5 |
| Qwen3-8B | 0.792 | 0.323 | 7.3 → 11.5 |
Token NLL differs by 2.89×; bits per byte, which controls for tokenizer granularity, differs by 2.84×. The data contains clearly novel behavior for Qwen3-8B and is already highly familiar to Qwen3.5-9B. The direction of the benefit changes with the student, so “high-quality teacher trajectories” cannot by itself determine whether SFT will help.4
#5. NLL reveals a candidate golden region
NLL does not ask whether a trajectory is correct. It asks how difficult it is for the student to predict the next reasoning segment, action, and stop signal given the teacher-produced history.
NLL asks a student-relative question
Putting v14.d, TMax, and the later SWE trace experiments side by side reveals a candidate separation:
| Student × candidate traces | Teacher-forced NLL / approximate training signal | Known training result |
|---|---|---|
| Qwen3.5-9B × TMax-SFT 1K | 0.274 | 41.9 → 35.5 (−6.4pp) |
| Qwen3.5-4B × v14.d train-300 | 0.307 | no expected closed-loop gain |
| AxiaoDBL Qwen3.5-4B × DeepSeek-V4-Flash GA | early / epoch-one ≈ 0.55 | 44.8 → 55.0 (+10.2pp) |
| Qwen3-8B × SWE-Lego resolved-500 | 0.738 | 7.6 → 42.2 (+34.6pp) |
| Qwen3-8B × TMax-SFT 1K | 0.792 | 7.3 → 11.5 (+4.2pp) |
| Qwen3-8B × AgentForge random-500 | 0.812 | 8.0 → 38.2 (+30.2pp) |
The three Qwen3-8B results are particularly useful. With the backbone fixed, the data associated with successful training clusters at 0.738–0.812. Restricting SWE-Lego and AgentForge to wholly untruncated trajectories gives 0.758 and 0.824, so truncation does not create the ordering.124
This is not a causal curve. The 500 SWE-Lego and AgentForge examples are samples from public trajectory families, not either complete training mixture; the recipes also differ in masking, curricula, and harnesses. AxiaoDBL's 0.55 is trainer update loss rather than a strictly matched offline token NLL. Most importantly, AgentForge at 0.812 sits just above 0.8. The defensible claim is therefore a candidate operating region with a 0.5–0.8 core and a soft upper edge near 0.85. The “golden region” is a hypothesis for the next experiment, not an established law.
The same MiniMax v14.d data has NLL 0.668 under Qwen3-8B, 2.17× the Qwen3.5-4B value. Qwen3.5-4B has similarly low NLLs of 0.319 and 0.355 on Kimi-K2.6 and DeepSeek-V4-Flash Preview SBV traces, but we do not have matched causal SFT outcomes.1215 These cross-scores reinforce that NLL measures student × exact traces, not absolute teacher or issue-pool quality.
#6. Why NLL is cheap, useful, and insufficient
Agentic SFT does not directly optimize issue resolution. It maximizes the likelihood of the teacher's next reasoning segment, tool call, and termination action on histories visited by the teacher:
On a fixed dataset, cross-entropy approximates minimizing , but only at states . At deployment, the student sees histories induced by its own actions. One bad search or edit can move it into states absent from the teacher data.19
SFT gains therefore come from interface alignment, action-prior reshaping, and reusable procedure transfer. A strong student may already possess the first two; a teacher's full long-horizon procedure may also exceed what it can express reliably. Low training loss then proves “more like the teacher,” not “better under the student's own rollouts.”
Existing selection methods approximate training value from different directions. SWE-Lego, SWE-Master, and DEITA first construct correct, clean teacher pools; IFD uses student likelihood to estimate sample difficulty; LESS-style methods use warm-up training and per-example gradient influence to estimate relevance to a target task; GRAPE, token rank, and RSR add distributional fit and informative alignment.1172123
NLL's advantage is not completeness. It is low cost, protocol fidelity, and suitability for large-scale preflight. Once the student is frozen, candidate trajectories need only one gradient-free teacher-forcing forward pass. There is no warm-up checkpoint to train, no per-example gradients to store, and no separate target validation set to define. Before expensive rollout expansion or full-parameter training, NLL can quickly remove data that is absolutely correct but relatively redundant for this student.
The limitation is that NLL measures residual information only. High NLL can arise from noise, incompatible formatting, or a teacher policy the student cannot express; the lowest NLL can simply identify behavior the student already knows. A stronger target is therefore: after quality gating, find trajectories with moderately elevated NLL, reasonable rank on important tokens, and a direct relationship to actual student failure states. The operating region moves with model, tokenizer, template, mask, and domain; cross-tokenizer comparisons should also report BPB or nats per character.
#7. The rollout teacher can matter more than the issue source
The AxiaoDBL author states that its issue pool is a subset of PrimeIntellect/SWE-Lego-Real-Data-Verified and that the rollout teacher is DeepSeek-V4-Flash.79 Prime derives from SWE-rebench and repairs test IDs before validating F2P/P2P with gold patches: a strong traditional quality filter.
Issue source cannot explain the entire difference. Exact instance-ID overlap between the v14.d manifest and the 4,323 resolved Prime Verified issues gives:
| Overlap measure | Count / share |
|---|---|
| Exact overlapping issues | 1,736 |
| Share of Prime Verified issues | 40.2% |
| Share of v14.d Rebench issues | 61.8% |
| Share of all v14.d manifest issues | 16.5% |
| Matching v14.d trajectories | 2,635 / 16,124 (16.3%) |
The two datasets do not inhabit entirely different issue worlds. A larger difference is the behavior generated on similar issues: MiniMax-M2.5 versus DeepSeek-V4-Flash GA/0731, followed by different action masking and curriculum choices.16
The exact checkpoint matters. DeepSeek reports that V4-Flash-0731 retains the architecture and scale of Preview but receives new post-training. Terminal Bench rises from 61.8 to 82.7, NL2Repo from 39.4 to 54.2, DeepSWE from 7.3 to 54.4, and Toolathlon from 49.7 to 70.3.10
The SWE-Router DeepSeek-V4-Flash SBV traces were generated before the 0731 release and belong to the Preview period.12 Our Qwen3.5-4B NLL of 0.355 on those traces cannot substitute for AxiaoDBL's GA-teacher data. Community information identifies the AxiaoDBL teacher as the GA/0731-era model, although the public discussion retains only the mutable DeepSeek-V4-Flash alias and not an immutable revision hash.8
The unit of an agentic dataset is not an issue or even a resolved trajectory. It is issue × exact teacher checkpoint × harness × reasoning effort × sampling × loss mask.
The cheapest test is not the teacher name or its benchmark score. It is the student's likelihood on the exact candidate traces.
#8. From teacher quality to a student-relative residual
SWE-Lego improves trajectory quality through execution verification, erroneous-action masking, and difficulty curricula. SWE-Master combines reward, format, length, BoN difficulty, and LLM-judge filters. Orchard adds credit assignment that recovers valuable segments from failed trajectories.13
A strong student needs five dimensions of selection:
| Dimension | Question | Candidate signals |
|---|---|---|
| Correctness | Did the patch actually resolve the issue? | F2P/P2P, gold-patch validation |
| Process quality | Are actions wrong, inefficient, or contaminated? | action mask, format/length filters, LLM judge |
| Issue difficulty | Is the issue too easy, intermediate, or impossible for this student? | student pass@k, IFD |
| Student-relative information | Is this behavior already familiar? | NLL, BPB, token rank, RSR, self–teacher gap |
| Deployment-state relevance | Does supervision cover states the student reaches? | student rollout prefixes, failure states, behavior classes |
Correctness and process quality decide whether the data can be learned safely. The remaining dimensions decide whether this student should spend updates on it. The NLL preflight in the previous section only locates a potentially useful residual cheaply; IFD, LESS, GRAPE, RSR, token rank, and student pass@k continue the work by estimating difficulty, absorbability, target relevance, and real state coverage.172123
NLL can remove data that is absolutely correct but relatively redundant. Token rank and execution checks help reject data that is surprising because it is wrong or unreachable. Student rollout states determine whether supervision corresponds to mistakes the deployed model will actually make.
This is student-relative residual learning:
Do not compress the strongest teacher's entire behavior into the student. Within verified behavior, teach the learnable residual between the student's current policy and the target policy.
#9. Why the data should become approximately on-policy
Pure teacher rollout is off-policy. The teacher chooses which files to inspect, which commands to run, and which intermediate states to enter. The student imitates only on those states. A resolved patch does not guarantee that the teacher states are where the student most needs help. With a long SWE-agent horizon, one early divergence can make the second half of a teacher trajectory irrelevant to the student's actual history.
DAgger queries an expert on states induced by the learner. GKD similarly lets a language-model student generate sequences first, then obtains teacher supervision on those student-generated sequences; it can also replace fixed forward-KL distillation with other objectives.19 Fully on-policy expert annotation is expensive in a real SWE environment, but several approximations are practical:
- rollout the student first and select issues that are intermittently solvable or have recoverable failures;
- start teacher correction from a student failure prefix, bad edit, or validation loop instead of restarting from a clean state;
- retain correct student actions and supervise only the teacher's key residual actions;
- recompute NLL and rank after intermediate checkpoints so that the selected data moves with the student policy.
These methods do not eliminate every limitation of forward KL. They shift the state distribution in the training expectation from toward : the model learns at places it will actually reach, and the supervision concentrates on gaps in its current policy. For an already strong agent, that is closer to capability extension than simply enlarging a static dataset of teacher successes.
#10. A cheap preflight before training
Before paying for thousands of rollouts or a full SFT run:
- Apply traditional quality filters first. Require an executable environment, resolved patch, valid F2P/P2P, no leakage, and valid actions.
- Freeze the scoring protocol. Use the exact student checkpoint, training tokenizer/chat template, complete historical reasoning, and the same loss mask.
- Score 100–1,000 candidate trajectories. Report token-weighted mean, per-trajectory quantiles, target-token count, and BPB across tokenizers.
- Anchor against student rollouts. A teacher NLL close to the student's self-rollout NLL may contain little new information; a moderately higher value with reasonable rank and verified behavior is more promising.
- Estimate student-relative difficulty. Divide issues into always-pass, mixed, and always-fail using pass@k; inspect mixed and recoverable failures first.
- Decompose by behavior. Score reasoning, tool actions, first effective edit, verification, and submission separately.
- Run a small closed-loop pilot. Track pass@1, submit rate, first-edit turn, empty patch, no-progress turns, and turn/context caps before scaling.
The inexpensive signal set is therefore: pass@k for prior student capability; NLL and token rank for novelty and absorbability; student rollouts for the deployment state distribution; and a small closed-loop pilot for whether teacher-forcing gains transfer into real behavior.
#11. Next step — Coming soon
The next experiments will test the residual-learning hypothesis directly rather than comparing more teacher names.
Experiment 1: Is NLL sufficient, or is informative alignment required? With one Qwen chat template, full historical reasoning, and one loss mask, decompose assistant targets into reasoning, tool name/arguments, edits, tests/verification, and submission/termination. Compute NLL, token rank, and RSR for each category.
Experiment 2: Bucket quality-controlled data. On execution-verified trajectories from one issue pool, use <0.5, 0.5–0.8, 0.8–1.1, and >1.1 NLL bins, then match assistant-token budget, repository, length, and teacher across bins. Split each bin again by low versus high token rank. This directly tests whether 0.5–0.8 is a stable operating region, a soft boundary, or an accident of the current small evidence set.
Experiment 3: Add approximately on-policy issue and state selection. Use student pass@k to split always-pass, mixed, and always-fail issues; add teacher corrections that begin from student failure prefixes. Compare static teacher-success SFT, student-difficulty filtering, and failure-state residual SFT.
Experiment 4: Measure capability, not only loss. Hold issue count, assistant target tokens, repository distribution, steps, and teacher constant. Measure invalid tool calls, localization and first effective edit, reproduction/regression tests, premature submission, turn/context caps, checkpoint-level ID/OOD results, and cross-harness behavior.
A falsifiable prediction follows: after quality gating, medium-high-NLL, low-rank trajectories originating from mixed or failed student states will improve closed-loop performance more than lowest-NLL, highest-NLL, or random data. The gain should first appear in the student's existing failure modes—not merely in formatting or response length.
#Conclusion
Traditional SWE data filtering remains essential, but it only ensures that a trajectory is good. Extending an already strong SWE agent also requires asking whether the trajectory is informative for this student, learnable by it, and relevant to states it will actually visit.
v14.d shows the more unsettling failure: 16K rigorously filtered successful trajectories, aligned protocols and masks, and steadily improving train and validation losses still did not produce a better closed-loop policy. TMax then turns the isolated result into a pattern: the same SFT data strengthens Qwen3-8B and weakens Qwen3.5-9B. NLL and BPB expose that difference in familiarity before training; the positive SWE-Lego, AgentForge, and AxiaoDBL results turn a 0.5–0.8 core with a soft edge near 0.85 into a concrete region to test.
The objective should not be merely:
Find more correct trajectories.
It should become:
Within correct trajectories, find the part this student has not learned, can learn, and will need during its own rollout.
Data engineering for strong agents should continue moving from teacher-centric quality filtering toward student-relative residual learning. NLL is not the final answer, but it is cheap and direct. Combined with token rank, student difficulty, and student-generated states, it is closer to the quantity we actually care about than teacher reputation, issue source, or training loss alone.
#Method notes
- The repaired v14.d evaluation preserves all historical assistant reasoning. Assistant turns include malformed action responses. Teacher-forced NLL targets assistant reasoning, tool-call markup, and assistant end tokens; system, user, and tool responses are context only.
- The public audit supplement records the v14.c → v14.d data protocol, MS-SWIFT template, and loss-mask checks.26
- Claude Opus 4.7 SBV traces have anomalously low assistant-turn and token counts and are excluded from the argument because the cause may be the harness or capture process.
- Kimi-K2.6 and DeepSeek Preview NLLs measure Qwen3.5-4B familiarity with candidate trace distributions; they are not causal SFT results.12
- All three Qwen3-8B random/resolved-500 measurements use the official Qwen3 chat template and target only complete assistant bodies, tool calls, and end tokens. SWE-Lego and Open-SWE have high truncation rates at 40,959 tokens, so the article also reports a non-truncated-trajectory sensitivity result.25
- GRAPE and RSR are validated on general instruction tuning and reasoning-trajectory selection. Here they motivate a hypothesis for SWE-agent data selection rather than serve as existing SWE causal evidence.17
- Figure design uses direct labels, low-saturation color, light grids, and whitespace inspired by Transformer Circuits research articles.14
#References
- Tao et al. SWE-Lego: Pushing the Limits of Supervised Fine-tuning for Software Issue Resolving. arXiv:2601.01426, 2026.
- Peng et al. Orchard: An Open-Source Agentic Modeling Framework. arXiv:2605.15040, 2026.
- Song et al. SWE-Master: Unleashing the Potential of Software Engineering Agents via Post-Training. arXiv:2602.03411, 2026.
- Ivison et al. TMax: A Simple Recipe for Terminal Agents. arXiv:2606.23321, 2026.
- This project: SFT v14.d W&B run and API verification record.
- This project: v14.d SWE-bench Verified 1–500 final distribution.
- AxiaoDBL. qwen3.5-4b-swebench-sft model card.
- AxiaoDBL. Teacher model and issue pool discussion.
- PrimeIntellect. SWE-Lego-Real-Data-Verified dataset card.
- DeepSeek. DeepSeek-V4-Flash-0731 model card.
- DeepSeek API. Model updates and release history.
- SWE-Router. swebench-verified-deepseek-v4-flash dataset.
- MemoryAsModality. swebench-verified-kimi-k2p6-traces dataset.
- Anthropic Transformer Circuits. Characterizing interference weights in a tiny language model.
- This project: v14.d, TMax, Kimi, and DeepSeek teacher-forced NLL records.
- This project: v14.d × Prime SWE-Lego Verified exact issue overlap.
- Zhang et al. The Best Instruction-Tuning Data are Those That Fit. arXiv:2502.04194, 2025/2026.
- Yang et al. Which Reasoning Trajectories Teach Students to Reason Better? A Simple Metric of Informative Alignment. arXiv:2601.14249, 2026.
- Agarwal et al. On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes. ICLR, 2024.
- Ross et al. A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning. AISTATS, 2011.
- Li et al. From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning. arXiv:2308.12032, 2023.
- Xia et al. LESS: Selecting Influential Data for Targeted Instruction Tuning. arXiv:2402.04333, 2024.
- Liu et al. What Makes Good Data for Alignment? A Comprehensive Study of Automatic Data Selection in Instruction Tuning. ICLR, 2024.
- Klear Team. Klear-AgentForge: Forging Agentic Intelligence through Posttraining Scaling. arXiv:2511.05951, 2025.
- This project: Qwen3-8B × SWE-Lego / AgentForge / Open-SWE random-500 teacher-forced NLL record.
- This project: v14.d data-protocol and loss-mask audit.