Boxz Studio

Study Models and evaluation

Small Factual Updates in LLMs

LoRA-SFT / LoRA-DPO Evaluation Study

A completed experimental study of small factual updates in LLMs, comparing LoRA-SFT and LoRA-DPO across update insertion, hallucination-like side effects, Mean BS scores, and general capability retention.

Status
Completed study
Model
Qwen2.5-3B-Instruct
Methods
LoRA-SFT, LoRA-DPO
Retention checks
MMLU, HellaSwag, GSM8K
Trajectory chart comparing SFT and DPO update success against hallucination-like rate across checkpoints.
On this page

research study completed post-training

A small fact is small only at the surface. In a model, the update becomes a question of where the edit ends.

00 Research problem

Many model-update discussions treat factual correction as a simple insertion problem: an old value is wrong, a new value is available, and the model should return the new value. This experiment starts from a narrower but harder premise. The edited fact is deliberately small: same subject, same relation, one changed object value, usually a number or date.

That narrowness is the point. It removes many excuses. If the task is only to replace one value, then the real research question becomes sharper:

Can lightweight post-training absorb a corrected value without scattering uncertainty into adjacent behavior?

The study compares two LoRA-based post-training routes on that question. LoRA-SFT receives supervised examples that point directly to the updated value. LoRA-DPO receives preference pairs where the updated value is preferred over the stale value. Both methods share the same base model family, adapter mechanism, and checkpoint schedule. The difference under inspection is the training objective.

Evaluation stack for small factual updates: targeted success, side-effect rate, and general capability retention.
Figure 01Evaluation is treated as a stack. A factual correction has to pass targeted insertion, side-effect, and general-retention checks before it can be read as a stable update.

01 Experimental design

The update cases are derived from WikiFactDiff-style ReplaceObject entries. Each case keeps the subject and relation fixed while replacing an old object value with a new one. The public version focuses on numeric and date-like replacements with bounded magnitude, because those cases make the output easy to falsify.

The model under study is Qwen/Qwen2.5-3B-Instruct. Both training routes use the same LoRA configuration: rank 16, alpha 32, dropout 0.05, and target modules covering attention and MLP projection layers. Checkpoints are evaluated at steps 200 / 400 / 600 / 800 / 834.

The training surface contains 13,326 examples per format. SFT uses chat-style instruction-answer examples. DPO uses chosen/rejected preference pairs, with the updated value as the preferred answer and the stale value as the rejected answer. Targeted evaluation uses a held-out, relation-stratified set of 75 checks per checkpoint.

This setup makes the comparison narrow on purpose. The adapter mechanism is held constant; the update objective changes.

01 / Insert

Does the model move toward the new literal value?

02 / Disturb

Does the same move increase hallucination-like behavior?

03 / Retain

Does broader benchmark behavior stay near the base model?

02 Evaluation logic

The experiment has one central rule: update success alone is insufficient. A method can move toward the updated value while also making the model more willing to produce unsupported or unstable answers. A method can also look weaker on targeted insertion while preserving broader behavior.

The targeted layer reports three values:

  • Update success: the adapted output is judged closer to the updated value than the base model output.
  • Hallucination-like rate: the adapted output is judged to increase hallucination-like behavior relative to the base output.
  • Mean BS: the average side-effect severity score from the judge output.

The general layer then asks whether the update procedure perturbs broader model ability. It uses MMLU aggregate accuracy, HellaSwag normalized accuracy, and GSM8K flexible exact match, reported as percentage-point deltas against the base checkpoint.

The logical order matters. If the first metric improves while the second metric worsens, the update has become a risk frontier rather than a clean edit.

Training examples 13,326 per SFT / DPO format
Targeted checks 75 per checkpoint / relation-stratified
SFT peak update success 37.33% step 400
DPO peak update success 22.67% step 400
SFT final side-effect rate 46.67% step 834
DPO best average retention delta +1.39pp step 800 / selected benchmarks

03 Targeted result: the update-risk trajectory

The targeted evaluation shows a clear split between the two objectives. SFT reaches higher update success, but it also moves into a higher side-effect region as training continues. DPO reaches lower update success, but its hallucination-like rate stays lower across all measured checkpoints.

Trajectory chart comparing SFT and DPO update success against hallucination-like rate across checkpoints.
Figure 02SFT moves farther toward insertion while accumulating side-effect pressure. DPO forms a smaller, cleaner edit trajectory.
StepNSFT successDPO successSFT hallucination-likeDPO hallucination-likeSFT Mean BSDPO Mean BS
2007528.00%14.67%29.33%17.33%0.5870.373
4007537.33%22.67%32.00%21.33%0.6270.453
6007526.67%17.33%44.00%18.67%0.7330.413
8007528.00%12.00%44.00%18.67%0.7330.453
8347533.33%12.00%46.67%22.67%0.8000.533

The most important checkpoint is step 400. SFT reaches its peak targeted update success at 37.33%; DPO also reaches its peak there, but at 22.67%. The same checkpoint shows the cost of that difference: SFT’s hallucination-like rate is 32.00%, while DPO’s is 21.33%.

Later checkpoints sharpen the pattern. SFT’s targeted success fluctuates while its hallucination-like rate rises to 46.67% by step 834. DPO remains more conservative: lower insertion, lower side-effect pressure, lower Mean BS at every measured checkpoint.

04 General result: retention is a second test

The targeted result would be easy to over-read without the general benchmark layer. A factual edit can appear successful inside the narrow update set while still deforming broader behavior. The selected general benchmarks therefore act as a retention check: MMLU for broad knowledge, HellaSwag for commonsense continuation, and GSM8K for arithmetic reasoning under flexible extraction.

General capability retention chart showing average benchmark delta for DPO and SFT across checkpoints.
Figure 03General benchmark movement is limited in magnitude, but it has method-specific shape. DPO peaks later on selected average delta, while SFT weakens at the final checkpoint.
MethodStepMMLUHellaSwagGSM8KdMMLUdHellaSwagdGSM8KAvg delta
Base066.6 +/- 0.668.0 +/- 4.767.0+0.0+0.0+0.0+0.00
DPO20066.4 +/- 0.665.0 +/- 4.870.0-0.2-3.0+3.0-0.06
DPO40066.5 +/- 0.667.0 +/- 4.769.0-0.1-1.0+2.0+0.29
DPO60066.7 +/- 0.668.0 +/- 4.769.0+0.1+0.0+2.0+0.71
DPO80066.8 +/- 0.668.0 +/- 4.771.0+0.2+0.0+4.0+1.39
DPO83466.7 +/- 0.668.0 +/- 4.770.0+0.1+0.0+3.0+1.04
SFT20066.9 +/- 0.667.0 +/- 4.770.0+0.3-1.0+3.0+0.77
SFT40066.5 +/- 0.667.0 +/- 4.769.0-0.1-1.0+2.0+0.31
SFT60066.6 +/- 0.667.0 +/- 4.768.0-0.0-1.0+1.0-0.01
SFT80066.6 +/- 0.665.0 +/- 4.870.0-0.0-3.0+3.0-0.01
SFT83466.7 +/- 0.665.0 +/- 4.867.0+0.1-3.0+0.0-0.95

DPO reaches the best selected average benchmark delta at step 800: +1.39 pp, driven mainly by GSM8K while MMLU and HellaSwag remain close to base. SFT reaches its best selected average earlier at step 200: +0.77 pp. By step 834, SFT falls to -0.95 pp, mainly because HellaSwag remains down and GSM8K returns to base.

The general benchmark layer leaves the targeted result intact and changes what it means. SFT remains the stronger insertion operator. DPO remains the more conservative operator. The broader evaluation shows why those labels are incomplete without retention and side-effect measurement.

05 What the experiment says

The cleanest reading is this:

SFT behaves like a stronger edit instrument. It pushes harder toward the updated literal value. That makes it attractive if the only target is insertion. The cost appears in side-effect pressure: later SFT checkpoints show rising hallucination-like rate and rising Mean BS.

DPO behaves like a more conservative preference instrument. It does less targeted insertion. It also stays lower on hallucination-like rate and Mean BS across the checkpoint series, and it reaches the strongest selected average retention score in the general benchmark layer.

The mechanism-level reading is that the two objectives apply different pressure to the same adapter surface. SFT trains direct completions and can make the corrected literal value easier to retrieve. DPO trains a contrast between preferred and rejected completions, which appears to regularize against aggressive overwrite but also undershoots exact insertion. The experiment therefore makes objective choice visible as a behavioral shape rather than a training detail.

The result is therefore a map of the update-risk frontier. The same intervention can be judged differently depending on whether the evaluator asks for literal correction, behavioral stability, or retention under broader tasks.

That is the governance relevance of the experiment. Factual updating is often described as maintenance. In practice, it is also an audit problem. A responsible update pipeline needs to say what changed, what remained stable, and what new failure surface appeared after the correction.

06 Reproducibility boundary

The current study is strongest as a controlled comparison of objectives under a shared base model and adapter mechanism. Its targeted held-out set contains 75 relation-stratified checks per checkpoint. The judge layer uses strict JSON fields and deterministic identifiers, but judge dependence remains a methodological constraint.

The public artifacts preserve the code, dataset construction surface, training scripts, evaluation scripts, benchmark outputs, generated tables, and result summaries. The page above presents the research reading; the external repositories provide the inspectable material layer.

07 Materials

  • GitHub repo for code, data-processing scripts, training/evaluation scripts, generated result tables, and summaries.
  • Hugging Face artifacts for the published artifact surface of the small-change SFT/DPO experiment.