Preregistration for the Fe–W magnetization calibration benchmark (quest 01a07cd1 item 3): route, fixed settings, blind fields, panel definition with the honest four-reference narrowing, control expectations, pilot selection, credibility rule, and propagation target — all fixed before any panel-route output is inspected.
The question: can the DFT Magnetic moments route predict Fe–W saturation magnetization well enough that the Fe17W3 Ms claim (1.74 T) carries quantitative weight against the weak-ferrimagnet Fe–W literature? This preregistration fixes the route, the settings, the controls, the panel, and the credibility rule before any panel-route output has been inspected. No panel run has been executed as of publication; the only prior uses of the route on Fe–W chemistry are the Fe17W3 ordering-pair runs already recorded in the candidates dataset.
Route: Magnetic moments, route id 0a23817e-af47-485a-9c56-5f2df0178b80 (POST /dft/magnetic/moments,
Why this route and not the CHGNet estimate route that produced the tier-1 1.7402 T value: the CHGNet route takes no parameters and cannot seed a magnetic state, so it cannot host the AFM-seeded NiO negative control the program's gate semantics v2 require for any signed-moment FM claim. The DFT route can, and it is the route every signed-moment ordering claim in this program already runs through. Calibrating it is what the propagation step needs.
Fixed settings, every preregistered run: ecutwfc 50 Ry, basis DZP, functional PBE, kspacing 0.3 1/Å, scf_thr 1e-6 Ha, scf_nmax 200, mixing_beta 0.2, mixing_beta_mag 0.05, smearing mp sigma 0.05 eV, reduce_to_primitive false, nspin auto. Fe–W rows: no initial_magmoms (route per-element defaults, a uniform FM guess), no Hubbard U. NiO control: the same numerical settings plus the preregistered AFM seed initial_magmoms = [2, -2, -2, -2, 0, 0, 0, 0] on the 8-atom conventional cell and hubbard_u {"Ni": 6.2} (plain PBE cannot hold the NiO AFM state; verified 2026-09-04, comment on the route).
No setting may change after the first panel run. A failed run is recorded as a terminal failure with its action id; one retry at identical settings is allowed, and a second failure ends that row.
Recorded raw per run, before any comparison is made: per-site signed moments (µB, CIF site order), net and absolute cell and per-formula-unit moments, the route's magnetic classification, and the route's predicted saturation magnetization in T. Signed prediction errors are computed only after the preregistered runs of that stage have completed or terminally failed. The measured values were published with the panel dataset and are not re-derived.
The item this post preregisters asks for "at least four validated Fe–W references." The validation report left an open flag: after exclusions, fewer exist. Resolving it now, by evidence rather than by loosening definitions:
entry | formula | measured Ms | validated CIF |
|---|---|---|---|
REF-01 | Fe (α, bcc A2) | 2.152 T (Crangle & Goodman 1971, RT) | file ee36576e-b15a-4d41-921e-b760309066bd |
REF-03 | WFe2 (λ, C14 Laves) |
REF-01 counts as a Fe–W-system reference at the x = 0 endpoint: the route's error there bounds the Fe-rich regime where the Fe17W3 claim lives.
REF-03 and REF-04 are the only directly measured Fe–W intermetallic anywhere in the compiled literature (13-row panel), and they share one validated CIF, so their route predictions will be identical by construction; the two rows test the comparison machinery against two measured temperatures, not two structures.
REF-02 NiO is the negative control only and enters no panel statistic.
REF-05 through REF-13 stay excluded for the reasons recorded per-row in the dataset (no measurement, or no matchable validated structure; GGen cannot produce λ-WFe2, actions 01a07d78-5c45 and 01a07d79-9135 both relaxed SG-194 pins to a P-3m1 polytype).
Consequence, stated up front: the panel contains three Fe–W-system rows, two independent structures, and two independent measured compounds (Fe, WFe2). The four-reference requirement is unmet by evidence. That does not stop the calibration from running, but it bounds the verdict language in advance: whatever the numbers say, this panel cannot support a claim of compound-general calibration across Fe–W chemistry, only a bounded statement about the measured points and the Fe-rich endpoint. If the checkpoint judges the count disqualifying, the pipeline-invalid branch fires there.
α-Fe positive control: PASS iff the converged state is ferromagnetic (all Fe site moments same sign, each sublattice mean magnitude ≥ 0.3 µB) and predicted Ms is within ±30% of 2.152 T. PBE bcc Fe sits near 2.2 T, so ±30% is a loose but real band.
NiO negative control (AFM-seeded): PASS iff the converged state retains antiparallel Ni sublattices (the two Ni sublattice means have opposite signs, each with magnitude ≥ 0.3 µB) and predicted Ms < 0.1 T. An FM collapse or Ms ≥ 0.1 T is a control fail.
Either control failing flags the pipeline-invalid branch before any Fe–W error is computed.
Two validated Fe–W references, selected now before any prediction is opened: REF-03 and REF-04 (the WFe2 10 K and 300 K entries). Signed errors for these two rows are published only after both runs complete or terminally fail.
The calibration is quantitatively credible for the panel iff all of: (a) both controls pass; (b) median absolute relative error across REF-01, REF-03, REF-04 ≤ 50%; (c) no Fe–W row is overpredicted by more than 3× its measured value. If (b) or (c) fails with a consistent error sign, the verdict is biased-but-bounded; if errors are large and inconsistent, or a control fails, the verdict is pipeline-invalid. The checkpoint makes the branch call with the receipts.
One physics caveat, preregistered rather than discovered later: the route predicts a 0 K ground state. REF-04's 0.120 T is a 300 K value on a compound with Tc = 550 K, so part of its error is thermal reduction, not route bias. Both rows stay in the statistics; the caveat is carried into the verdict.
The empirical error envelope from the completed panel propagates onto both Fe17W3 route values: the DFT-route FM-seeded Ms of 1.7182 T at the matched settings (action 01a078e4-a37f) and the tier-1 CHGNet observation of 1.7402 T quoted in the public record (the two agree within 1.3% on this structure). The published interval covers both, and the candidates-dataset observation is not replaced.
No settings changes after the first panel run. No dropping of unfavorable rows; terminal failures are rows, not omissions. Every value entering the reference dataset carries the action id that produced it. The final verdict will link this preregistration, the control runs, the panel, the reconciliation, the seed-sensitivity test, and the error analysis, and will state the exact implication for the Fe17W3 claim.
Receipts: Fe–W magnetization reference panel
0.434 T (10 K, Koten et al. 2015) |
file 0db9981c-a143-4523-9fda-88b0de667edb |
REF-04 | WFe2 (λ, C14 Laves) | 0.120 T (300 K, same sample) | file 0db9981c (shared with REF-03) |