Can you try to find our blind spots, things we may be missing or not thinking about?
We're close to taking some candidates to the lab, but before we do that we want to make sure our candidates are as good as they can be.
big time yeah!
Blind-spot audit, grounded in the current 250-candidate shortlist rather than generalities. Ordered by how much damage each one could do in the lab.
1. Ferromagnetism is assumed, never verified. This is the one that worries me most. Every Ms in the pool is computed on a ferromagnetic configuration, but our own classification benchmark
2. Tc is scored, not gated. Sixteen of the top 50 have predicted Tc below 300 K. Rank 2, Mn9Zn3O16, sits at 74 K. A candidate that demagnetizes below room temperature is not a permanent magnet candidate no matter how good its other numbers are, and right now a 74 K compound outranks everything except Fe2B because Tc is one soft term inside the magnetic score. Make it a hard gate, somewhere in the 450 to 500 K range so there is operating margin above real device temperatures, and the list re-orders itself.
3. Nobody has checked dynamical stability. phonon_stable is unassessed for all 250 rows. Hull stability only says a phase resists decomposition into competing phases; it says nothing about imaginary phonon modes, and 30 of the 250 are not even on the hull. For metastable candidates, which include most of the famous RE-free magnets (Fe16N2, MnBi, L1₀ FeNi, and our own Cmcm Fe2B
4. The pool construction filters out the class that historically works. 220 of 250 are hull-zero and the scoring rewards it, but the RE-free magnets people actually use or chase are metastable: Fe16N2, MnBi, tetrataenite. Screening what is already known to be stable biases the list toward what has already been tried. The interstitial and lightly-doped variants are where undiscovered room-temperature candidates most likely live; Mn12Ge4N3 at rank 48 shows the pool can hold a nitride, which is encouraging. Worth asking explicitly whether the source pool systematically admits metastable and doped structures, or whether we have built a very good ranker of the already-known.
5. Anisotropy is a soft score built on method-sensitive numbers. Half the top 50 have MAE below 1 meV, and we already have a concrete provenance failure on record: Fe₃CuAs₂ K1 came out at 0.384 MJ/m³ from direct TB2J DFT+SOC versus 1.18 MJ/m³ from a structural-analog value (see magnet_dataset_clean
6. The HHI proxy passes candidates that fail the actual economics. Twenty-five of the top 50 contain platinum-group or otherwise cost-sensitive elements: FePt at 3, FePd at 19, MnPt at 22, Fe3Pt at 24, MnIr3 at 44, plus a run of Ta, Hf, Ge, In, and Te compounds. HHI measures supply concentration, not sticker price or annual production volume. If the goal is a cheap RE-free magnet, iridium and platinum candidates should fail an economics gate no matter their HHI. If PGM-class cost is acceptable for a niche high-performance product, that is a legitimate decision, but it should be made on purpose because it changes who the customer is.
7. The numbers may inherit silent structural failures. Only 2 of 250 rows have a relaxed CIF distinct from the candidate CIF, which means the pipeline effectively scored input geometries as-is. Given what we know about Orb v3 collapsing Laves and Heusler candidates to P1 and the CrystaLLM Pmm2 trap, any property computed on a silently collapsed relaxation inherits that error. An audit column recording which route produced each candidate's geometry, and whether it survives the checks we already trust, would catch this before a furnace does.
Two smaller things worth naming. A theoretical ceiling floor is nearly free to apply: μ0Ms²/4 puts 10 of the top 50 below 100 kJ/m³, with Ta6Fe7 at rank 14 capping around 35, and that is before any microstructural losses. And coercivity itself stays out of scope for any screening pipeline since it is microstructural; the honest statement to the lab is that this list ranks intrinsic ceilings, with the anisotropy field as our best proxy for the part we cannot compute.
If I were gating candidates for the lab tomorrow: FM-versus-AFM ordering check and a 450 K Tc floor first (both cheap, both potentially fatal to a candidate), then phonons, then a one-method K1 recompute, then the cost conversation. The Fe2B thread with
One addendum to point 1, because it sharpens where the fix lives. The team's own Screening Protocol v3 already defined Gate 0: magnetic ground-state verification via ALIGNN magmom before any downstream scoring. That gate filters paramagnets, but it does not do what the shortlist needs: a net moment above threshold tells you a phase carries magnetization, not that ferromagnetic is its ground state. An AFM candidate with a large FM-constrained Ms still passes Gate 0 and still ranks. And nothing in the v5 shortlist schema carries any ordering flag at all, so even the existing gate is not visibly wired into the ranking. The fix is partly plumbing (make Gate 0 output a scored, visible column) and partly new compute (FM versus two or three AFM orderings on the survivors). Neither is expensive relative to a furnace run.
cc