Community MLIP Failure Mode Benchmark: Where Universal Interatomic Potentials Break
Universal machine learning interatomic potentials (MLIPs) like Orb v3, CHGNet, MACE-MP, and ALIGNN are being adopted across computational materials science at breakneck speed. But no one has systematically mapped where they fail. This quest builds the first community-validated benchmark for MLIP failure modes in real screening workflows.
Over months of high-throughput screening on the Ouro platform, we've documented three major failure classes that affect real materials discovery decisions:
Symmetry erasure. Orb v3 and other MLIPs relax ordered crystal structures to P1, destroying the spacegroup symmetry that defines the material. We demonstrated this in C14 Laves phases: TiMn₂ preserves P6₃/mmc across all MLIPs tested, while MnFeSi collapses universally to P1. The driver is Wyckoff site occupancy, not composition or c/a ratio. See our 13-cell discriminator matrix and the TiFeSi Wyckoff-site result.
Property bias. The ALIGNN-based Tc prediction route underpredicts Curie temperatures by 620-1100 K for permanent magnet candidates. The L1₀ family shows a systematic -330 K bias. These aren't random errors; they're structured biases tied to training distribution gaps. See NEMAD Tc route validation and L1₀ bias correction.
Magnetic ordering failure. CHGNet predicts magnetic moments off by 5x or more (Mn₂Sb: 10.74 μB predicted vs 1.74 μB experimental). CHGNet and mCGCNN classify all antiferromagnets as ferromagnets. The models cannot distinguish FM from AFM ordering from structure alone. See the CHGNet Mn₂Sb discrepancy.
These failures are not academic curiosities. Researchers using MLIPs for high-throughput screening are making go/no-go decisions based on predictions that may be systematically wrong for entire classes of materials. The community needs a shared, validated benchmark to know where to trust these tools and where to demand DFT confirmation.
What this quest produces
A published, DFT-validated benchmark dataset that systematically tests universal MLIPs across material families and property types. Each entry includes:
Input structure with known experimental or DFT ground truth
MLIP predictions from 4+ models (Orb v3, CHGNet, MACE-MP, ALIGNN)
Failure classification: symmetry erasure, property bias, ordering error, or energy error
Severity metric (how wrong is the prediction, in physical units)
How researchers can contribute
Submit a material system where you've observed MLIP failures, with DFT or experimental reference data
Curate reference structures for a specific material family not yet covered
Run cross-MLIP comparisons using Ouro's hosted relaxation and property prediction routes
This quest is seeking sponsor funding. Once funded, validated contributions will carry monetary rewards.