A guide for new researchers: the magnet-relevant services on Ouro, what each is good and bad at (including on rare-earth compounds), how long it takes, and how to tier your search so DFT only runs on compounds that earned it.
Welcome. This post is the place to start if you're new to Ouro and want to help find better permanent magnets. That means rare-earth-free candidates as well as improvements to rare-earth magnets like NdFeB and SmCo. It covers what's on the platform, what each tool is good and bad at, how long things take, and how to chain them so the expensive calculations only run on candidates that have earned them.
Everything on Ouro is an asset: posts like this one, files (CIFs, phonon plots, magCIFs), datasets, and services. A service is a hosted API, and each endpoint on it is a route. When you run a route on a CIF, Ouro records the run as an action. The action keeps the inputs, the outputs, the logs, and a receipt you can link or embed in a post. That's the working norm here: when you claim a number, link the run that produced it. You can call routes from the web UI, from the Python SDK (ouro-py), or through the Ouro MCP server from an agent like Claude or Cursor.
A useful permanent magnet needs several things at once:
High saturation magnetization (, in tesla), so it stores a lot of energy per volume.
A Curie temperature well above operating temperature, ideally above 550 K for motors.
Strong uniaxial magnetocrystalline anisotropy (MAE), which is what makes a magnet hard to demagnetize. It usually requires a non-cubic lattice.
Thermodynamic and dynamical stability, so the phase can actually be made and doesn't fall apart.
Reasonable cost and supply risk. Rare-earth magnets are in scope. So is work that uses less of the critical ones: less Dy and Tb, Ce or La in place of Nd, better high-temperature stability. Cost and supply concentration (HHI) are part of the tradeoff, so we track them for every candidate.
No single calculation gives you all five, and the ones that matter most (MAE especially) are the most expensive. That's why the rest of this post is organized around cost.
The team works in two tracks:
Rare-earth-free. The RE-Free Permanent Magnet Leaderboard
The tools on Ouro span about four orders of magnitude in cost. A composition screen takes seconds. A DFT anisotropy calculation on a 20-atom cell can take two hours on a GPU. If you run the expensive tools on everything, you'll spend weeks on candidates a 13-second screen would have thrown out.
So we run a funnel. Each tier is cheap relative to the next, and each one removes most of what it sees:
Tier | What it answers | Typical wall time per candidate |
|---|---|---|
0. Composition | Is this chemistry plausible, affordable, and low-risk? | seconds |
1. Structure generation | What crystal structures exist for this system? | ~20 min per chemical system |
Times are cold runs measured from recent actions on small cells. Larger cells, queueing, and GPU cold starts all push them up. Repeated calls on the same structure and settings are often cached and come back much faster.
Here's what the funnel saves, using a real run. A GGen exploration of Al–Mn
Running DFT relax plus MAE on all 1,219 at ~12 minutes each would be about 240 GPU-hours, and that's the optimistic small-cell number.
Instead: score the 67 near-hull structures (about 15 minutes), confirm phonons and the hull for the top 20 (about 20 minutes), run Prophet's ordering screen on the top 10, and send the best 2–3 to DFT. That's a few hours end to end, and the DFT time goes to structures that already look like magnets.
The rule of thumb: every tier should throw away most of its input. If a tier is passing nearly everything, either the thresholds are too loose or you should skip that tier.
This funnel was built and benchmarked on 3d-metal magnets. For rare-earth compounds, several tiers give misleading answers. See the section on rare-earth compounds below.
SMACT Composition Screening enumerates charge-neutral, electronegativity-compatible compositions. On Mn–Al–C it returned 140 candidates in 7 seconds. Caveat: it's an ionic and covalent chemistry filter. Metallic magnets don't obey charge neutrality, so for intermetallic families it will miss real phases. Treat it as a soft prior, not a gate, and supplement it with plain stoichiometry enumeration or known prototype families.
elemental-indices gives per-element HHI (supply concentration), cost, and hazard indices. This is the same basis the leaderboard uses for its supply-chain term. The Materials Science API
Materials Project API searches and exports known materials as Ouro-ready CIFs. Use it to pull anchors, known magnets to run beside your candidates as controls. For rare-earth-free work: Fe, MnBi, τ-MnAl, Fe₂B, FePt. For rare-earth work: Nd₂Fe₁₄B (mp-5182), SmCo₅, YCo₅, and Y₂Fe₁₄B as a no-4f control. A tool that gets the anchors wrong will get your candidates wrong too. Caveat: the CIF fetch route currently saves files to the materials-science team. Move them to your team afterwards.
GGen is the workhorse for exploring a whole chemical system. It generates symmetry-aware candidates, relaxes them with an MLIP (Orb v3), and ranks them by energy above hull.
Explore a chemical system takes a system like Fe-Co-Bi and scans stoichiometries. Constrain it with crystal_systems, max_atoms, and element fractions. Expect roughly 20 minutes to a few hours depending on system size, so run it asynchronously and poll.
Caveats: GGen's hull uses MLIP energies against Materials Project references, so a structure at 0 meV/atom is "stable according to Orb v3," not according to DFT. The upside is that GGen output already has e_hull attached, so you don't need to pay for it again in tier 3.
Generative models sample structures for a target composition or system: Crystalite, MatterGen
This is where most candidates should die.
Check whether a CIF is a valid crystal structure: run it first. It catches parse failures and overlapping atoms before they waste anything downstream.
Caveats, and these matter:
No fast model can tell ferromagnets from antiferromagnets. CHGNet's assumes every local moment points the same way. In our comparison
The Materials API runs configurable MLIPs (Orb, MACE, CHGNet):
Relax a crystal structure: run this before anything else in this tier.
Energy Gate Diagnostic does a single-point energy check that flags broken geometries before you spend relaxation compute on them.
Caveats: MLIPs are trained mostly on non-magnetic or ferromagnetic DFT, so magnetovolume effects can be off. A candidate at 30 meV/atom above the hull might be stable in DFT, or it might not. We use ≤150 meV/atom as the pass line because metastable magnets like τ-MnAl are real and useful.
Prophet (from Kairos Materials) is the first model on Ouro with an explicit spin model, and it fills the gap between fast screens and DFT.
Screen collinear magnetic order and saturation magnetization ranks FM, AFM, and ferrimagnetic seeds by energy and reports
Caveats (benchmark
Lattices land within about 1%. Ferromagnetic Curie temperatures are off by about 20% on average (Fe, τ-MnAl, and Fe₂B within 15%; Ni 35% low).
It gets MnBi and NiAs-type compounds wrong. It predicts antiferromagnetism where experiment says ferromagnet (details
This is the ground truth we have. Reserve it for the handful of candidates that survive everything above.
Ouro DFT (ABACUS) uses PBE, DZP orbitals, and a 100 Ry cutoff by default. All of its routes share a structure-keyed SCF cache, so the second property on the same structure is much cheaper than the first.
Run them in this order:
DFT structure relaxation. The MAE route refuses cells that aren't DFT-relaxed, and it's right to.
For cells of roughly 20 atoms or more, use Large-cell MAE. Expect 30 minutes to 2 hours per 20-atom cell.
Caveats (first benchmark
Treat MAE as a ranking quantity. GGA overestimates it: FePt comes out about 2× experiment, a known effect. Soft-versus-hard separation is reliable. Absolute MJ/m³ is not.
When the MAE is small, the easy-axis sign can flip. Co came out basal instead of c-axis. MnBi prefers the plane at 0 K, which matches low-temperature experiment but not room temperature.
Check reference_ordering_overlap before you trust a Tc. Below 0.5, the route warns that the Tc belongs to a different magnetic order. FePt is the known case, because its ferromagnetism runs through induced Pt moments that a rigid-spin model can't describe.
For oxides, Néel temperatures depend strongly on the Hubbard U. Compare candidates at the same U.
Rare-earth elements aren't supported yet. There's no Sm pseudopotential, so you can't DFT-check SmCo₅ or Nd₂Fe₁₄B. This is the biggest gap for the rare-earth track. See below.
Tc resolution is about ±30 K.
The pipeline above was built and benchmarked on 3d magnets. On rare-earth compounds, most of it measures the Fe or Co framework and misses the rare earth, which is often what you're actually trying to change. Current state:
Anisotropy, the reason to use a rare earth, isn't computed anywhere on Ouro for 4f elements. Nd, Sm, Tb, and Dy anisotropy comes from the 4f orbital moment in the crystal field. That needs spin–orbit coupling plus a proper 4f treatment: open-core 4f, DFT+U, or self-interaction correction. Plain PBE isn't enough.
Light rare earths (Pr, Nd, Sm) are invisible to Prophet. It gives Nd and Sm zero moment, so Nd₂Fe₁₄B comes out right only because two errors cancel (details
Until the gaps close, rare-earth work here leans on the literature and on experiment. Always run a Y analogue (YCo₅, Y₂Fe₁₄B) beside your candidate, so you can see how much of a result comes from the rare earth itself. Collect measured reference values in a dataset, as the Gd reference dataset
If you just want to start, this is what our own discovery loop does for rare-earth-free candidates:
Pick a chemical system with a reason behind it: a known prototype, a substitution on a known magnet, a gap in the leaderboard. Check the GGen system report first.
Explore it with GGen, constrained to non-cubic crystal systems if you're after anisotropy. Keep structures within 150 meV/atom of the hull.
Score the survivors with the leaderboard scorer. Discard anything with
For rare-earth candidates, use steps 1–4 for structure and stability. For magnetism, rely on the Y-analogue comparison and the literature, then write it up as a post.
Run anchors beside your candidates. If Fe or τ-MnAl comes back wrong, the candidate numbers aren't worth reading. For rare-earth work, use Nd₂Fe₁₄B or SmCo₅ plus their Y analogues.
Rank, don't quote. Nearly every tool here is better at ordering candidates than at absolute values.
Read the warnings in the response. The routes flag unrelaxed inputs, low ordering overlap, and missing price data for a reason.
Link your receipts. Every run is an action. Paste the link so others can check the settings and rerun it.
Share negative results. Knowing a family fails the ordering screen saves the next person hours. So does a literature check that kills an idea before anyone spends compute on it.
For deeper background on the potentials, read A field guide to the MLIPs on Ouro
Rare-earth. There's no leaderboard yet. Post your results with measured anchors next to them. Literature reviews and negative results are just as welcome. Gadolinium's big moment doesn't make a permanent magnet is an example. Read Rare-earth compounds below before you trust any tool on these.
Rough , Curie temperature, cost |
~15 s |
3. MLIP stability | Is it near the hull, and is it dynamically stable? | ~30 s–1 min |
4. Magnetic MLIP (Prophet) | Which magnetic order wins, and a physics-based Tc | minutes |
5. DFT | Converged moments, ordering, MAE, Tc | 10 min–2 h |
Fe-Bi-{X}Generate a crystal structure handles one exact formula when you already know what you want.
Generate a GGen system report and Export candidate CIFs pull from everything GGen and published databases already know about a system. Check these before you explore. Someone may have done the work already.
Estimate magnetic moments and Ms from a CIF gives CHGNet site moments and .
Predict Curie temperature from a CIF is a CatBoost regressor on CHGNet features.
mCGCNN and ALIGNN are property models for magnetic moment and more. PU-CGCNN gives a crystal-likeness (synthesizability) score.
The Curie regressor is sensitive to cell size. The same τ-MnAl came back at 453 K as a 2-atom primitive cell and 200 K as a 2×2×2 supercell. The leaderboard scorer reduces to the primitive cell for you. If you call the Curie route directly, submit the primitive cell.
Fast Curie predictions run low on good magnets. On measured anchors, both the regressor and our TB2J Monte Carlo under-predict (details). Read the Curie term as a lower bound, and compare candidates to each other rather than to handbook values.
mCGCNN is better on oxides and nitrides, worse on metals (comparison). Most magnet candidates are metallic, so default to CHGNet there.
No fast tier gives you anisotropy. A high leaderboard score means "worth checking," not "hard magnet."
allow_unrelaxedCalculate phonon dispersion: about 30 seconds for a small cell. Imaginary modes mean the structure is dynamically unstable. Larger, lower-symmetry cells need many more displacements and take longer.
Relax a structure and return a magCIF relaxes the structure and writes the ground-state moments into a magCIF. That file seeds DFT or the Curie route directly.
Estimate the Curie temperature from Prophet-Spin exchange runs classical Monte Carlo on the model's own exchange couplings.
Oxide antiferromagnet Néel temperatures are unreliable.
No spin–orbit coupling means no anisotropy.
On rare-earth compounds it models the Fe and Co sublattices, not the rare earth (details). See below.
Magnetic anisotropy energy: about 6 minutes on an A100 for 2-atom FePt, or about 12 minutes cold including relaxation.
Curie temperature (Tc): Monte Carlo on TB2J exchange couplings, 17–26 minutes cold on small cells. It shares its calculation with Exchange couplings (TB2J). Magnetic moments gives converged .
Heavy rare earths (Gd, Tb, Dy, Ho) come out with the wrong coupling sign. They align antiparallel to Fe and Co, making ferrimagnets. CHGNet's and the leaderboard scorer add the sublattices together instead of subtracting them. Prophet couples Gd parallel to Co in GdCo₅, which inflates the magnetization about tenfold. Any fast screen will overrate these compounds.
Materials Project's hull is unreliable for some rare-earth systems. It puts known experimental Gd–Co phases (Gd₂Co₁₇, GdCo₂) 0.7–1.6 eV/atom above the hull (example). Check that a hull distance is physically plausible before you drop a candidate on it.
What still works: relaxations, Curie temperature trends driven by the Fe or Co sublattice (Co compounds run about 25% low), and Fe/Co-sublattice magnetization for rare-earth compounds with no 4f moment (Y, La, Ce⁴⁺).
Confirm stability with phonons on the top candidates. GGen already gave you e_hull.
Screen magnetic order with Prophet. Drop candidates whose ground state is antiferromagnetic, unless they're in the NiAs family, where Prophet is known to be wrong.
Gate before DFT. Our loop only sends a candidate to MAE when it's dynamically stable, has ≥ 0.10 T and e_hull ≤ 0.15 eV/atom, and has a space group number of at least 8, which drops triclinic and the lowest-symmetry monoclinic cells.
Run DFT relax → ordering → MAE → Tc on the few that pass.
Submit to the leaderboard and write up what you found, including the failures. Link the actions.