Wall time, per-stage timing, GPU utilisation and cost for the /dft/magnetic/mae (TB2J) route. Cases are L1â‚€ FePt (2 and 16 atoms) and ferrimagnetic GdCoâ‚…. Runs compare A100 and CPU workers, the ELPA, genelpa and cusolver solvers, and MPI process layouts, before and after the TB2J kernel rewrite. Every run in a case must give the same MAE. MAE values here still use TB2J's former 5.1 eV band cut; the accuracy dataset has the corrected values. Costs use Modal list prices.
| note | kmesh | config | wall_s | formula | n_atoms | hardware |
|---|---|---|---|---|---|---|
| Stage 1: production TB2J before the rewrite | 8x8x6 | prod-a100 | 240.7 | FePt | 2 | A100 + 8 CPU |
| Stage 1: production TB2J before the rewrite | 8x8x6 | a100-cusolver | 254.9 | FePt | 2 | A100 + 8 CPU |
| Stage 1: production TB2J before the rewrite | 8x8x6 | a100-mpi8 | 185.8 | FePt | 2 | A100 + 8 CPU |
| Stage 1: production TB2J before the rewrite | 8x8x6 | cpu8-mpi8 | 288 | FePt | 2 | 8 CPU |
| Stage 1: production TB2J before the rewrite | 8x8x6 | cpu8-mpi8-kpar1 | 430.2 | FePt | 2 | 8 CPU |
| FePt after the kernel rewrite and H(R) cache | 8x8x6 | cpu8-mpi8 | 57.6 | FePt | 2 | 8 CPU |
| Stage 2: bigger cells, final TB2J kernel; MAE still with the 5.1 eV band cut | 4x4x3 | prod-a100 | 560.7 | FePt | 16 | A100 + 8 CPU |
| Stage 2: bigger cells, final TB2J kernel; MAE still with the 5.1 eV band cut | 4x4x3 | a100-cusolver | 283.8 | FePt | 16 | A100 + 8 CPU |
| Stage 2: bigger cells, final TB2J kernel; MAE still with the 5.1 eV band cut | 4x4x3 | cpu8-mpi8 | 340.5 | FePt | 16 | 8 CPU |
| Stage 2: bigger cells, final TB2J kernel; MAE still with the 5.1 eV band cut | 4x4x3 | cpu16-mpi16 | 320.9 | FePt | 16 | 16 CPU |
| Stage 2: bigger cells, final TB2J kernel; MAE still with the 5.1 eV band cut | 5x5x6 | prod-a100 | 417.3 | GdCo5 | 6 | A100 + 8 CPU |
| Stage 2: bigger cells, final TB2J kernel; MAE still with the 5.1 eV band cut | 5x5x6 | a100-cusolver | 309.9 | GdCo5 | 6 | A100 + 8 CPU |