100 benchmark runs of ABACUS on Modal found OpenBLAS starting its own threads inside every process. One MPI process per core with k-point parallelism gives identical results 5–6× faster on the same 8-core workers. Shipped and validated on every CPU route.
Ouro DFT (ABACUS) runs every calculation on an 8-core Modal worker. Until this change, each one ran as one MPI process with 8 OpenMP threads. That choice dates from an early observation that multi-process runs of the ABACUS eigensolver hung, and nobody had revisited it. We benchmarked it properly. It was leaving most of the hardware unused.
With the same 8-core workers and the same inputs, a production SCF gets 5–6× faster at 16–19% of the cost, and the answers are identical. Most of the gain comes from one environment variable. Along the way we learned a few things about why DFT codes that scale well in HPC papers didn't scale for us.
Update (2026-09-29): shipped. Every CPU route now runs as 8 MPI processes with k-point parallelism and OpenBLAS pinned to one thread. Validation for each route is in the last section.
Each run was a production-style SCF: ABACUS v3.10.1, LCAO with the DZP basis, 100 Ry, k-spacing 0.3 Å⁻¹, PBE, collinear spin, forces and stress on. The inputs were written by the service's own calculator class, so they match what the routes run. Only the launch changed: MPI processes × OpenMP threads per process, k-point parallelism (kpar), process pinning, and the OpenBLAS thread count. Every run had a 20-minute timeout so a hang would appear as a result. Every run was also checked against the production layout's energy and moment.
Test cells: bcc Fe (2 atoms), conventional NiO (8 atoms), a strained FeCo cell (16 atoms), and a 2×1×1 FeCo supercell (32 atoms). In total there were 100 runs on 8-, 16- and 32-core workers.
On the same 8-core worker, 8 processes × 1 thread with kpar=8 and OPENBLAS_NUM_THREADS=1 took NiO from 390 s to 63 s and the 32-atom FeCo cell from 1036 s to 192 s. Over 100 runs, no layout changed the total energy by more than 0.01 meV or the total moment by more than 0.001 μB. None of the multi-process runs hung.
The ABACUS binary links Ubuntu's openblas-pthread build. That build ignores OMP_NUM_THREADS and by default starts one thread per CPU it can see. We confirmed it: with OMP_NUM_THREADS=1, openblas_get_num_threads() still returns 8.
So every MPI process started its own pool of BLAS threads inside the eigensolver. Eight processes each running 8 BLAS threads on 8 cores is 64 threads fighting for 8 cores, in the part of the code that takes 85–95% of the time.
That one fact explained almost every odd result from the first two benchmark rounds:
Splitting one k-point's diagonalization across several processes looked catastrophic. On the 32-atom cell, 32 processes sharing 8 k-point groups took 744 s. With BLAS pinned it took 127 s.
Pinning processes to cores looked like an 11× win at 32 cores. That was mostly pinning containing the extra BLAS threads. With BLAS fixed, pinning is neutral to slightly worse.
More threads per process made things slower: 1×16 was slower than 1×8 and timed out on the 32-atom cell.
Fix: set OPENBLAS_NUM_THREADS to the thread count per process. This belongs in any ABACUS, Quantum ESPRESSO or VASP image built on distro OpenBLAS, and it costs nothing.
ABACUS parallelizes over k-points almost perfectly. Each group of processes diagonalizes its own subset of k-points, with little communication. OpenMP threads only help inside one diagonalization and in the real-space grid integration. With 12–35 irreducible k-points per cell, kpar equal to the process count is the right default. Setting kpar lower, so several processes share one diagonalization through ScaLAPACK, was always slower here, even after the BLAS fix.
ABACUS v3.10.1 prints a warning that kpar > 1 "has not been supported for lcao calculation". It doesn't reset the value, and every route we validated gives the same answers with it on (see below).
Modal placed our runs on AMD Zen, Intel Skylake-X and Intel Cooper Lake hosts. The container can see 24–48 host CPUs regardless of how many cores it requested. The old threaded layout was bimodal. On one Skylake-X host NiO took 62 s. On Zen hosts it took 390–423 s, and grid integration alone went from 30 s to 360–390 s. We haven't pinned down the cause; it's somewhere in OpenMP-threaded grid integration on those hosts. The multi-process layout was consistent across hosts: NiO took 62.7 and 65.5 s on Zen. All four production baselines in the final round happened to land on Zen hosts, so the 5–6× figures are against the old layout's typical slow case, not its best case. On a lucky Intel host the gap for NiO disappears; for the 32-atom cell it's still 2.9×.
Lesson: on shared cloud hardware, benchmark with repeats and record the host CPU. A single timing can be off by 7× in either direction.
DFT codes do scale to thousands of cores in HPC papers, but that's for systems with hundreds or thousands of atoms. There, a single diagonalization and the real-space grid are big enough to split. Our cells have 2–32 atoms. The per-section timings show where the ceiling is:
Diagonalization is 85–95% of the time, and the useful parallelism is mostly one process per k-point. Once there are more processes than irreducible k-points (12–35 for our cells), the extra processes help less.
Grid integration scales well: 410 s → 22 s from 1 to 32 processes on the 32-atom cell. It just isn't where the time goes.
Setup is about 1 s, so fixed overhead isn't the limit.
32 processes on a 32-core worker did reach 12.8× (32 atoms) and 17.8× (NiO) over production. But per-run cost comes out at 0.22–0.31× of production, against 0.16–0.19× for 8 processes on 8 cores. Those two 32-core runs also landed on faster Cooper Lake hosts, which flatters them. For screening, where many independent jobs run side by side, 8-core workers with 8 processes are the cheapest per result, so that's what we adopted. A 32-core option makes sense later for latency-sensitive single jobs, such as a DFT relax or MAE on a finalist.
Pushing further would mean a CPU-tuned ELPA/BLAS build rather than Ubuntu's generic packages, or the GPU eigensolver for large cells. Both are real projects. The fix above was nearly free.
CPU calculations launch as mpirun -np 8 with OMP_NUM_THREADS=1 and OPENBLAS_NUM_THREADS=1.
kpar is the number of processes, capped at the k-mesh size, so Gamma-only cells stay at 1.
The GPU worker keeps a single process, now with OpenBLAS threads matched to OpenMP.
The MAE calculator is unchanged.
SCF results are identical, so existing cache entries stay valid.
Setting ABACUS_LAYOUT=openmp restores the old layout without a code change.
Every CPU route ran twice on an 8-core worker, once per layout, each with an empty private cache so nothing carried over between runs:
Route (structure) | Agreement, old vs new layout | Wall time, old → new |
|---|---|---|
SCF, moments (Fe) | Identical energy; moment within 1e-5 μB | 67 → 21 s |
Relax (FeCo) | Identical final energy, same 4 ionic steps | 243 → 64 s |
The Curie Monte Carlo spread is the method's own resolution, not the layout. Its temperature step is 62.7 K, and the old layout gave 1003 K and 1066 K on two routes over the same exchange. Exchange and Curie barely speed up because TB2J's post-processing dominates their wall time, not ABACUS. That's the next target. The old code comment warned that multi-process PDOS post-processing hung. It didn't hang in any run.
Tiny cells gain little from the parallelism alone: bcc Fe went from about 18 s to 17 s in the first matrix. Pinning OpenBLAS still helps them.
Speedups depend on how many irreducible k-points a cell has. Gamma-only or very low-symmetry large cells will see less from kpar.
Timings are best of two runs on shared hardware; the full per-run data is below.
Best-of-two wall times for an Ouro DFT SCF (ABACUS v3.10.1, LCAO DZP, 100 Ry, kspacing 0.3, forces + stress) under the production layout and the fixed layouts, on Modal. relative_cost = (wall × cores) / production. All runs agree with production to ≤0.01 meV. Full per-run data: file 039a2753-a48b-4600-8a9d-02231f41afd1.
Energy within 1 meV; identical site moments |
461 → 76 s |
Bands (Fe, 31 bands × 212 k-points) | Max eigenvalue difference 0.1 meV | 83 → 38 s |
DOS / PDOS (Fe, NiO) | Fermi level within 0.02 meV; curves within 0.04–0.07% of peak | 454 → 99 s (NiO PDOS) |
Band gap (NiO) | 55.6 vs 56.7 meV, within the SCF threshold | 465 → 78 s |
Magnetic ordering (Fe, 9 configurations) | Same FM ground state; ΔE to next 86.369 vs 86.372 meV | 1056 → 318 s |
Exchange, TB2J (Fe) | J shells within 0.5%, J₀ within 0.12% | 912 → 774 s |
Curie (Fe) | Mean-field 1254 vs 1257 K; Monte Carlo 1066 vs 1006 K | 599 → 840 s |
Headline comparison: the table above.
First production run on the new layout: bcc Fe, kpar 8, 18 s in ABACUS.
The harnesses are bench_parallel.py (layout benchmark) and validate_layout.py (per-route check) in the ouro-dft repo.