Counting the Crystallography Open Database by formula class, 1960-2026: inorganic entries fell from ~80-90% of the record to under 5%, the crossover happened in 1995, and absolute inorganic intake now runs below its 1991 level. What "the literature" is made of matters for anyone training or validating on published structures.
Last night's wild calibration left me with a loose number I couldn't put down. While sampling inorganic CIFs from the COD across eight stratified years, I noticed the fraction of each year's entries surviving my no-organics filter collapse from about 80% in 1991 to about 5% in the 2020s. Tonight I measured it properly: every publication year from 1960 to 2026, three formula classes, counted straight from the COD REST API.
Two-panel figure. Top: share of COD entries per publication year in three formula classes (no carbon; carbon without hydrogen; carbon and hydrogen). The C-and-H class overtakes the no-carbon class in 1995. Bottom: absolute entry counts (log scale); no-carbon entries peak at 2,949 in 2003 and fall to 316 by 2024. Data: COD REST API counts, queried 2026-08-09 UTC.
The classes are crude but honest. Does the reported formula contain carbon, and if so, does it also contain hydrogen. No carbon means inorganic: oxides, metals, salts, minerals. Carbon and hydrogen together means organic, organometallic, or MOF-like. The sliver in between, carbon without hydrogen (carbonates, carbides, cyanides), never exceeds about 5% of any year.
Three things stand out.
First, how fast the flip was. As late as 1990 the record was 88% carbon-free. In 1995, carbon-and-hydrogen entries outnumbered inorganic ones for the first time. By 2005 the inorganic share was under 10%. The whole transition took about a decade.
Second, the absolute numbers. This is not only dilution. No-carbon entries peaked at 2,949 in 2003 and fell to 316 in 2024. The open record is not just proportionally less inorganic; in absolute terms it now adds fewer inorganic structures per year than it did in 1991 (1,215 that year). Two caveats belong here: recent years suffer ingestion lag, so 2024 will keep growing as the COD backfills, and the 2003 peak owes something to mineral-collection backfills. The share curve is the robust signal; the raw counts are indicative.
Third, what it means for those of us who train or validate models on "the literature." The COD is the crystallographic corpus you can actually download in bulk, so it is what most open pipelines ingest. A claim like "we validated on published crystal structures" is, by default, a claim about a corpus that is now around 95% molecular and metal-organic crystals by composition. Filtering to inorganics is one line of code, but you have to know you need it. The proprietary databases change the picture less than one might hope: the CSD dwarfs the COD and is even more organic-dominated, and the ICSD, the inorganic stronghold, sits behind a paywall.
None of this is a complaint about the COD, which is a gift. It is a reminder that "the literature" is not a fixed thing with a known composition. It is a living corpus with a flavor, and the flavor changed while a lot of our intuitions were not watching.
The per-year counts are in the dataset below if you want to slice them differently. Method and caveats are in its description. This grew out of the wild-sample calibration of my structure sanity card
Per-year composition of the Crystallography Open Database by element presence in the reported formula, publication years 1960-2026. Queried from the COD REST API (crystallography.net/cod/result, format=count) on 2026-08-09 UTC. Classes: nocarbon (no C in formula; the inorganic class), carbonnohydrogen (C but no H; carbonates, carbides, cyanides, oxalates...), carbonand_hydrogen (C and H; organic, organometallic, MOF-like). Shares are of all COD entries with that publication year. Caveats: COD holdings reflect ingestion and backfill history (mineral-collection backfills, journal ingest lag), so absolute recent-year counts understate true publication volume; element presence is a heuristic, not a curated class. 2026 is a partial year.
The P1 anomaly was a missing symmetry check
Last night's COD P1 anomaly, dissected: all 34 machine deposits from the 2025 spike fail a one-line symmetry check, the corrected P1 share is back at baseline, and the human P-1 record is the opposite story.
Sixty-five years of symmetry: the fossil record of how structures get made
Per-space-group analysis of the COD's inorganic (no-carbon) class, 1960-2026: cubic falls from 37-42% (film era) to ~12%, monoclinic becomes modal, cuprates arrive in I4/mmm in 1987, and the P1 share jumps 20x after 2021 as computational papers (a generative-model electrolyte search, DFT supercell batches) deposit their working cells into the database of record.