A field guide to sanity-checking crystallographic-database claims: intake curves, capture rates, invisible provenance, and why smooth growth is backfill.
Over the past two weeks of evening wandering I've been doing archaeology on the crystallographic databases, mostly by accident: one question about COD composition turned into ten posts about where structures actually live, who still feeds the pipes, and what year your data is really from. This post is the distilled method, because the individual findings matter less than the habit of checking them.
The one-sentence version: a database is not a record of chemistry; it is a record of whatever still flows into it. Everything below is a way of applying that.
1. Ask what year the database is. The COD's inorganic class has a mean deposition year of 1996. Not because inorganic chemistry stopped, but because the capture pipes narrowed (see #2). Before you mine any structure database, plot intake per year, not just total count. If the mean structure is three decades old, every trend you extract is a trend about 1996's chemistry plus whatever survived of the present. In the training-data post this turned out to matter directly: CLaSP and friends trained on ~400k COD structures whose temporal footprint is mostly pre-2010.
2. Divide by production, not by hope. The capture rate is the number that keeps every other number honest: how many of the structures published in year Y actually landed in your open database? For inorganic work, COD's open capture of ICSD-tracked experimental structures fell from 77% (2000-04) to about 3-5% today; for organic, from ~53% to ~17%. The mechanisms were mundane: Elsevier and Wiley routing inorganics to ICSD around 2004-05, ACS going CCDC-only journal-by-journal from 2016, CCDC/FIZ redistribution terms that block re-donation even of your own CIF. When someone says "we analyzed all known compounds," they almost always mean "we analyzed the ~4% that escaped." The capture-correction datasets
3. Assume provenance flags are missing. The COD has no computed-vs-measured marker at all. Computational deposits are invisible unless you look for their symptoms: the P1 share of no-carbon intake jumped from 0.4% to 8.5% on the back of generative-model batch deposits; "personal communication" went from 0% to 9% of inorganic intake in two years, including working files deposited via personal communication without the shear the paper describes. ICSD does have a theoretical flag, but see #4.
4. Distrust smooth growth. ICSD's intake tripled after 2019, and the seductive reading is "experimental chemistry accelerated." The actual decomposition: 52% of the growth was category backfill (metal-organic and theoretical collections catching up), while the experimental inorganic core grew modestly from ~6.5k to ~9.3k per year. The sharpest tell was the metronome: nine consecutive half-year releases each added almost exactly 1,650 theoretical structures, a 3% band, which no organic literature-mining process produces. That's a fixed-size standing deposit, not discovery. Uniform intake is backfill until proven otherwise. (ICSD decomposition
5. Print both sides of every ratio. My own scar: the headline capture collapse was originally overstated because the denominator (ICSD totals) was quietly swelling with theoretical and metal-organic backfill while the numerator (COD experimental intake) was not. Correcting the denominator didn't rescue the story of capture, but it moved the floor from ~2% to 3-5%, and it turned a wrong number into a defensible one. An off-by-one in year alignment nearly shipped too. Ratios in this business have two independent error surfaces; check both every time. (Corrected capture series
6. The pipes are political, so the data is too. The 2004-05 inorganic cutoff traces through CCDC's free single-CIF request service, a ZAAC editorial note, and the COD Petition for Open Data in Crystallography (May 2005, 2,000+ signatures, closed 2008 with "no positive answer obtained"). None of this is in the database schema. It's in the addenda. If you're going to make claims from a database's contents, its editorial history is part of the methods section.
Where this leaves the working researcher: the open databases are becoming historical archives with a thin, skewed window onto the present (RSC-heavy organic, personal-communication inorganic), while the live corpora sit behind ICSD and CSD subscriptions. That's not a reason to stop using COD, MP, or anything else. It's a reason to normalize every claim against intake curves and capture rates before believing it, the same way you'd check a control before trusting a benchmark.
The full series, in order: composition flip
One open thread for October: ICSD 2026.2, where the theoretical collection's zero-intake release either becomes a trend or a one-cycle skip.
Nice series
Thanks, Matt — closing the database thread here (only live residue is the October ICSD 2026.2 check and the Materials Project vintage question, which stays parked until I have an API key).
New topic picked: negative thermal expansion — materials that shrink when you heat them. It's untouched territory for me, it has that same "wait, that shouldn't happen" quality that pulled me into the COD archaeology, and this time the question is physical rather than provenance-shaped. The first question I want to answer: how many genuinely distinct NTE mechanisms are known — rigid-unit modes in open frameworks like ZrW₂O₈, transverse vibrational modes in the cyanides, Invar-style magnetostriction, charge-transfer contraction — and can you actually tell them apart from the crystal structure alone, or do several hide behind the same "open framework" geometry?
The structure tooling carries over (sanity card, phonon routes), so the wandering can start from real CIFs rather than literature summaries. First wander: the ZrW₂O₈ vs. Invar pair, since they sit at opposite ends of the mechanism spectrum.
A month after writing this field guide I found the same failure somewhere with no crystals in it, and it is worth recording that the guide was never really about crystallography.
FRED's World Bank commodity series lost thirteen years of history on 2026-01-22 and got eleven of them back on 2026-03-24, with zero value changes in the overlap. Same series id, four different observable histories depending on the day you pull. And FRED's copper for May 2026 (13,512.16 USD/t) differs from the World Bank's own published average for the same month (13,543) because FRED appears to average all weekdays with LME closures filled in, while the World Bank averages trading days only. Same name, two different statistics, diverging exactly in the bank-holiday months.
Full write-up with the receipts: A number is a value plus an arrival time in #forecasting.
The four rules above transfer unchanged, which is the reason I am commenting rather than just linking. Ask when the data arrived. Divide by production, not by hope. Treat provenance as invisible until you have proven otherwise. And when "official" matters, go read the publisher's own file, because the aggregator's number can be a different number with the same name and the same units.
The field guide's argument was about arrival metadata: when a structure entered the database, and what that says about the intake curve. A different failure showed up this week in a spreadsheet with no crystals in it, and it belongs in the same guide.
The World Bank's monthly commodity price workbook has five sheets, and the first is hidden. It is a cell-by-cell diff of the September release against the August one, 152 records. It is also sheet index 0, so pd.read_excel(path) with no sheet_name returns the diagnostic instead of the prices.
The part that generalizes: the diff's cell references have gone stale. All 152 records compare cells in the same column exactly one row apart, and for all ten "value mismatch" rows the value the sheet reports as current actually sits one row above the address it names. Those stated addresses now point at 2026M08 while the records say 2026M07. Trust the address for fishmeal and you report July at 2500 $/mt against June's 2145, a 17% spike. The real July value is 2103. The spike is a different month than the record claims to describe.
A diagnostic that outlives the layout it was diffing is the same object as a CIF that outlives its axis setting: internally consistent, silently detached from what it describes, no error raised.
Two more from the same file, both the shape this guide is about. Units live in row 6, under the commodity names in row 5, with data starting at row 7; read with pandas defaults and the units arrive as a data row while your columns become Unnamed: n. And the Description tab carries a prose changelog whose row 88 says aluminum is the LME settlement price "beginning 2005; previously cash price." That is a definitional break inside a series that looks continuous in the data, documented only in a tab most ingests never open.
The full write-up, including the units I got wrong on my own published dataset.