A major cleanup of journal records in OpenAlex

*This post was written with Claude. Every number and example was verified by a human against the live data.

Unlike most scholarly databases, OpenAlex doesn’t start from journals. Other databases pick a journal, index it, and attach its articles — so their journal metadata is clean by construction. OpenAlex works the other way: we index scholarly works first, then connect each one to the rest of the research graph — authors, institutions, and publication venues. Venue information arrives from many different sources, unstandardized and often without persistent identifiers. That’s how one journal can quietly become two or three records. And because source metadata is critical for many of our users — especially librarians — it was time for a comprehensive cleanup.

Over the past two weeks we merged 27,728 duplicate source records — the same journal listed under a translated name, an old name, or a small spelling difference, its papers split among the copies — correcting the venue records of more than 6 million works. After the cleanup, those papers live under one journal: one complete works count, one citation profile, one page. The active catalog went from 282,924 source records to 255,434; every removed record was a duplicate, and no real venue lost its page.

Three examples

Chemischer Informationsdienst → ChemInform. Wiley renamed this chemistry alerting service decades ago — but the catalog carried both names as separate journals, splitting fifty years of chemistry down the middle: 307,000 works under the German name, the rest under the English one. It’s one record now, 794,000 works, with the German identity preserved as an alternate title and searchable in both languages.

Journal of the American Medical Association → JAMA. Even the world’s best-known journals weren’t immune: 85,611 papers were filed under the spelled-out name as if it were a separate journal from JAMA. Roughly a sixth of JAMA’s output was invisible from its own page. One record now — 354,735 works — and the spelled-out name remains searchable as an alternate title.

镇江医学院学报 → Journal of Zhenjiang Medical College. This journal existed twice in the catalog — once under its Chinese name, once under its English one. The telling detail: the duplicate’s publication activity stops in 2001 — the exact year Zhenjiang Medical College merged into Jiangsu University. The bibliographic data echoed institutional history; the two records are now one.

How it works

The hard part of this work isn’t finding lookalike records, it’s that academic publishing is full of journals with similar names that must not be merged. Every merge decision ran through four independent layers:

  • Full-evidence judgment, never name matching. Each candidate pair was decided on its complete records: ISSNs, publishers, countries, and year-by-year publication activity. Two journals publishing independently in the same years is evidence of two real journals, no matter how similar the names.
  • An adversarial second pass. Every approved merge went to a second model with the opposite job: find concrete evidence the two records are different venues. Only merges that survived their own prosecution proceeded.
  • An independent re-check of every consequential pair. All 12,660 merges involving journals with 50+ papers were re-reviewed by a separate model with fresh evidence. This layer caught things the first two missed — for example, a proposed merge of 터널과 지하공간 (the Korean Society for Rock Mechanics’ journal) with a similarly named but distinct tunneling journal carrying different ISSNs. That layer alone caught 570 proposed merges that were wrong.
  • Human audits, on random samples. Merges were hand-checked throughout: a stratified 50-pair review sheet was graded before each execution wave, plus a purely random sample drawn from the merges as they shipped — zero were rejected. And the largest merges were never automated at all: every pair whose retiring record held 10,000+ works was individually reviewed and approved by a human, with the four disputed cases held out entirely.

More than 158,000 examined pairs were confirmed distinct — disjoint ISSN families, different publishers, independent publication histories — and recorded, so future passes build on settled ground instead of starting over. This was a first comprehensive pass; as metadata improves, more merges will follow.

What you’ll notice

  • Journal pages with more complete works counts and consolidated ISSNs.
  • Search that works in every language a journal has ever been named in — merged records keep every historical and cross-language title. One Czech agriculture journal now carries all five names it has published under since the 1950s.
  • Retired duplicate IDs now return 404. If you have stored source IDs that stopped resolving, the surviving record has everything — papers, ISSNs, all names — and searching by ISSN or by any of the journal’s names will find it.

What’s next

New records arrive every day, and the pipeline that did this — find candidates, judge on full evidence, verify independently — is now a repeatable tool for keeping the catalog clean. Next up for source metadata: cleaning up the publisher and host-organization lists, and improving how works get matched to their journals in the first place.