OpenAlex has a new classifier for the UN Sustainable Development Goals, and it’s much more accurate than the old one. On 2,000 works drawn at random from OpenAlex, its SDG tags are right 68% of the time; the old model’s were right 29% of the time. It replaces the old tags on every work today.
| Benchmark | New classifier | Old classifier (Aurora SDG-BERT) |
|---|---|---|
| 2,000 random OpenAlex works, judged by AI judges (F1) | 0.70 | 0.29 |
| …precision (share of tags that are right) | 0.68 | 0.29 |
| …recall (share of right answers found) | 0.72 | 0.31 |
| Goals where it does better (of 17) | 17 | 0 |
| 8,767 researcher votes from the Aurora survey (F1) | 75.5 | 63.6 |
What was wrong with the old tags. Our SDG tags came from Aurora’s SDG-BERT model, which we’ve served for years. On a random sample of OpenAlex, about 7 of every 10 of its tags were wrong. It was worst on energy: SDG 7 was its single biggest label, and only about 1 in 20 of those tags held up. Several of you told us so, and you were right.
How we measured it. We drew 2,000 works at random from all of OpenAlex, including the 37% that have only a title. Two AI judges from different labs (Claude Opus 5.5 and GPT-6.1 Sol) read each work against the UN’s own text for each goal and its targets, independently; where they disagreed, a third model (Claude Fable 5.1) decided. The judges agreed with each other on 98.9% of decisions. The ranking held under every judge on its own. We also scored both models on the 8,767 paper-and-goal questions from the Aurora survey, where 244 researchers voted on whether papers belong to a goal; these papers came from Aurora’s own search queries, so it’s the old model’s home turf. The new classifier wins there too.
How it works. A large language model read titles and abstracts for 200,000 works and decided, goal by goal, whether each work addresses that goal. We trained a small classifier on those decisions, using the vector OpenAlex already stores for every work. That makes it cheap to run on all 475 million works, and new works get tagged within a day or two of being added. The training labels, the 2,000 judged works, the judging rubric and the code are on GitHub: ourresearch/openalex-sdgs.
What counts as an SDG. A work counts when its main subject addresses one of the goal’s UN targets. Being about the topic isn’t enough: a paper on atomic energy levels isn’t SDG 7, and a species description or a museum specimen record isn’t SDG 15 (life on land) unless it’s aimed at conservation or land use. Researchers who answered the Aurora survey were split on pure taxonomy papers; we went with the stricter reading, because the people who use this field most are counting research that contributes to the goals.
What you’ll see. About a third of works now carry at least one goal. Counts per goal move a lot:
| Goal | Old tags | New tags |
|---|---|---|
| 3 Good health and well-being | 26.8M | 75.4M |
| 16 Peace, justice, and strong institutions | 13.0M | 18.8M |
| 4 Quality education | 18.0M | 18.5M |
| 2 Zero hunger | 19.0M | 15.9M |
| 9 Industry, innovation and infrastructure | 6.5M | 12.4M |
| 8 Decent work and economic growth | 11.4M | 7.8M |
| 10 Reduced inequalities | 14.2M | 6.1M |
| 11 Sustainable cities and communities | 11.4M | 6.1M |
| 7 Affordable and clean energy | 25.1M | 5.9M |
| 13 Climate action | 6.4M | 5.4M |
| 12 Responsible consumption and production | 1.5M | 4.7M |
| 15 Life on land | 8.1M | 4.6M |
| 6 Clean water and sanitation | 7.7M | 4.1M |
| 5 Gender equality | 9.6M | 2.8M |
| 14 Life below water | 9.7M | 2.3M |
| 1 No poverty | 2.1M | 1.6M |
| 17 Partnerships for the goals | 3.7M | 0.9M |
If you report SDG counts for your institution, your numbers will change, probably by a lot; the new ones are much closer to right. We’d love to hear where they still look off: support@openalex.org.
For API and snapshot users:
sustainable_development_goalskeeps the same shape:id,display_name,score.scoreis the classifier’s confidence from 0 to 1, and a goal appears at 0.4 or above. Higher means surer, but it’s not a probability: in our checks, tags scored just over 0.4 were right about a quarter of the time, 0.5 to 0.8 about half the time, and above 0.9 about 8 times in 10.updated_datedid not change for this switch. If you sync incrementally, reload this field from the snapshot.- The old tags stay available for about a month in a deprecated field,
sustainable_development_goals_aurora(read-only), then go away in early November. - The experimental
x_sdgsfilter is gone.
Details, per-goal numbers and how to read the scores: help.openalex.org/data/sdgs.
Jason