You can now filter all of OpenAlex by study design: randomized controlled trials, systematic reviews, meta-analyses and more. 21 million works, in every field and language, carry one. Our randomized-trial tags are 99.7% precise, and on the same papers they’re more precise than PubMed’s, long the gold standard.
When you’re searching, you can focus on just the highest-quality evidence. And if you build evidence tools, you get study design for the whole literature, free, through the OpenAlex API.
Seven designs, every field
The designs come from PubMed’s own vocabulary, and we add nothing to it. What’s new is reach. PubMed tags the journals it indexes; we tag every work with an abstract. Half of the randomized trials we tag, 306,000, are in journals PubMed’s MEDLINE doesn’t index.
Here’s a graph how many works are tagged with each study design, in both OpenAlex and PubMed:

Every randomized trial is also a clinical trial, and every meta-analysis is also a systematic review, so a work can count twice. PubMed’s Observational Study tag dates only from 2014. Works without an abstract get no design.
Find the trials, see the evidence
Filter to one design. Of 34,395 papers that mention semaglutide, 533 are randomized trials; one filter returns just those, on openalex.org or in the API with filter=study_designs.id:randomized-controlled-trial. AI agents that search OpenAlex can do the same.
Group by design to see what kind of evidence a field has. Long COVID research is mostly observational: 31 observational studies for every randomized trial. Semaglutide already has more meta-analyses (1,341) than randomized trials.
Tagged papers on each topic, by design. Papers an OpenAlex search for the topic returns.

Share of each topic’s papers that carry a design; protocols left out. We tag conservatively, so real counts run higher, trials most of all.
Built on PubMed, tuned for precision
PubMed’s publication types are carefully indexed and trusted across medicine. But its tags sometimes describe the study a paper came from rather than the paper itself, so a methods paper or a later analysis of trial data can carry the trial’s tag. We tag only what the paper itself reports, and only when we’re confident.
How often each tag is right
Share of tags an independent judge confirms, on the same PubMed-indexed papers.

Where the tags differ
Here are some examples of papers that PubMed tags as RCTs, but we (correctly) don’t apply the RCT.
| Paper | What it is |
|---|---|
| Bland and Altman, Lancet 1986: measuring agreement | A statistics methods paper |
| TOAST stroke classification, 1993 | Definitions written for a trial |
| The Implicit Association Test, 1998 | Lab experiments that built a test |
| Oncotype DX gene test, NEJM 2004 | A gene test on tumors from one arm of an earlier trial |
| SAPS II severity score, JAMA 1993 | A risk score built from a cohort |
Our model is tuned for precision, not recall–when we say it’s an RCT, it’s an RCT, with 99.9% accuracy. PubMed captures some RCTs that our stingy approach misses, but at the cost of some accuracy.
Judge: Claude Opus 5.5, under written definitions, on 8,308 papers sampled so that every estimate is weighted to the whole index; nothing was tuned on that sample. Re-reading 436 open-access papers in full barely moved the verdicts.
Check our work
Everything behind these numbers is public: the code, the models, the samples and every judge’s verdict. One command reruns every table. github.com/ourresearch/openalex-study-designs