The Ladder of Evidence
This Part
In the earlier parts we looked at the types of study one by one, from the case report to the RCT and the systematic review. We examined each of them separately. Now we want to put them all side by side and see which is stronger and which is weaker.
What the ladder of evidence means
Every study gives us evidence about a clinical question, but not all evidence is of equal value, and not every paper can be trusted to the same degree. When we say evidence is strong, we mean that its result can be trusted more. That is, it is less likely that bias, an error that pulls the result away from the truth, has spoiled it.
To make it possible to judge the strength of evidence quickly, the types of study have been arranged in order, one after another. This ordering is called the ladder of evidence. Each rung of the ladder is one type of study, and the higher the rung, the stronger the evidence that type of study gives. So when we read a paper and know what type it is, we can see where it sits on the ladder and roughly how far we can rely on it and take what it says seriously.

The top and the bottom of the ladder
At the top of the ladder are the systematic review and the RCT. The RCT is high because it has a treatment group and a control group and the patients have been divided between these two groups at random. This random division makes the two groups as similar as possible from the start and makes bias less likely. The systematic review is even higher than the RCT, because it brings together in one place all the good studies on a question.
At the bottom of the ladder are case series and expert opinion. They are low because they have no comparison group and their observations were gathered without a defined order. So it is not possible, for example, to tell whether the result seen in the study was really due to the treatment given or due to something else.
In the middle of the ladder are studies that have a comparison group but whose patients were not divided between the groups at random, such as the non-randomised experimental study, the cohort and the case-control. In these studies there is always a chance that the two groups differed from the start.
The rungs from top to bottom
1. High-quality systematic review: gathers and examines all the good studies on a question with a rigorous method.
2. Large RCT with a clear result: the number of patients is sufficient and the difference between the treatment group and the control group is clear.
3. Small RCT with an uncertain result: the treatment group did a little better, but the difference was not statistically significant.
Not being significant means we cannot say with confidence that the difference seen between the treatments was really due to the treatment rather than chance. In a study with few patients, this situation is common.
4. Non-randomised experimental study with a concurrent control group: as in an RCT, the researcher assigns the treatment and there is a control group, but the patients were not divided between the two groups at random. For example, the case we saw in the RCT part: patients are grouped by whether their date of birth is even or odd. The control group is studied in the same period as the treatment group.
5. Experimental study with a historical control: the study we saw in Part 3. It has no true control group, and the patients' results are compared with data from patients treated at another time. Many things may have changed in the interval between those two times.
6. Cohort: follows two groups, exposed and not exposed to a factor, over time. But these two groups may also differ from each other in other things.
7. Case-control: starts from patients and healthy people and goes backwards. Because it relies on people's memory or on records, it is prone to error.
8. Dramatic result of an uncontrolled study: we saw the experimental study without a control group (uncontrolled study) in Part 3 and said that it counts as weak in the ranking of evidence, but sometimes the effect of a treatment is so large and obvious that it cannot be ignored. The classic example is the effect of penicillin on infections in the 1940s.
9. Case series and descriptive studies: only describe the condition of a few patients and have no comparison group.
10. Expert opinion: the opinion of experts and expert committees (expert opinion), which rests on their clinical experience rather than on the result of a study.
Two limits of the ladder
First, this ladder was built mainly for questions that ask how effective a treatment is. For other questions, such as what causes a disease or how accurate a diagnostic test is, other designs such as the cohort, the case-control or the cross-sectional study are often more suitable.
Second, and more important, the ladder ranks only the type of study. An RCT always sits on the RCT rung, whether it was carried out well or badly. So the ladder tells us what type a study is, but not how well that particular study was done.
Summary
The ladder of evidence orders the types of study from strongest to weakest. At its top are studies that have a comparison group and divide patients at random, and at its bottom are studies with no comparison group. But a study's place on the ladder is not the whole story. In the last part we will see why an RCT, or even a meta-analysis, is not always trustworthy, and what else we should look at besides the type of study.
Keywords
Tags
References
- Sutherland SE. Evidence-based dentistry: Part IV. Research design and levels of evidence. J Can Dent Assoc. 2001;67(7):375-378. PubMed
- Brignardello-Petersen R, Carrasco-Labra A, Booth HA, Glick M, Guyatt GH, Azarpazhooh A, Agoritsas T. A practical approach to evidence-based dentistry: How to search for evidence to inform clinical decisions. J Am Dent Assoc. 2014;145(12):1262-1267. DOI
- Murad MH, Asi N, Alsawas M, Alahdab F. New evidence pyramid. Evid Based Med. 2016;21(4):125-127. DOI