The Type of Study Is Not the Whole Story
This Part
In this series we looked at the types of study one by one. In the previous part we put them all side by side on the ladder of evidence: the systematic review and the RCT at the top, case series and expert opinion at the bottom. At the end of that part we said that the ladder ranks only the type of study and does not tell us how well that study was done. In this final part we see what that means, and, when we read a paper, what else we should look at besides the type of study.
Where the ladder falls short
With the ladder, we judge how far a study can be trusted from its type alone. This approach is simple and often gives the right answer. But two studies of the same type are not necessarily equally good. An RCT may have been carried out badly, and an observational study may have a result about which there is almost no doubt.
That is why a method called GRADE (Grading of Recommendations Assessment, Development and Evaluation) was developed. GRADE measures how much confidence we can have in the available evidence on a question. In this method the type of study is only the starting point. After that, several other things are examined that can lower or raise this confidence. The result is reported at four levels: high, moderate, low and very low certainty. If you come across the word GRADE, or one of these four levels, in a systematic review or a guideline, this is what is meant.
What lowers certainty
Poor conduct. In the RCT part we saw that randomisation and blinding must be done properly. If they were not, being an RCT is not enough on its own.
An imprecise result (imprecision). No study states the effect of a treatment exactly, as a single definite number. What a study gives us is a range: it says the true effect of the treatment most likely lies somewhere between these two numbers. This range is called the confidence interval.
Suppose a study with a small number of patients has shown that a mouthwash reduces gingival inflammation, but its range says the true effect could be anything from “greatly reduces inflammation” to “slightly increases inflammation” (hypothetical example). This study does not actually tell us whether the mouthwash is useful or not, because both possibilities fit inside its range. Now suppose another study with a very large number of patients says the true effect lies between “slightly reduces inflammation” and “moderately reduces inflammation” (hypothetical example). This result is precise, because its range is narrow and wherever in it we take the true effect to be, the mouthwash is useful.
So the rule is simple: the wider this range, the less the result can be relied on. And if the range is so wide that it contains both benefit and harm, the study has in effect not answered our question.
Indirectness. Every study is designed to answer a specific question: in which patients, with which treatment, compared with what, and to measure which outcome. When we read a paper, we first need to understand exactly what that study’s own question was, and then see whether that question is the same as ours. For example, if a study measured the effect of a fluoride varnish on caries in children and we are making a decision about adult patients, the study answers our question only indirectly (hypothetical example). To use its result we have to assume that the same holds in adults, and the greater this gap, the less confident we are in that result for our own patient.
A real example: in a systematic review, five RCTs had examined whether very tight blood-glucose control in hospitalised patients reduces mortality. The result slightly favoured tight control. But in most of these studies randomisation and blinding had not been done properly. The range of the result was also so wide that the approach might have reduced mortality considerably, and it might equally have increased it. So although this evidence came from five RCTs, it cannot be trusted much.
What raises certainty
Sometimes it is the other way round. For example, we know that hip replacement benefits a patient who has severe osteoarthritis of the hip and finds walking difficult. This treatment has never been tested in an RCT and its evidence comes from observational studies. Yet we are highly confident about its benefit. So being lower on the ladder does not always mean less trust.
To show exactly this, in a new version of the evidence pyramid (the same ladder drawn as a triangle), the lines that separate the layers are not straight but wavy. That is, a study may sit above or below the usual place of its type, depending on its quality.

A systematic review has conditions too
In Part 8 we saw that the systematic review sits at the top of the ladder. But a systematic review is built by putting other studies together. So if the studies inside it are weak, the result of the review is weak too.
To judge a systematic review we look at two things. First, that the review itself was done properly, that is, it found all the relevant studies and selected them carefully. Second, what type the studies inside it are and how well they were carried out.
In a systematic review one more thing matters: consistency, that is, whether the studies put together in the review reached results close to one another. If one saw a large effect and another saw no effect at all, the pooled result can be trusted less.
A real example: a meta-analysis had put together a large number of case series on a severe injury of the aorta and concluded that one surgical approach had lower mortality than the others. But all the studies in it were case series, meaning none of them had a proper comparison group. Such a meta-analysis is not equivalent to a meta-analysis of several good RCTs, even though both are meta-analyses.
For this reason, in the new pyramid the systematic review has been taken off the top of the pyramid and drawn as a magnifying glass. That is, a systematic review is not a rung of the ladder but a tool with which we look at the other studies. However good the magnifying glass is, what we see through it is only the studies we have placed under it.
The DentCast Evidence Score (DES)
If you look at some DentCast articles, you will see a letter at the top, from A to E, and below them an explanation of where that grade came from. We call this assessment DES (DentCast Evidence Score). Its logic is the same as what we saw in this part: the type of study is the starting point, and the quality of its conduct raises or lowers it.
This assessment is done in three steps. First, the study receives a base score according to its type, that is, its place on the ladder. This ladder is separate for each type of question, because the right design for a treatment question differs from the right design for a diagnostic question.
Second, it is examined how well the study was carried out, for example in an RCT whether randomisation and blinding were done properly. The more flaws are found, the more the score is reduced. So an RCT with several serious flaws may receive a lower grade than a good cohort. A systematic review that did not assess the quality of the studies inside it also receives a low grade.
Third, if the paper did not state clearly what it should have, for example who funded the study, the score is reduced again.
In the end the paper is placed in one of five grades, from A (strong evidence) to E (at the level of opinion). The name of the journal and the reputation of the authors have no effect on this grade. If only the paper’s abstract was available, the grade is counted as provisional.
When a paper is in front of us
With everything we have seen in this series, when we read a paper we ask these few questions:
- What type of study is this, and where does it sit on the ladder?
- Was it carried out well? If it is an RCT, were randomisation and blinding done properly?
- Is the range of its result narrow, or so wide that it includes both benefit and harm?
- Is the question the study answered the same as our question? Was it done in patients like ours?
- If it is a systematic review or meta-analysis, what type are the studies inside it, and do their results agree with one another?
Summary
The type of study and its place on the ladder are the starting point, not the end of the job. Poor conduct and an imprecise result can weaken even RCT evidence, and if a study answered a question different from ours, its result can be carried over to our own patient with less confidence. On the other hand, evidence from lower rungs sometimes gives high certainty. Systematic reviews and meta-analyses, too, are only as trustworthy as the studies inside them. The DES assessment at DentCast is built on this same logic.
Keywords
Tags
References
- Sutherland SE. Evidence-based dentistry: Part IV. Research design and levels of evidence. J Can Dent Assoc. 2001;67(7):375-378. PubMed
- Brignardello-Petersen R, Carrasco-Labra A, Booth HA, Glick M, Guyatt GH, Azarpazhooh A, Agoritsas T. A practical approach to evidence-based dentistry: How to search for evidence to inform clinical decisions. J Am Dent Assoc. 2014;145(12):1262-1267. DOI
- Murad MH, Asi N, Alsawas M, Alahdab F. New evidence pyramid. Evid Based Med. 2016;21(4):125-127. DOI