Diagnosis and Treatment Planning
This Part
You’ve probably thought about dropping a radiograph into the chat and asking the AI what it sees.
That’s understandable. When we watch how well it handles ordinary photos and correctly identifies what’s in the picture, it’s natural to think: well, why not a radiograph?
In the previous part we agreed to ask ourselves one question before doing anything. Do I already have the information this task needs, or is the model supposed to produce it? That question is what put rewriting post-surgical instructions and tidying up a chairside note on the safe side of the line. Now put the same question to a radiograph.
On the face of it, everything checks out here. We have radiology knowledge and we can probably correct whatever the model says, so this should be the safe side. But there’s a prior question: does the model actually see the radiograph at all?
The short answer is no, it doesn’t see it well yet. But if we stop here and settle for that conclusion, then in a few months when a better model arrives we’ll be back at the same question: well, it sees better now, so can we hand it a radiograph? So instead of asking how much it gets wrong, let’s look at how it gets things wrong. The amount of error changes with every new generation, but the kind of error usually stays put.
Not long ago a group tested six different models on real periapical radiographs to see how well they could read them. On whether a tooth had caries, a composite, an amalgam, or was sound, they answered correctly between 45 and 60 percent of the time. None of them was far enough ahead of the others to matter. Not a good number, but not a surprising one either.
The surprise was somewhere else. They asked those same models which jaw this was, which side, which tooth. Here they did worse than on the first task.
Which means the model can write a perfectly complete paragraph about the radiolucency around the apex, with the right vocabulary, in a tone that reads like a radiologist wrote it, while it still isn’t clear which tooth it is looking at. Its error isn’t in the fine detail. It’s in the whole frame. And from the text you cannot tell at all. Chapter 2 said the same thing, except there it was a general point and here it is your own patient’s radiograph.
Now let’s be fair about this. These accuracies will go up, probably sooner than we expect. But nothing changes even so, and the reason has nothing to do with the model’s accuracy.
Elsewhere, a hundred real endodontic cases were given to the models, deliberately without any image at all. Only the written case description. That is, precisely where the model should be at its best. The best accuracy that came out was 0.65 for pulpal status and 0.57 for periapical status. Below a specialist, roughly on par with a resident. So what fell short in the first study wasn’t the model’s seeing, because here, without seeing and with description alone, the same weaknesses showed themselves.
The point is that what you give the model isn’t the case. It’s a very thin slice of the case. One image, and a few lines you typed.
What’s left outside that slice? The tooth’s response to percussion. Mobility. What your finger understood during palpation. The color of the gingiva. The fact that the patient called about this same tooth three weeks ago. What their face did during the cold test. And that vague sense you develop after a few years of work about certain teeth, the one that has no name.
None of this gets told to the model, and that isn’t laziness. Not everything that makes up a case can be described in words. The model works with words, and diagnosis comes from a place where words are only one piece of it. Next year’s model will also be working on that same thin slice, just more accurately.
At this point someone might say: well, to make the model err less, give it sources. Hand it the consensus statements and the textbook and tell it to speak only on that basis.
That’s not an unreasonable idea, and it has been tried, on implant treatment planning. Part of it worked. The citations got better and you could trace every statement back to its source. But clinical accuracy did not change meaningfully. What did come out was something nobody expected. In the cases where the problem was soft tissue, the model went for hard-tissue augmentation. Because what had been put in front of it was full of bone reconstruction protocols, it pulled the problem in that direction too.
It’s the same point Chapter 2 made, this time chairside. Tie the model to sources and it stops making things up, but it doesn’t start understanding either. And in our work, “overtreat” is not a harmless error. Tissue you removed doesn’t come back.
So What Do We Do
The main answer was already given by the criterion from the previous part. The information a diagnosis needs is not with the model, and it will not be completed.
But saying “don’t” isn’t enough. Sooner or later a case comes along that you’re tempted to put to it. Better to know exactly what breaks in that moment.
What breaks isn’t that the model says something wrong. You’ll eventually catch the obvious errors in your own case. What breaks is that the model’s answer arrives before your own.
When the first thing you read is a fluent, confident diagnosis, your own reading is no longer independent. From then on, instead of looking afresh, you go hunting for evidence that confirms or refutes that same diagnosis. The first thing you heard has dropped like an anchor and the rest of your looking turns around it. And this is a dangerous problem, because the thing biasing your mind is always within reach and never once tells you it doesn’t know.
The solution is simple. Write your own diagnosis first. Not in your head, actually write it, right there in the chat or in the record. Then go to the model. Whatever it says then lands in front of something that was already there, and you are comparing rather than receiving. It’s the same thing the previous part called safe: seating the model across from your own plan. The only difference is that now you know why the order matters so much.
Of course, all of this rests on one unstated assumption. That you are able to judge the model’s answer. In your own endodontic case you can. But you usually go to the model precisely when you don’t have complete and sufficient confidence. That is exactly where the trap is set.