Two photographs come back from a forensic analysis with the same grade: F. The first is carrying GPS coordinates, a full camera serial number and the owner’s name in the metadata. The second was flagged by a statistical test measuring texture consistency across the image.
Those are not the same finding. The first is a fact you can read straight out of the file — it is either there or it is not, and no amount of argument changes it. The second is an inference from pixel statistics, and ordinary photographs trip it all the time: flat skies, studio backdrops, anything shot against a white wall.
Report both as “F” and you have thrown away the most useful thing you knew.
The two questions a grade is really answering
Any forensic result answers two questions at once, and most tools squeeze them into a single output.
How much did we find? More findings, or more serious ones, push the result further. That is the grade, and it is what everyone reports.
How much can the checks be trusted? That is a separate axis entirely. It has nothing to do with how bad the finding is. A check that reads a fact out of a file is very hard to fool. A check that infers something from statistics can be wrong on a perfectly innocent image, and the rate at which it is wrong is a property of the check, not of the file in front of it.
Once you separate them, things that looked equivalent stop being equivalent. A high-confidence C is a worse problem than a low-confidence F, and you would act on them differently. Collapse the two axes into one letter and the reader cannot tell which situation they are in.
Why you cannot reuse the detector’s own score
There is an obvious shortcut here, and it is wrong. Most detectors emit a number — a probability, a distance, a percentage. It is tempting to show that number as the confidence.
It is not one. A detector’s score says how strongly this particular file tripped the test. Confidence has to say how often this test is right. Those are different quantities, and only the second can be established by running the detector against files whose answers you already know.
Presenting the first as the second is how tools end up displaying “94% confident” about a check nobody has ever validated. The number is real; the claim attached to it is invented. In snapWONDERS’ forensic analysis, deriving a confidence from a detector’s own output is explicitly forbidden for exactly this reason.
The answer nobody publishes: “we have not measured that”
Here is where it gets uncomfortable. Once confidence has to come from measurement, you quickly find checks where no measurement exists. Nobody has run them against a labelled set. Nobody knows how often they are wrong.
There are three things you can do with a check like that. Quietly drop it. Guess a confidence and hope. Or show the finding and state plainly that its reliability is unmeasured.
We do the third. Some results carry a label that says, in effect, we have not measured how reliable this is, so we are not putting a number on it. It is not a fourth, lowest rating. It is a statement about the state of our own evidence.
That reads as an admission, and it is. It is also more useful than the alternative, because a fabricated confidence looks exactly like a real one and a reader has no way to tell them apart.
Nothing on a report should be decorative
The other rule worth stealing is simpler: if a check appears on the report, it has to count for something.
It is easy to end up with checks that are displayed but ignored — computed, printed, contributing nothing. That is the worst position available. The reader sees a finding and reasonably assumes it mattered, while nothing in the system is watching whether that check still works. A signal nobody scores is a signal nobody notices breaking.
So everything shown contributes. Where a check is weak it contributes very little — in some cases moving a result by a hundredth of a point, which is close to nothing, and deliberately so. But close to nothing and nothing are different claims, and the weight now says which one is meant.
What this costs, honestly
Two things, and both are worth being straight about.
It is more work. Every check needs a defensible weight and a recorded reason for its confidence, and “it seems about right” does not qualify. Several checks that had sat quietly on reports for a long time turned out, on measurement, to fire on almost every file — which means they were never telling anyone anything.
And it makes the tool look less certain. A report saying “high confidence” on some rows and “not measured” on others reads as less authoritative than one that says nothing about reliability at all. I think that is the correct trade. The confidence was always uneven; the only question was whether the reader got to see it.
If you are building anything that grades, classifies or scores, the question is worth asking of your own output: does your number say how much you found, or how much you should be believed? If it is trying to say both, it is probably saying neither clearly.
Kenneth Springer is the founder of snapWONDERS, a digital forensic analysis platform for images and video. The two-axis grading described here — a severity for what was found, a separate confidence for how much the check can be trusted — is how snapWONDERS reports every forensic result. snapWONDERS forensic analysis — no account required.

