Research validation
Benchmarks
What this build has been measured against, read from the evidence it ships with. Every module reaches a terminal state with a stated reason, and every difference is attributed to a cause rather than left as a bare failure.
Overlapping dataset evaluation— sealed expert ground truth exists for every subject and the engine has run against it, which is more than case runtime evidence; and ds004199 shares a contributing site with the MELD development cohort, which is less than an independent evaluation.
- Execution
- Permitted in preproduction research— The engine completing on a case says the software ran. It says nothing about whether the answer was right, and a cohort of successful executions is not a performance result.
- Dataset independence
- Overlaps
- External-validation claim
- Prohibited
- Clinical enablement
- Prohibited
ds004199 and the MELD development cohort share University Hospital Bonn, so any figure computed here measures detection and memorisation together. No analysis of this data can separate them, so no quantity of additional cases turns this into external validation. This is now measured rather than assumed, and it is the strongest form of the overlap rather than the weakest: ds004199 is single-site -- its own dataset_description.json records ethics approval from the Ethics Committee of the University of Bonn and names its pipeline bonn_fcd -- so it is not that some of the evaluated subjects might come from a training-contributing site, it is that all twenty-eight do. No subset of this cohort could be carved out as independent. What remains unknown is whether these particular scans were in MELD's training data: Bonn contributing to that cohort is a fact about the institution, ds004199 is a separately collected series, and subject-level overlap is possible, undemonstrated, and not excludable without MELD's own subject manifest, which is unpublished.
nothing here is prospective, nothing here was read by a clinician under clinical conditions, and no outcome was measured. This evidence cannot support clinical use and is not a step toward it.
The five evidence tiers, and where this one sits
CASE_RUNTIME_EVIDENCEthe engine ran on real data and produced an output. A statement about software, not about diagnostic performance.
DATASET_EVALUATIONoutputs were compared against sealed ground truth across a defined cohort, on a dataset independent of the training data.
OVERLAPPING_DATASET_EVALUATIONthis cohortthe same comparison, on a dataset that shares sites with the training cohort. The result measures detection and memorisation together and cannot separate them.
INDEPENDENT_EXTERNAL_VALIDATIONperformance measured on data from sites that contributed nothing to training, under a pre-registered analysis.
CLINICAL_VALIDATIONprospective evidence that use of the system changes clinical decisions or outcomes. Nothing in this repository approaches it.
- Lesion-level sensitivity
- 61.5%16 of 26 subjectsthe expert lesion maska subject counts as detected when at least one predicted cluster overlaps the expert mask by one voxel or more
- Median Dice
- 0.160over 26 subjectsthe expert lesion maskmedian overlap between the prediction and the expert mask, over the subjects scored against it
- False-positive clusters per patient
- 0.38mean over 26 subjectsthe expert lesion maskmean number of predicted clusters per subject that touch no part of the expert mask
63.0%17 of 27 subjects
Admitting the one contested case moves lesion sensitivity from 0.615 to 0.630. The headline is the primary figure; this is how far it moves under the one judgement call it rests on.
The named acquisition is genuinely absent, but a shipped image carries the mask's exact shape and its affine to 3.2e-4, and its sidecar SeriesDescription names the missing acquisition's own coil. A mask cannot be drawn on a grid nobody was looking at, so the shipped image is that scan under a different `acq` label. Registered through it the case is a hit, which a world-coordinate overlay would have reported as a miss.
The evidence is a grid match rather than the filename the rest of the cohort resolves by, and the admissibility policy that produced the primary figure should not be relaxed for the one case it excludes. Published beside it, not instead of it.
sub-00125, sub-00095, sub-00100. An empty candidate set is a model that produced nothing to look at. It is not a normal MRI and must never be rendered as one: every subject in this cohort has a surgically or histologically anchored lesion, so an empty prediction is a miss by construction.
- Total in locked split
- 28
- Eligible
- 28
- Executed
- 28
- Completed
- 28
- Excluded from evaluation
- 1
- Still executing
- 0
- Ineligible (data)
- 0
- Failed (model)
- 0
- Interrupted (infrastructure)
- 0
- Not started
- 0
Sealed expert lesion masks, hemisphere and lobe are recorded for every subject in the locked split. The split contains no lesion-negative subject, which is why the negative half of the confusion matrix is undefined rather than unmeasured.
- Labels published at
- 25 Aug 2026, 21:20:04
- Sealed subjects
- 28
- Quarantined completions
- 1
The boundary is the first commit that added app/config/detection/cohort-evidence.json, which is when the expert lesions became reachable from a build. A prediction frozen before it is sealed inference; one produced after it is not, whatever the ledger calls it.
sub-00004, sub-00009, sub-00020, sub-00034, sub-00043, sub-00048, sub-00061, sub-00062, sub-00064, sub-00066, sub-00074, sub-00092, sub-00095, sub-00099, sub-00100, sub-00103, sub-00145. Anything those attempts produce is runtime evidence and is excluded from every metric on this page. An attempt that finished before the label was committed is unaffected and still counts.
| Metric | Status | Value, or why there is none |
|---|---|---|
Eligible patients 28 in the locked test split | Computed | 28 |
Analysed (engine reached a verdict) of 28 eligible | Computed | 28 |
Admissible for evaluation of 28 completed | Computed | 27 |
Quarantined (no prediction predates the published label) completions excluded from every metric below | Computed | 1 |
Scored against the expert mask of 27 admissible | Computed | 26 |
Not scored against the expert mask | Partial |
sub-00048: measured, but not by the route the primary figure accepts: its expert mask names an acquisition the dataset does not ship, and the moving image was identified by grid identity instead. The comparison is published in the sensitivity analysis rather than folded into the primary. |
Lesion-level sensitivity 26 subjects scored against the mask | Computed |
a subject counts as detected when at least one predicted cluster overlaps the expert mask by one voxel or more. Reference standard: the expert lesion mask -- the radiologist-drawn *_FLAIR_roi.nii.gz, carried into the prediction's space through a rigid FLAIR->T1 registration recomputed locally from the two public images. Denominator 26 subjects. |
Lesion-level sensitivity, largest cluster only | Computed |
the harsher question: the largest predicted cluster, which is what a reader looks at first, overlaps the expert mask. Reference standard: the expert lesion mask -- the radiologist-drawn *_FLAIR_roi.nii.gz, carried into the prediction's space through a rigid FLAIR->T1 registration recomputed locally from the two public images. Denominator 26 subjects. |
Lesion-level PPV (cluster level) every predicted cluster across the scored subjects | Computed |
A predicted cluster counts as a true positive when it overlaps the expert mask by one voxel or more. A cluster is a connected component of at least 10 voxels: MELD works on a surface mesh, and the map back into the volume leaves one-voxel specks that are not places a reader gets sent to. The threshold is not tuned and not delicate: the cohort's components run 1, 1, ... 5, 6 and then nothing until 204, so any minimum from 7 to 204 yields the same 28 clusters. The discarded voxels are still counted against the mask. Reference standard: the expert lesion mask -- the radiologist-drawn *_FLAIR_roi.nii.gz, carried into the prediction's space through a rigid FLAIR->T1 registration recomputed locally from the two public images. Denominator 26 subjects. |
False-positive clusters per patient 26 subjects scored against the mask | Computed |
How many extra places per subject a reader would be sent for nothing. Same cluster definition as the row above: a connected component of at least 10 voxels. Counting specks instead would have reported 46 candidates across the cohort where MELD's own tables list 28, and all eighteen extras -- every one of them six voxels or smaller -- would have landed in this column. Reference standard: the expert lesion mask -- the radiologist-drawn *_FLAIR_roi.nii.gz, carried into the prediction's space through a rigid FLAIR->T1 registration recomputed locally from the two public images. Denominator 26 subjects. |
Dice against expert mask 26 subjects | Computed |
Median is the headline because the distribution is bimodal: subjects the model localises score 0.13 to 0.60, and subjects it misses score exactly zero. A mean over those two populations describes neither. Reference standard: the expert lesion mask -- the radiologist-drawn *_FLAIR_roi.nii.gz, carried into the prediction's space through a rigid FLAIR->T1 registration recomputed locally from the two public images. Denominator 26 subjects. |
IoU against expert mask 26 subjects | Computed |
Reference standard: the expert lesion mask -- the radiologist-drawn *_FLAIR_roi.nii.gz, carried into the prediction's space through a rigid FLAIR->T1 registration recomputed locally from the two public images. Denominator 26 subjects. |
Voxel sensitivity 26 subjects | Computed |
The fraction of expert-mask voxels the prediction covers. Not the same quantity as lesion-level sensitivity, and much lower: MELD localises a lesion far more often than it outlines one. Reference standard: the expert lesion mask -- the radiologist-drawn *_FLAIR_roi.nii.gz, carried into the prediction's space through a rigid FLAIR->T1 registration recomputed locally from the two public images. Denominator 26 subjects. |
Voxel positive predictive value 26 subjects | Computed |
Reference standard: the expert lesion mask -- the radiologist-drawn *_FLAIR_roi.nii.gz, carried into the prediction's space through a rigid FLAIR->T1 registration recomputed locally from the two public images. Denominator 26 subjects. |
Centroid distance 26 subjects | Computed |
Distance between the centre of mass of all predicted voxels and the centre of mass of the expert mask. The mean is four times the median because a subject whose prediction lands in the wrong hemisphere contributes 70 to 110 mm; that is a real miss, not an outlier to be trimmed. Reference standard: the expert lesion mask -- the radiologist-drawn *_FLAIR_roi.nii.gz, carried into the prediction's space through a rigid FLAIR->T1 registration recomputed locally from the two public images. Denominator 26 subjects. |
Hemisphere agreement 23 subjects with a non-empty prediction | Computed |
the sign of the x coordinate of the predicted centroid, against the hemisphere recorded in participants.tsv. Reference standard: the recorded hemisphere and lobe -- the hemisphere and lobe recorded in ds004199's participants.tsv, anchored in surgery and histopathology. Coarser than the mask, and available for subjects whose mask cannot be registered. |
Empty candidate sets of 26 subjects scored | Computed |
An empty candidate set is a model that produced nothing to look at. It is not a normal MRI and must never be rendered as one: every subject in this cohort has a surgically or histologically anchored lesion, so an empty prediction is a miss by construction. |
True positives (patient level) of 26 lesion-positive subjects scored | Computed | 16 |
False negatives (patient level) of 26 lesion-positive subjects scored | Computed | 10 |
True negatives (patient level) | Undefined for this cohort | UNDEFINED for this cohort, not unmeasured. All 28 ds004199 subjects here carry an expert-delineated lesion, so the cohort contains no lesion-negative subject and the true-negative cell does not exist. Any figure printed in this row would be invented. It is a property of the composition of ds004199, not a property of the model: measuring it needs a control cohort of subjects without FCD, which this dataset is not and does not contain. |
False positives (patient level) | Undefined for this cohort | UNDEFINED for this cohort, not unmeasured. All 28 ds004199 subjects here carry an expert-delineated lesion, so the cohort contains no lesion-negative subject and the true-negative cell does not exist. Any figure printed in this row would be invented. It is a property of the composition of ds004199, not a property of the model: measuring it needs a control cohort of subjects without FCD, which this dataset is not and does not contain. |
Specificity | Undefined for this cohort | UNDEFINED for this cohort, not unmeasured. All 28 ds004199 subjects here carry an expert-delineated lesion, so the cohort contains no lesion-negative subject and the true-negative cell does not exist. Any figure printed in this row would be invented. It is a property of the composition of ds004199, not a property of the model: measuring it needs a control cohort of subjects without FCD, which this dataset is not and does not contain. |
Negative predictive value | Undefined for this cohort | UNDEFINED for this cohort, not unmeasured. All 28 ds004199 subjects here carry an expert-delineated lesion, so the cohort contains no lesion-negative subject and the true-negative cell does not exist. Any figure printed in this row would be invented. It is a property of the composition of ds004199, not a property of the model: measuring it needs a control cohort of subjects without FCD, which this dataset is not and does not contain. |
Accuracy | Undefined for this cohort | UNDEFINED for this cohort, not unmeasured. All 28 ds004199 subjects here carry an expert-delineated lesion, so the cohort contains no lesion-negative subject and the true-negative cell does not exist. Any figure printed in this row would be invented. It is a property of the composition of ds004199, not a property of the model: measuring it needs a control cohort of subjects without FCD, which this dataset is not and does not contain. |
AUROC | Undefined for this cohort | UNDEFINED for this cohort, not unmeasured. All 28 ds004199 subjects here carry an expert-delineated lesion, so the cohort contains no lesion-negative subject and the true-negative cell does not exist. Any figure printed in this row would be invented. It is a property of the composition of ds004199, not a property of the model: measuring it needs a control cohort of subjects without FCD, which this dataset is not and does not contain. |
Calibration of the cluster score | Not computable | MELD reports a per-cluster 'confidence' and it is not a probability. It was never fitted to an outcome frequency on this cohort or any cohort here, so 40.45 does not mean a 40% chance of lesion and no reliability curve can be drawn from it without first calibrating on held-out data this evaluation does not have. The score is used only to rank clusters within one subject, which is the one thing an uncalibrated monotone score can honestly do. |
Lobe correctness | Not computable | MELD reports a Desikan-Killiany region and ds004199 records a lobe. The mapping is many-to-one and genuinely ambiguous for several regions, so resolving it by guessing would produce a correctness rate that means nothing. Hemisphere agreement, above, is the part of this question that has an unambiguous answer. |
Lesion-level sensitivity, sensitivity analysis the primary denominator plus the one contested case | Computed |
NOT the primary figure. The named acquisition is genuinely absent, but a shipped image carries the mask's exact shape and its affine to 3.2e-4, and its sidecar SeriesDescription names the missing acquisition's own coil. A mask cannot be drawn on a grid nobody was looking at, so the shipped image is that scan under a different `acq` label. Registered through it the case is a hit, which a world-coordinate overlay would have reported as a miss. The evidence is a grid match rather than the filename the rest of the cohort resolves by, and the admissibility policy that produced the primary figure should not be relaxed for the one case it excludes. Published beside it, not instead of it. Reported because a result that moves from 0.615 to 0.630 when the contested case is admitted has had its robustness demonstrated; one reported alone has had it asserted. |
Dice, sensitivity analysis | Computed |
The contested case is a hit with a low Dice, so admitting it raises sensitivity and lowers the median overlap. Both move, in opposite directions, and neither movement is large. |
Hemisphere agreement, sensitivity analysis | Computed |
Reference standard: Reference standard: the recorded hemisphere and lobe -- the hemisphere and lobe recorded in ds004199's participants.tsv, anchored in surgery and histopathology. Coarser than the mask, and available for subjects whose mask cannot be registered. |
Cluster-floor sensitivity 26 subjects, floors 1 to 100 | Computed |
the floor changes false-positive clusters per patient and cluster PPV, and nothing else. Lesion sensitivity is identical at every value from 1 to 100 -- no discarded component was ever the one that found a lesion -- and Dice, centroid distance and hemisphere agreement are computed from the voxel mask rather than the cluster decomposition, so the floor cannot reach them at all. The published constant sits on the plateau rather than on a slope, which is what makes it a measured choice rather than an arbitrary one. |
Sites in the evaluated cohort read from ds004199's own dataset_description.json | Computed |
every evaluated subject comes from a site that contributed to the cohort MELD Graph declares, so no subset of this result is independent and no further case from this dataset could make it so. The tier is OVERLAPPING_DATASET_EVALUATION at its strongest reading, not its weakest. STILL UNKNOWN: whether these particular scans were in MELD's training data. Bonn contributing to that cohort is a fact about the institution; ds004199 is a separately collected series. Subject-level overlap is possible, is not demonstrated, and cannot be excluded from anything recorded here -- it would need MELD's own subject manifest, which is not published. |
Measurement reproducibility the whole cohort measured twice, five registrations refitted | Computed |
The FLAIR->T1 fit is stochastic, so the same cohort measured twice does not return bit-identical floats. Measured rather than assumed: every detected/missed verdict, every denominator and every cluster count was identical across two full runs, and the largest movement in any case Dice was 0.0027. Numbers below the third decimal are fit noise and should not be read. |
Registration provenance every subject whose mask was carried into prediction space, primary and sensitivity analysis together | Computed |
An earlier pass reported this as unrecoverable, on the grounds that the transform was written inside the pipeline container and never exported. That was a misreading of what the transform is. The prediction comes back on the raw T1's own grid -- identical shape and affine, checked per subject -- so the transform needed is FLAIR->T1, a function of two images the public dataset ships. Recomputed locally it reproduces the one exported case to Dice 0.9928, and nothing had to survive the container. |
QC exclusions cases whose QC output was captured | Computed |
|
Infrastructure interruptions never attributed to a case; each one halted the cohort and left the subject retryable | Computed | 6recorded faults |
Runtime per completed case 28 completed reconstruction(s) | Computed |
|
Training-site independence recorded dataset and model provenance | Computed |
|
| Subject | State | Runtime | Evaluation | Notes |
|---|---|---|---|---|
| sub-00043 | Completed | 4.4 h | sealed | claimed again after its label was published. |
| sub-00066 | Completed | 4.2 h | quarantined | claimed again after its label was published. |
| sub-00048 | Completed | 4.4 h | sealed | claimed again after its label was published. |
| sub-00064 | Completed | 4.6 h | sealed | claimed again after its label was published. |
| sub-00061 | Completed | 3.6 h | sealed | claimed again after its label was published. |
| sub-00092 | Completed | 4.5 h | sealed | claimed again after its label was published. |
| sub-00099 | Completed | 5.2 h | sealed | claimed again after its label was published. |
| sub-00032 | Completed | 6.1 h | sealed | |
| sub-00121 | Completed | 4.8 h | sealed | |
| sub-00136 | Completed | 5.3 h | sealed | |
| sub-00125 | Completed | 6.1 h | sealed | |
| sub-00134 | Completed | 4.9 h | sealed | |
| sub-00006 | Completed | 5.9 h | sealed | |
| sub-00144 | Completed | 5.6 h | sealed | |
| sub-00020 | Completed | 7.9 h | sealed | claimed again after its label was published. |
| sub-00142 | Completed | 5.7 h | sealed | |
| sub-00107 | Completed | 6.0 h | sealed | |
| sub-00074 | Completed | 6.4 h | sealed | claimed again after its label was published. |
| sub-00128 | Completed | 4.6 h | sealed | |
| sub-00108 | Completed | 5.8 h | sealed | |
| sub-00009 | Completed | 5.6 h | sealed | claimed again after its label was published. |
| sub-00004 | Completed | 3.8 h | sealed | claimed again after its label was published. |
| sub-00095 | Completed | 5.9 h | sealed | claimed again after its label was published. |
| sub-00034 | Completed | 6.4 h | sealed | claimed again after its label was published. |
| sub-00103 | Completed | 5.2 h | sealed | claimed again after its label was published. |
| sub-00145 | Completed | 5.3 h | sealed | claimed again after its label was published. |
| sub-00062 | Completed | 5.2 h | sealed | claimed again after its label was published. |
| sub-00100 | Completed | 5.5 h | sealed | claimed again after its label was published. |
- Commit
- 3c4604fad6b1565980e9dee59833e097d423b98e
- Engine / dataset
- meld_graph on ds004199 (test)
- Ledgers read
- case_states.json, case_states.worker-host.json, case_states.worker-vm.json
- Generated
- 27 Aug 2026, 03:43:39
Loading benchmark evidence…