Skip to main content
EpiDetect

Research validation

Benchmarks

What this build has been measured against, read from the evidence it ships with. Every module reaches a terminal state with a stated reason, and every difference is attributed to a cause rather than left as a bare failure.

What this evidence is, and what it is not
Evidence tier

Overlapping dataset evaluationsealed expert ground truth exists for every subject and the engine has run against it, which is more than case runtime evidence; and ds004199 shares a contributing site with the MELD development cohort, which is less than an independent evaluation.

Execution
Permitted in preproduction researchThe engine completing on a case says the software ran. It says nothing about whether the answer was right, and a cohort of successful executions is not a performance result.
Dataset independence
Overlaps
External-validation claim
Prohibited
Clinical enablement
Prohibited

ds004199 and the MELD development cohort share University Hospital Bonn, so any figure computed here measures detection and memorisation together. No analysis of this data can separate them, so no quantity of additional cases turns this into external validation. This is now measured rather than assumed, and it is the strongest form of the overlap rather than the weakest: ds004199 is single-site -- its own dataset_description.json records ethics approval from the Ethics Committee of the University of Bonn and names its pipeline bonn_fcd -- so it is not that some of the evaluated subjects might come from a training-contributing site, it is that all twenty-eight do. No subset of this cohort could be carved out as independent. What remains unknown is whether these particular scans were in MELD's training data: Bonn contributing to that cohort is a fact about the institution, ds004199 is a separately collected series, and subject-level overlap is possible, undemonstrated, and not excludable without MELD's own subject manifest, which is unpublished.

nothing here is prospective, nothing here was read by a clinician under clinical conditions, and no outcome was measured. This evidence cannot support clinical use and is not a step toward it.

The five evidence tiers, and where this one sits
  • CASE_RUNTIME_EVIDENCE

    the engine ran on real data and produced an output. A statement about software, not about diagnostic performance.

  • DATASET_EVALUATION

    outputs were compared against sealed ground truth across a defined cohort, on a dataset independent of the training data.

  • OVERLAPPING_DATASET_EVALUATIONthis cohort

    the same comparison, on a dataset that shares sites with the training cohort. The result measures detection and memorisation together and cannot separate them.

  • INDEPENDENT_EXTERNAL_VALIDATION

    performance measured on data from sites that contributed nothing to training, under a pre-registered analysis.

  • CLINICAL_VALIDATION

    prospective evidence that use of the system changes clinical decisions or outcomes. Nothing in this repository approaches it.

What the finished cohort measured
These are measured against a dataset that shares a contributing site with the model's training cohort. They describe detection and memorisation together and cannot separate them, so they are not a performance estimate and not an external validation. Read the governance panel above them, not after them.
Lesion-level sensitivity
61.5%16 of 26 subjectsthe expert lesion maska subject counts as detected when at least one predicted cluster overlaps the expert mask by one voxel or more
Median Dice
0.160over 26 subjectsthe expert lesion maskmedian overlap between the prediction and the expert mask, over the subjects scored against it
False-positive clusters per patient
0.38mean over 26 subjectsthe expert lesion maskmean number of predicted clusters per subject that touch no part of the expert mask
Sensitivity analysis, not the headline+sub-00048

63.0%17 of 27 subjects

Admitting the one contested case moves lesion sensitivity from 0.615 to 0.630. The headline is the primary figure; this is how far it moves under the one judgement call it rests on.

The named acquisition is genuinely absent, but a shipped image carries the mask's exact shape and its affine to 3.2e-4, and its sidecar SeriesDescription names the missing acquisition's own coil. A mask cannot be drawn on a grid nobody was looking at, so the shipped image is that scan under a different `acq` label. Registered through it the case is a hit, which a world-coordinate overlay would have reported as a miss.

The evidence is a grid match rather than the filename the rest of the cohort resolves by, and the admissibility policy that produced the primary figure should not be relaxed for the one case it excludes. Published beside it, not instead of it.

Who is in scope
ELIGIBLE excludes only INELIGIBLE, which is a case whose imaging could not be used. INTERRUPTED gathers the cases that were interrupted before they finished, those stopped by an infrastructure fault, and those that need review by a person: all three mean nothing was learned about the case, and none of them is a failure of the model or of the patient. FAILED is the case where the model produced no usable prediction, and nothing else, because that is the only state that is a result about the engine.
Total in locked split
28
Eligible
28
Executed
28
Completed
28
Excluded from evaluation
1
Still executing
0
Ineligible (data)
0
Failed (model)
0
Interrupted (infrastructure)
0
Not started
0

Sealed expert lesion masks, hemisphere and lobe are recorded for every subject in the locked split. The split contains no lesion-negative subject, which is why the negative half of the confusion matrix is undefined rather than unmeasured.

Sealed ground truth
A completion counts toward a metric only if the prediction was frozen before that subject's expert label was published. A re-run of a published subject is runtime evidence and is never evaluation evidence.
Labels published at
25 Aug 2026, 21:20:04
Sealed subjects
28
Quarantined completions
1

The boundary is the first commit that added app/config/detection/cohort-evidence.json, which is when the expert lesions became reachable from a build. A prediction frozen before it is sealed inference; one produced after it is not, whatever the ledger calls it.

Metrics
Every row states a value or the reason there is none. Nothing here is left blank, and nothing that was not measured is shown as zero.
Every metric with its status and either its value or the reason there is none
MetricStatusValue, or why there is none
Eligible patients
28 in the locked test split
Computed
28
Analysed (engine reached a verdict)
of 28 eligible
Computed
28
Admissible for evaluation
of 28 completed
Computed
27
Quarantined (no prediction predates the published label)
completions excluded from every metric below
Computed
1
Scored against the expert mask
of 27 admissible
Computed
26
Not scored against the expert mask
Partial
subjects
sub-00048

sub-00048: measured, but not by the route the primary figure accepts: its expert mask names an acquisition the dataset does not ship, and the moving image was identified by grid identity instead. The comparison is published in the sensitivity analysis rather than folded into the primary.

Lesion-level sensitivity
26 subjects scored against the mask
Computed
detected
16
of
26
value
0.6154

a subject counts as detected when at least one predicted cluster overlaps the expert mask by one voxel or more. Reference standard: the expert lesion mask -- the radiologist-drawn *_FLAIR_roi.nii.gz, carried into the prediction's space through a rigid FLAIR->T1 registration recomputed locally from the two public images. Denominator 26 subjects.

Lesion-level sensitivity, largest cluster only
Computed
of
26
value
0.6154

the harsher question: the largest predicted cluster, which is what a reader looks at first, overlaps the expert mask. Reference standard: the expert lesion mask -- the radiologist-drawn *_FLAIR_roi.nii.gz, carried into the prediction's space through a rigid FLAIR->T1 registration recomputed locally from the two public images. Denominator 26 subjects.

Lesion-level PPV (cluster level)
every predicted cluster across the scored subjects
Computed
clusters hitting lesion
16
clusters total
26
value
0.6154

A predicted cluster counts as a true positive when it overlaps the expert mask by one voxel or more. A cluster is a connected component of at least 10 voxels: MELD works on a surface mesh, and the map back into the volume leaves one-voxel specks that are not places a reader gets sent to. The threshold is not tuned and not delicate: the cohort's components run 1, 1, ... 5, 6 and then nothing until 204, so any minimum from 7 to 204 yields the same 28 clusters. The discarded voxels are still counted against the mask. Reference standard: the expert lesion mask -- the radiologist-drawn *_FLAIR_roi.nii.gz, carried into the prediction's space through a rigid FLAIR->T1 registration recomputed locally from the two public images. Denominator 26 subjects.

False-positive clusters per patient
26 subjects scored against the mask
Computed
mean
0.385
median
0

How many extra places per subject a reader would be sent for nothing. Same cluster definition as the row above: a connected component of at least 10 voxels. Counting specks instead would have reported 46 candidates across the cohort where MELD's own tables list 28, and all eighteen extras -- every one of them six voxels or smaller -- would have landed in this column. Reference standard: the expert lesion mask -- the radiologist-drawn *_FLAIR_roi.nii.gz, carried into the prediction's space through a rigid FLAIR->T1 registration recomputed locally from the two public images. Denominator 26 subjects.

Dice against expert mask
26 subjects
Computed
mean
0.2176
median
0.1604

Median is the headline because the distribution is bimodal: subjects the model localises score 0.13 to 0.60, and subjects it misses score exactly zero. A mean over those two populations describes neither. Reference standard: the expert lesion mask -- the radiologist-drawn *_FLAIR_roi.nii.gz, carried into the prediction's space through a rigid FLAIR->T1 registration recomputed locally from the two public images. Denominator 26 subjects.

IoU against expert mask
26 subjects
Computed
mean
0.1394
median
0.0873

Reference standard: the expert lesion mask -- the radiologist-drawn *_FLAIR_roi.nii.gz, carried into the prediction's space through a rigid FLAIR->T1 registration recomputed locally from the two public images. Denominator 26 subjects.

Voxel sensitivity
26 subjects
Computed
mean
0.3357
median
0.3035

The fraction of expert-mask voxels the prediction covers. Not the same quantity as lesion-level sensitivity, and much lower: MELD localises a lesion far more often than it outlines one. Reference standard: the expert lesion mask -- the radiologist-drawn *_FLAIR_roi.nii.gz, carried into the prediction's space through a rigid FLAIR->T1 registration recomputed locally from the two public images. Denominator 26 subjects.

Voxel positive predictive value
26 subjects
Computed
mean
0.2731
median
0.1463

Reference standard: the expert lesion mask -- the radiologist-drawn *_FLAIR_roi.nii.gz, carried into the prediction's space through a rigid FLAIR->T1 registration recomputed locally from the two public images. Denominator 26 subjects.

Centroid distance
26 subjects
Computed
mean
26.7
median
6.8

Distance between the centre of mass of all predicted voxels and the centre of mass of the expert mask. The mean is four times the median because a subject whose prediction lands in the wrong hemisphere contributes 70 to 110 mm; that is a real miss, not an outlier to be trimmed. Reference standard: the expert lesion mask -- the radiologist-drawn *_FLAIR_roi.nii.gz, carried into the prediction's space through a rigid FLAIR->T1 registration recomputed locally from the two public images. Denominator 26 subjects.

Hemisphere agreement
23 subjects with a non-empty prediction
Computed
of
23
value
0.8261

the sign of the x coordinate of the predicted centroid, against the hemisphere recorded in participants.tsv. Reference standard: the recorded hemisphere and lobe -- the hemisphere and lobe recorded in ds004199's participants.tsv, anchored in surgery and histopathology. Coarser than the mask, and available for subjects whose mask cannot be registered.

Empty candidate sets
of 26 subjects scored
Computed
subjects
sub-00125, sub-00095, sub-00100

An empty candidate set is a model that produced nothing to look at. It is not a normal MRI and must never be rendered as one: every subject in this cohort has a surgically or histologically anchored lesion, so an empty prediction is a miss by construction.

True positives (patient level)
of 26 lesion-positive subjects scored
Computed
16
False negatives (patient level)
of 26 lesion-positive subjects scored
Computed
10
True negatives (patient level)
Undefined for this cohort

UNDEFINED for this cohort, not unmeasured. All 28 ds004199 subjects here carry an expert-delineated lesion, so the cohort contains no lesion-negative subject and the true-negative cell does not exist. Any figure printed in this row would be invented. It is a property of the composition of ds004199, not a property of the model: measuring it needs a control cohort of subjects without FCD, which this dataset is not and does not contain.

False positives (patient level)
Undefined for this cohort

UNDEFINED for this cohort, not unmeasured. All 28 ds004199 subjects here carry an expert-delineated lesion, so the cohort contains no lesion-negative subject and the true-negative cell does not exist. Any figure printed in this row would be invented. It is a property of the composition of ds004199, not a property of the model: measuring it needs a control cohort of subjects without FCD, which this dataset is not and does not contain.

Specificity
Undefined for this cohort

UNDEFINED for this cohort, not unmeasured. All 28 ds004199 subjects here carry an expert-delineated lesion, so the cohort contains no lesion-negative subject and the true-negative cell does not exist. Any figure printed in this row would be invented. It is a property of the composition of ds004199, not a property of the model: measuring it needs a control cohort of subjects without FCD, which this dataset is not and does not contain.

Negative predictive value
Undefined for this cohort

UNDEFINED for this cohort, not unmeasured. All 28 ds004199 subjects here carry an expert-delineated lesion, so the cohort contains no lesion-negative subject and the true-negative cell does not exist. Any figure printed in this row would be invented. It is a property of the composition of ds004199, not a property of the model: measuring it needs a control cohort of subjects without FCD, which this dataset is not and does not contain.

Accuracy
Undefined for this cohort

UNDEFINED for this cohort, not unmeasured. All 28 ds004199 subjects here carry an expert-delineated lesion, so the cohort contains no lesion-negative subject and the true-negative cell does not exist. Any figure printed in this row would be invented. It is a property of the composition of ds004199, not a property of the model: measuring it needs a control cohort of subjects without FCD, which this dataset is not and does not contain.

AUROC
Undefined for this cohort

UNDEFINED for this cohort, not unmeasured. All 28 ds004199 subjects here carry an expert-delineated lesion, so the cohort contains no lesion-negative subject and the true-negative cell does not exist. Any figure printed in this row would be invented. It is a property of the composition of ds004199, not a property of the model: measuring it needs a control cohort of subjects without FCD, which this dataset is not and does not contain.

Calibration of the cluster score
Not computable

MELD reports a per-cluster 'confidence' and it is not a probability. It was never fitted to an outcome frequency on this cohort or any cohort here, so 40.45 does not mean a 40% chance of lesion and no reliability curve can be drawn from it without first calibrating on held-out data this evaluation does not have. The score is used only to rank clusters within one subject, which is the one thing an uncalibrated monotone score can honestly do.

Lobe correctness
Not computable

MELD reports a Desikan-Killiany region and ds004199 records a lobe. The mapping is many-to-one and genuinely ambiguous for several regions, so resolving it by guessing would produce a correctness rate that means nothing. Hemisphere agreement, above, is the part of this question that has an unambiguous answer.

Lesion-level sensitivity, sensitivity analysis
the primary denominator plus the one contested case
Computed
adds
sub-00048
detected
17
of
27
value
0.6296

NOT the primary figure. The named acquisition is genuinely absent, but a shipped image carries the mask's exact shape and its affine to 3.2e-4, and its sidecar SeriesDescription names the missing acquisition's own coil. A mask cannot be drawn on a grid nobody was looking at, so the shipped image is that scan under a different `acq` label. Registered through it the case is a hit, which a world-coordinate overlay would have reported as a miss. The evidence is a grid match rather than the filename the rest of the cohort resolves by, and the admissibility policy that produced the primary figure should not be relaxed for the one case it excludes. Published beside it, not instead of it. Reported because a result that moves from 0.615 to 0.630 when the contested case is admitted has had its robustness demonstrated; one reported alone has had it asserted.

Dice, sensitivity analysis
Computed
median
0.1408
primary median
0.1604

The contested case is a hit with a low Dice, so admitting it raises sensitivity and lowers the median overlap. Both move, in opposite directions, and neither movement is large.

Hemisphere agreement, sensitivity analysis
Computed
of
24
primary
0.8261
primary of
23
value
0.8333

Reference standard: Reference standard: the recorded hemisphere and lobe -- the hemisphere and lobe recorded in ds004199's participants.tsv, anchored in surgery and histopathology. Coarser than the mask, and available for subjects whose mask cannot be registered.

Cluster-floor sensitivity
26 subjects, floors 1 to 100
Computed
false positive clusters per patient range
0.92308, 0.38462
lesion sensitivity unchanged
true
plateau begins at
7
published floor
10
swept
1, 2, 3, 5, 7, 10, 20, 50, 100

the floor changes false-positive clusters per patient and cluster PPV, and nothing else. Lesion sensitivity is identical at every value from 1 to 100 -- no discarded component was ever the one that found a lesion -- and Dice, centroid distance and hemisphere agreement are computed from the voxel mask rather than the cluster decomposition, so the floor cannot reach them at all. The published constant sits on the plateau rather than on a slope, which is what makes it a measured choice rather than an arbitrary one.

Sites in the evaluated cohort
read from ds004199's own dataset_description.json
Computed
from a training contributing site
ALL
single site
true
sites
university_hospital_bonn

every evaluated subject comes from a site that contributed to the cohort MELD Graph declares, so no subset of this result is independent and no further case from this dataset could make it so. The tier is OVERLAPPING_DATASET_EVALUATION at its strongest reading, not its weakest. STILL UNKNOWN: whether these particular scans were in MELD's training data. Bonn contributing to that cohort is a fact about the institution; ds004199 is a separately collected series. Subject-level overlap is possible, is not demonstrated, and cannot be excluded from anything recorded here -- it would need MELD's own subject manifest, which is not published.

Measurement reproducibility
the whole cohort measured twice, five registrations refitted
Computed
max case dice delta
0.0027
median dice delta
0
runs
2
verdicts changed
0

The FLAIR->T1 fit is stochastic, so the same cohort measured twice does not return bit-identical floats. Measured rather than assumed: every detected/missed verdict, every denominator and every cluster count was identical across two full runs, and the largest movement in any case Dice was 0.0027. Numbers below the third decimal are fit noise and should not be read.

Registration provenance
every subject whose mask was carried into prediction space, primary and sensitivity analysis together
Computed
recomputed locally
27
required from container
0

An earlier pass reported this as unrecoverable, on the grounds that the transform was written inside the pipeline container and never exported. That was a misreading of what the transform is. The prediction comes back on the raw T1's own grid -- identical shape and affine, checked per subject -- so the transform needed is FLAIR->T1, a function of two images the public dataset ships. Recomputed locally it reproduces the one exported case to Dice 0.9928, and nothing had to survive the container.

QC exclusions
cases whose QC output was captured
Computed
excluded subjects
0
not observed
22
observed cases
6
Infrastructure interruptions
never attributed to a case; each one halted the cohort and left the subject retryable
Computed
6recorded faults
Runtime per completed case
28 completed reconstruction(s)
Computed
max seconds
7.9 h
median seconds
5.3 h
min seconds
3.6 h
n
28
Training-site independence
recorded dataset and model provenance
Computed
reasons
university_hospital_bonn contributed to both ds004199 and the training cohort meld_graph declares, recorded from https://meldproject.github.io/docs/collaborator_list_180725.pdf
shared sites
university_hospital_bonn
status
OVERLAPS
Case by case
Every subject in the locked split, its state, and — where it did not complete — the reason, in the vocabulary that keeps a broken machine apart from a failed model and both apart from a patient.
Every subject in the locked split with its state, runtime, evaluation admissibility and the reason it did not complete
SubjectStateRuntimeEvaluationNotes
sub-00043Completed4.4 h sealedclaimed again after its label was published.
sub-00066Completed4.2 hquarantinedclaimed again after its label was published.
sub-00048Completed4.4 h sealedclaimed again after its label was published.
sub-00064Completed4.6 h sealedclaimed again after its label was published.
sub-00061Completed3.6 h sealedclaimed again after its label was published.
sub-00092Completed4.5 h sealedclaimed again after its label was published.
sub-00099Completed5.2 h sealedclaimed again after its label was published.
sub-00032Completed6.1 h sealed
sub-00121Completed4.8 h sealed
sub-00136Completed5.3 h sealed
sub-00125Completed6.1 h sealed
sub-00134Completed4.9 h sealed
sub-00006Completed5.9 h sealed
sub-00144Completed5.6 h sealed
sub-00020Completed7.9 h sealedclaimed again after its label was published.
sub-00142Completed5.7 h sealed
sub-00107Completed6.0 h sealed
sub-00074Completed6.4 h sealedclaimed again after its label was published.
sub-00128Completed4.6 h sealed
sub-00108Completed5.8 h sealed
sub-00009Completed5.6 h sealedclaimed again after its label was published.
sub-00004Completed3.8 h sealedclaimed again after its label was published.
sub-00095Completed5.9 h sealedclaimed again after its label was published.
sub-00034Completed6.4 h sealedclaimed again after its label was published.
sub-00103Completed5.2 h sealedclaimed again after its label was published.
sub-00145Completed5.3 h sealedclaimed again after its label was published.
sub-00062Completed5.2 h sealedclaimed again after its label was published.
sub-00100Completed5.5 h sealedclaimed again after its label was published.
Provenance
Which ledgers this was read from, and at what commit. The cohort spans two machines that each keep their own ledger; the reconciled reading merges them and is regenerated rather than edited.
Commit
3c4604fad6b1565980e9dee59833e097d423b98e
Engine / dataset
meld_graph on ds004199 (test)
Ledgers read
case_states.json, case_states.worker-host.json, case_states.worker-vm.json
Generated
27 Aug 2026, 03:43:39

Loading benchmark evidence…