The models are ours, and they are scored.
We trained the segmentation models ourselves, on a public dataset that permits it. On cases the models never saw, the four chambers and the myocardium reach a median Dice of 0.956 or better across 119 scans. Everything below says how that number was produced, and what it does not mean.
The models
Two models do the segmentation. A localiser finds the heart and the aorta in the whole volume. A chamber model then separates the myocardium and the four chambers inside it. Both are nnU-Net v2 networks that we trained. The architecture is open and widely used; the trained weights are our own work, and they are what the measurements depend on.
We did not license a third-party chamber model, and we do not run one. That matters for what we can tell you about it: we hold the training log, the case lists and the held-out scores, so every figure on this page can be traced rather than taken on trust.
What they were trained on
The chamber model was trained on 484 cases of the TotalSegmentator v1 dataset, across all acquisition phases rather than contrast alone. That dataset is published under Creative Commons Attribution 4.0, which permits derived work provided the source is credited. It is credited here, and this is the only cohort in our work whose derived output we may release.
Wasserthal, J. et al. TotalSegmentator: Robust Segmentation of 104 Anatomic Structures in CT Images. Radiology: Artificial Intelligence (2023). Dataset licensed CC BY 4.0.
Why the holdout is genuinely held out
A score only means something if the model never saw the cases it was scored on. We checked that from the training artefacts rather than assuming it. The candidate list held 604 cases. Training used 387, validation another 97. The remaining 120 were never seen, and their overlap with the validation set is zero.
The arithmetic closes exactly. That is the point: it is checkable, and you can ask us for the case lists.
What it scores
Dice measures how far two outlines agree, where 1.000 is exact. Volume error is the difference between the measured volume and the reference, as a percentage of the reference.
| Structure | Dice, median | 10–90 | Volume error % |
|---|---|---|---|
| Myocardium | 0.956 | 0.904–0.970 | +0.1 |
| Left atrium | 0.979 | 0.960–0.986 | −0.1 |
| Left ventricle | 0.975 | 0.951–0.983 | −0.1 |
| Right atrium | 0.971 | 0.931–0.981 | −0.1 |
| Right ventricle | 0.973 | 0.944–0.981 | −0.2 |
119 HELD-OUT CASES · ONE REFUSED BY THE BORDER GATE, NOT SCORED
The localiser is scored separately, on 122 held-out cases: a median Dice of 0.970 for the heart and 0.979 for the aorta.
What this does not establish
Three limits, stated because a careful reader would find them anyway.
- It is not a comparison. No other chamber model has been scored on these cases. The claim is that ours is measured, not that it is better.
- It is resampled data. The public dataset ships at 1.5 mm, so performance on native-resolution clinical scans is not established by this run.
- It is a single reconstructed phase. These are not end-diastolic volumes and must not be read against end-diastolic reference ranges.
Every value carries its record
Each delivery ships the model version, the parameters it ran with, the hash of the volume it read, and the de-identification record. A measurement you cannot trace back to the scan and the method that produced it is not evidence, and we do not ship one.
Questions
Ask for the case lists, the training log or the scoring artefacts: anamaria.chioran@corpxanalytics.com