Research2025 · First-author journal article · Intelligent Oncology

A contour has to work for a person.

Published comparison of abdominal organ contours from nnU-Net, Auto3DSeg, SwinUNETR, and ground truth
Published Figure 3 · nnU-Net, Auto3DSeg, SwinUNETR, and ground-truth contours · Source

Good overlap scores are useful, but they do not tell the whole story of an automatically drawn organ. This head-to-head study pairs quantitative evaluation with blinded physician review.

Quantitative performance and expert review of deep-learning frameworks for abdominal-organ segmentation
Original public recordings

3 ways to watch the work.

Preview images load from this site. YouTube connects only after you press play.

Research · 14:46

Using Deep Learning models to auto-contour abdominal organs in CT scans

A talk on deep-learning contouring of abdominal organs, associated with the 2023 Physics Undergraduate Conference.

Uploaded March 24, 2023

Research · 12:24

Abdominal organs auto-segmentation to reduce patient-wait times in radiation oncology treatments

A CUPC 2022 presentation about automated abdominal-organ segmentation in radiation oncology.

Uploaded October 30, 2022
Event: October 29, 2022

Research · 4:18

CUMPC 2022 Presentation

A short Canadian Undergraduate Medical Physics Conference presentation on deep-learning abdominal multi-organ segmentation.

Uploaded August 24, 2022
Event: August 25, 2022

Comparing approaches fairly

The project compared nnU-Net, MONAI Auto3DSeg, and SwinUNETR on a shared abdominal computed-tomography (CT) segmentation task. The two automated machine-learning (AutoML) frameworks and the transformer-based model were evaluated on the same data.

Automating organ segmentation can reduce repetitive work, but a useful comparison needs to address both the geometry of an output and its acceptability to a clinical reader.

Beyond a single score

The study used 122 training images and 72 holdout images. Three physicians performed a blinded review of 30 cases, complementing the quantitative analysis.

That second layer matters because an aggregate metric can hide differences that become obvious when a physician inspects a contour in context.

What the comparison found

Both AutoML frameworks outperformed SwinUNETR in the reported evaluation. Average Dice scores were 0.924 for nnU-Net, 0.902 for Auto3DSeg, and 0.837 for SwinUNETR. Dice measures spatial overlap; the evaluation also used surface Dice and the 95th-percentile Hausdorff distance.

Physicians preferred nnU-Net over MONAI Auto3DSeg. That expert-review result adds a different kind of evidence to the geometric measurements.

These findings describe the evaluated dataset, implementations, and review protocol. They should not be read as a universal ranking for every anatomy, acquisition protocol, or clinical setting.

Publication

The first-author article was published in Intelligent Oncology in 2025, volume 1, issue 2, pages 160–171. A free full-text version is available through PubMed Central.

Sources & further reading

Research summary updated October 5, 2026. Findings describe the linked study and do not establish clinical readiness beyond its evaluation.