1. Time to BCR in years
A predicted time-to-biochemical recurrence (BCR) risk score for the case, learned under weak supervision from patient-level outcomes only.
Interpretable AI for hypothesis-free discovery in prostate cancer histopathology. OWL benchmarks how faithfully a model's explanation holds up when it predicts biochemical recurrence — not only how well it predicts.
Nearly 30% of patients experience recurrence after radical prostatectomy, and deep learning can already predict that risk directly from digitised slides, often beyond what current clinical grading captures. What those models will not tell you is which morphology drives the prediction, or whether they lean on a clinically meaningful structure or on a spurious correlation.
OWL builds upon earlier challenges from the same group. PANDA showed that models can reproduce a pathologist's task, Gleason grading, at expert level, but only within the ceiling of existing human-defined patterns. LEOPARD dropped the annotations and predicted recurrence from outcome data alone, with strong predictive accuracy across international cohorts, but the winning models stayed black boxes. OWL keeps the discovery setting and adds the missing question: is the explanation any good?
For every case, a submitted algorithm must produce four things — a prediction, where it looked, and why.
A predicted time-to-biochemical recurrence (BCR) risk score for the case, learned under weak supervision from patient-level outcomes only.
A predicted biochemical recurrence (BCR) risk score ("low", "medium", "high") for the case, learned under weak supervision from patient-level outcomes only. This score should be derived from the participant's algorithm's own time-to-recurrence predictions.
A spatial attribution map highlighting the regions relevant to the prediction, as an 8-bit unsigned integer TIFF.
A free-form PDF of at most 5 pages with a human-readable interpretation of the decision for that case: concept attributions, textual explanation, and so on.
The difficulty is real: whole-slide images are gigapixel-scale and heterogeneous, the predictive patterns can be sparse and spread across scales, and in genuine biomarker discovery there is by definition no interpretability ground truth to compare against.
H&E-stained prostatectomy whole-slide images with biochemical recurrence status and time-to-event or censoring information, from three medical centres on two continents.
| Split | Cases | Availability |
|---|---|---|
| Training | 508 | Public, via the AWS Open Data Initiative under CC BY-NC-SA |
| Test | 861 | Private, internal and external cohorts, hidden from participants |
All patient data are fully anonymised, and collection, sharing and research use were approved by the institutional review boards and ethics committees of the participating institutions. Validation and test data and labels stay hidden, and the private cohorts have not been used to develop publicly available foundation models — which keeps contamination and memorisation out of the benchmark.
No single metric can decide whether an explanation is faithful, clinically meaningful and useful for discovery, so OWL combines automatic evaluation with a structured expert reader study.
Reader studies run through the Grand Challenge reader-study functionality; all algorithms are executed in a controlled environment for reproducibility. Confidence intervals come from bootstrapping over patient cases, and significance from a two-sided permutation test.
Create a verified account, then join the challenge — joining confirms that you accept the challenge rules.
508 training cases are available through the AWS Open Data Initiative.
Models are developed on your own hardware. Wrap the trained algorithm and its weights into a Docker container: the platform runs it in an offline AWS environment, so nothing can be downloaded at inference time.
Debug (3 training cases, several submissions per day, logs visible) to check that cloud and local predictions match; Test (1 submission per team).
Top-ranking teams present their methods during the one-hour OWL session in the morning program, and take part in a joint benchmark publication with up to 5 co-authors per submission.
The full and authoritative rules live on the challenge website; this is the short version.
Separate from the CVPath paper deadlines, which are listed on the CVPath page.
| Phase | What happens |
|---|---|
| 27.08.2026 - 20.11.2026 Launch | Starter kit, baseline, data loaders, Docker templates and evaluation scripts released; registration opens on Grand Challenge |
| 20.09.2026 - 20.11.2026 Debug | Sanity check of submissions. |
| 20.10.2026 - 20.11.2026 Test | One submission per team on the hidden test cohort; models above C-index 0.60 qualify for the interpretability stage |
| Expert reader study | AI and pathology panels assess the qualifying submissions |
| Results | Results announcement, ranking and presentations by the top-ranking teams in the OWL session, on the CVPATH workshop day |
The phase dates for this edition are published on the challenge website, together with any change to the schedule.
Bringing together computational pathology, explainable AI and clinical uropathology. The panel of clinical evaluators is larger than the list below and is finalised before the manual evaluation phase.
Radboud University Medical Center, The Netherlands
Challenge administration
Radboud University Medical Center, The Netherlands
Coordination, data
Radboud University Medical Center, The Netherlands
Baseline
Radboud University Medical Center, The Netherlands
beta testing
Radboud University Medical Center, The Netherlands
Baseline
University of Tübingen, Germany
AI evaluation
Amazon AGI, Germany
AI evaluation
Medical University of Vienna, Austria
AI evaluation
University Hospital Cologne, Germany
Data, clinical evaluation
Instituto de Medicina Patológica, Brazil
Data, clinical evaluation