What has been tested, what the tests found, and where the model is known to be weak — assembled from the published validation artifacts.
Parameter recovery: simulate centers whose truth is known, run the pipeline, and see how well it recovers the ordering.
A 95% interval should contain the truth 95% of the time. Measured against held-out later SRTR releases.
Simulation-based calibration: draw parameters from the prior, simulate data, refit, and check the rank statistics are uniform. A failure here would mean priors, likelihood and sampler disagree.
Per-center Spearman correlation between predicted access and observed SRTR transplant rates.
Children are modeled from pediatric SRTR cohorts and validated against SRTR's own risk-adjusted pediatric tier — the one pediatric ground truth that is not an input to the pipeline.
A design question tested rather than assumed: does shrinking estimates across SRTR releases beat using the latest one?
Every published artifact, newest first. A stale date is itself a finding.
The score is a weighted sum of eight categories. Those weights are a judgement, not a measurement — so the question is whether a different reasonable judgement gives a different answer.
A weight only moves the ranking through a sub-score that varies between centers. A category scoring the same everywhere adds the same constant to every total and cannot reorder anything — however large its weight. So the weights the scoring panel displays are not the same thing as what the ordering is built from.
The tool asks for several things about you and returns a ranked list of centers. Those are two different questions: whether an answer changes your numbers, and whether it changes which center comes out on top. They do not have the same answer, and the difference is not currently obvious from the interface.
The full register lives in the repository; these are the ones that most change how the output should be read.