Skip to content

Cohort check codes

What medh5 dataset check reports. These describe a cohort — a set of files considered together — not a file.

That is why they are a separate code space from the diagnostic codes, and deliberately so: every file in a cohort can be individually valid while the cohort is unusable, because two sites used different label sets, or a subject appears in both train and test. No per-file validator can see either problem, so no per-file code exists for them.

The codes are tooling-level rather than normative: the specification does not require an implementation to emit them, and a minor version may add to them freely. The table is generated from medh5/dataset/check.py at build time.

medh5 dataset check cohort.json --deep
Code Meaning
C101 the cohort uses more than one label set
C102 a class id means different things in different label sets
C103 a sample declares no label set
C201 a split claim's manifest digest is not this manifest's
C202 one subject's or grouping key's samples claim different partitions of one split
C203 a sample carries no split claim
C204 a group holds part of a subject, so the split is not subject-safe
C301 a class is examined in only part of the cohort
C302 a class appears in no sample
C401 a file changed after the manifest was written
C402 a file in the manifest no longer exists
C501 the cohort mixes de-identified and non-de-identified samples