Cohort check codes¶
What medh5 dataset check reports. These describe a cohort — a set of files
considered together — not a file.
That is why they are a separate code space from the diagnostic codes, and deliberately so: every file in a cohort can be individually valid while the cohort is unusable, because two sites used different label sets, or a subject appears in both train and test. No per-file validator can see either problem, so no per-file code exists for them.
The codes are tooling-level rather than normative: the specification does not
require an implementation to emit them, and a minor version may add to them
freely. The table is generated from
medh5/dataset/check.py
at build time.
| Code | Meaning |
|---|---|
C101 |
the cohort uses more than one label set |
C102 |
a class id means different things in different label sets |
C103 |
a sample declares no label set |
C201 |
a split claim's manifest digest is not this manifest's |
C202 |
one subject's or grouping key's samples claim different partitions of one split |
C203 |
a sample carries no split claim |
C204 |
a group holds part of a subject, so the split is not subject-safe |
C301 |
a class is examined in only part of the cohort |
C302 |
a class appears in no sample |
C401 |
a file changed after the manifest was written |
C402 |
a file in the manifest no longer exists |
C501 |
the cohort mixes de-identified and non-de-identified samples |
Related¶
- Build and split a cohort — the task these checks close.
- Diagnostic codes — the per-file
E/Wspace. medh5 dataset check— flags and JSON output.