Storage¶
The on-disk layout, how bytes get chosen, and how the file proves it is intact. The normative statement is spec §2, §13 and §14; this page is the operational half.
Layout¶
An ordinary HDF5 file. h5ls, HDFView, h5py, MATLAB and Julia all open it.
case_0001.medh5 (root attrs: medh5_version, medh5_kind,
│ medh5_profiles, content_id, digest_algo,
│ created, generator)
├── meta the sample document: JSON, UTF-8, one dataset
├── grids/<grid_id> empty groups; geometry lives in attributes
├── images/<image_id> arrays, chunked and compressed
├── annotations/<ann_id>/ per-kind datasets, plus a header in attributes
├── transforms/<transform_id>/
└── index/<ann_id>/ derived: foreground coords, counts, bboxes
meta is a scalar variable-length UTF-8 string, always. HDF5 filters do not
apply to variable-length data — it lives in the global heap — so a compressed
meta is not expressible. A vocabulary large enough for that to matter uses a
referenced label set (form: "ref", §5.1) instead.
Everything a reader needs to interpret an array is in meta plus the grid
attributes. There is no sidecar and no dataset-wide schema to keep in sync.
Codec profiles¶
Recompression is in place by default: each file is rewritten atomically
under the new profile. --out writes beside the source instead, and takes a
single input and a destination filename — not a directory, and not several
files:
$ medh5 recompress case.medh5 --profile archive --out cold/case.medh5
$ for f in cohort/*.medh5; do
medh5 recompress "$f" --profile archive --out "cold/$(basename "$f")"
done
| Profile | Images | Labels | For |
|---|---|---|---|
training |
blosc2 lz4:1 shuffle | blosc2 lz4:1 shuffle | fastest decompression; the hot dataloader path |
balanced |
blosc2 zstd:3 shuffle | blosc2 zstd:3 bitshuffle | general use — the default |
archive |
blosc2 zstd:9 bitshuffle | blosc2 zstd:9 bitshuffle | smallest on disk; cold storage and distribution |
portable |
gzip:4 + HDF5 shuffle | gzip:4 + HDF5 shuffle | readable without hdf5plugin |
Under balanced, labels get bitshuffle where images get byte shuffle: label
planes are low-entropy integers, and bit-level transposition compresses them
much better than byte-level does. training byte-shuffles both for speed, and
archive bit-shuffles both for size.
portable exists because a collaborator with a plain h5py install should be
able to open the file at all. Blosc2 needs the filter plugin; gzip is in every
HDF5 build.
Datasets under 64 KiB raw are stored contiguous and uncompressed: chunking
and a filter pipeline cost more than they save at that size, and a contiguous
read is one seek. (W902 warns only from 1 MiB, so the policy never trips its
own warning.)
from medh5.storage.codecs import PROFILES, resolve_profile, dataset_kwargs
PROFILES["archive"].description
Chunking¶
The chunk is the real unit of I/O: reading one voxel reads a whole chunk. Two forces pull against each other — sizing to the L3 cache keeps a patch inside cache after decompression, sizing to the training patch keeps read amplification low — and the optimiser resolves them by starting at the patch, growing toward the cache budget, and stopping before the chunk is much larger than the patch.
patch_hint is how you tell it what you will read. Without one it assumes a
64-voxel cube, clipped to the grid, and you get a reasonable answer.
L3 size is detected per-core where the platform allows and falls back to ~1.375 MiB otherwise. Chunks are held between 512 KiB and 4 MiB.
Stacked encodings chunk per plane. A layers, bitmask or probmap
annotation — and a displacement field's components — is chunked
(1, *spatial_chunk), so reading one plane does not decompress the others
(§14.1). Combined with reading a multi-class dense() by plane rather than by
class, a 200-class annotation packed into four layers is four reads and not
two hundred — which is where the 64³ patch time went from 117 ms to 4 ms.
--rechunk re-derives the chunk shape as well as the codec, by the writer's
rule: from each grid's patch_hint (or chunk_hint), and (1, *spatial_chunk)
for the stacked encodings and a displacement field's components.
Integrity¶
Every object carries a SHA-256 over its decompressed content, with the object's sample-root-relative path, dtype and shape fed into the hash first, so two arrays with the same bytes and different shapes do not collide.
The root carries content_id: a Merkle digest over those object digests plus
the metadata attributes that define the sample.
s.verify().ok
s.verify(partial=["images/CT_tp0"])
s.content_id
s.compute_content_id() # recompute rather than read
Two properties follow from digesting content rather than encoding:
Recompression preserves content_id. Every stored byte changes; no digest
does. A cohort re-encoded for training is still verifiably the same data.
The root is not a substitute for the objects. Because content_id hashes
digests, editing a dataset without restamping it breaks that object's digest
and leaves the root matching. verify checks both, which is why it reports
per-object mismatches rather than a single yes/no.
$ medh5 fix case.medh5 # diagnose
$ medh5 fix case.medh5 --rebuild-index
$ medh5 fix case.medh5 --rewrite-digests --reason "rebuilt by an external tool"
--rewrite-digests is not repair — see CLI.
Atomic writes¶
medh5.create and medh5.amend build a temporary file beside the target and
os.replace it into position on a clean exit. A file appears complete or not
at all; an exception aborts and leaves nothing behind.
amend is copy-on-write, and it copies through every object it does not
understand — including ones written by a future minor version. Amending never
silently drops what this reader cannot read.
The sampling index¶
Per voxel annotation, per class: a bounded sample of foreground coordinates,
the exact voxel count, and a tight bounding box. That makes foreground patch
sampling O(1) in the volume instead of a scan, and gives dataset stats its
class counts for a few hundred bytes instead of a decompression pass.
A reader loads the class table and counts once per open, and a draw reads the
one coordinate it picked — O(1) in the class count too: 0.10 ms at 63 classes.
The index's datasets are stored with the file's label codec, and the coarse
occupancy map is chunked one class per chunk — the unit a reader asks about.
An index carries the digest of the annotation it derives from. When they
disagree the index is stale, readers must ignore it, and the validator
raises W905:
A stale index is not a file error. It is a cache that needs rebuilding, and the format says so rather than making the file invalid.
Collections¶
One .medh5c shard holding many samples, for filesystems that dislike a million
small files:
$ medh5 pack cohort/*.medh5 -o shard.medh5c
$ medh5 ls shard.medh5c
$ medh5 unpack shard.medh5c -o restored/
import medh5
medh5.pack(paths, "shard.medh5c")
with medh5.open_collection("shard.medh5c") as c:
c["case_0001"].images["CT"].read() # an ordinary Sample
A .medh5c holds many sample roots under samples/<key>. Each member is a
sample root, so every reader, validator and loader works on it unchanged.
Packing is a container operation: chunks move as raw bytes, nothing is
decompressed, and content_id is preserved. Unpacking reproduces the original
files chunk for chunk — there is a test that compares them at that level,
because comparing through the value API would decompress and recompress and
prove nothing.
sample ⊂ collection is strict containment: a collection is not a sample and
does not pretend to be one. medh5 validate dispatches on the kind.
Reading it without medh5¶
The point of plain HDF5:
import h5py, json
with h5py.File("case_0001.medh5") as f:
doc = json.loads(f["meta"][()])
doc["identity"]["subject_id"]
dict(f["grids"]["ct_tp0"].attrs) # spacing, origin, direction
f["images"]["CT_tp0"][10:20] # needs hdf5plugin for blosc2
import hdf5plugin before opening if the file uses a blosc2 profile;
--profile portable avoids that requirement entirely.
Related¶
- Tune performance — chunking, codecs and the index as levers.
- Check a file before training on it —
verifyandmedh5 fix. - Specification §2, §13, §14 — the normative statement.