Skip to content

Storage

The on-disk layout, how bytes get chosen, and how the file proves it is intact. The normative statement is spec §2, §13 and §14; this page is the operational half.

Layout

An ordinary HDF5 file. h5ls, HDFView, h5py, MATLAB and Julia all open it.

case_0001.medh5                      (root attrs: medh5_version, medh5_kind,
│                                     medh5_profiles, content_id, digest_algo,
│                                     created, generator)
├── meta                             the sample document: JSON, UTF-8, one dataset
├── grids/<grid_id>                  empty groups; geometry lives in attributes
├── images/<image_id>                arrays, chunked and compressed
├── annotations/<ann_id>/            per-kind datasets, plus a header in attributes
├── transforms/<transform_id>/
└── index/<ann_id>/                  derived: foreground coords, counts, bboxes
$ medh5 tree case_0001.medh5      # the same listing, annotated with spec roles

meta is a scalar variable-length UTF-8 string, always. HDF5 filters do not apply to variable-length data — it lives in the global heap — so a compressed meta is not expressible. A vocabulary large enough for that to matter uses a referenced label set (form: "ref", §5.1) instead.

Everything a reader needs to interpret an array is in meta plus the grid attributes. There is no sidecar and no dataset-wide schema to keep in sync.

Codec profiles

$ medh5 recompress cohort/*.medh5 --profile archive

Recompression is in place by default: each file is rewritten atomically under the new profile. --out writes beside the source instead, and takes a single input and a destination filename — not a directory, and not several files:

$ medh5 recompress case.medh5 --profile archive --out cold/case.medh5
$ for f in cohort/*.medh5; do
      medh5 recompress "$f" --profile archive --out "cold/$(basename "$f")"
  done
Profile Images Labels For
training blosc2 lz4:1 shuffle blosc2 lz4:1 shuffle fastest decompression; the hot dataloader path
balanced blosc2 zstd:3 shuffle blosc2 zstd:3 bitshuffle general use — the default
archive blosc2 zstd:9 bitshuffle blosc2 zstd:9 bitshuffle smallest on disk; cold storage and distribution
portable gzip:4 + HDF5 shuffle gzip:4 + HDF5 shuffle readable without hdf5plugin

Under balanced, labels get bitshuffle where images get byte shuffle: label planes are low-entropy integers, and bit-level transposition compresses them much better than byte-level does. training byte-shuffles both for speed, and archive bit-shuffles both for size.

portable exists because a collaborator with a plain h5py install should be able to open the file at all. Blosc2 needs the filter plugin; gzip is in every HDF5 build.

Datasets under 64 KiB raw are stored contiguous and uncompressed: chunking and a filter pipeline cost more than they save at that size, and a contiguous read is one seek. (W902 warns only from 1 MiB, so the policy never trips its own warning.)

from medh5.storage.codecs import PROFILES, resolve_profile, dataset_kwargs
PROFILES["archive"].description

Chunking

The chunk is the real unit of I/O: reading one voxel reads a whole chunk. Two forces pull against each other — sizing to the L3 cache keeps a patch inside cache after decompression, sizing to the training patch keeps read amplification low — and the optimiser resolves them by starting at the patch, growing toward the cache budget, and stopping before the chunk is much larger than the patch.

w.add_grid("ct", shape=..., spacing=..., patch_hint=(96, 96, 96))

patch_hint is how you tell it what you will read. Without one it assumes a 64-voxel cube, clipped to the grid, and you get a reasonable answer.

L3 size is detected per-core where the platform allows and falls back to ~1.375 MiB otherwise. Chunks are held between 512 KiB and 4 MiB.

Stacked encodings chunk per plane. A layers, bitmask or probmap annotation — and a displacement field's components — is chunked (1, *spatial_chunk), so reading one plane does not decompress the others (§14.1). Combined with reading a multi-class dense() by plane rather than by class, a 200-class annotation packed into four layers is four reads and not two hundred — which is where the 64³ patch time went from 117 ms to 4 ms.

$ medh5 recompress case.medh5 --profile training --rechunk

--rechunk re-derives the chunk shape as well as the codec, by the writer's rule: from each grid's patch_hint (or chunk_hint), and (1, *spatial_chunk) for the stacked encodings and a displacement field's components.

Integrity

Every object carries a SHA-256 over its decompressed content, with the object's sample-root-relative path, dtype and shape fed into the hash first, so two arrays with the same bytes and different shapes do not collide.

The root carries content_id: a Merkle digest over those object digests plus the metadata attributes that define the sample.

$ medh5 verify cohort/*.medh5
$ medh5 verify case.medh5 --partial images/CT_tp0
s.verify().ok
s.verify(partial=["images/CT_tp0"])
s.content_id
s.compute_content_id()          # recompute rather than read

Two properties follow from digesting content rather than encoding:

Recompression preserves content_id. Every stored byte changes; no digest does. A cohort re-encoded for training is still verifiably the same data.

The root is not a substitute for the objects. Because content_id hashes digests, editing a dataset without restamping it breaks that object's digest and leaves the root matching. verify checks both, which is why it reports per-object mismatches rather than a single yes/no.

$ medh5 fix case.medh5                       # diagnose
$ medh5 fix case.medh5 --rebuild-index
$ medh5 fix case.medh5 --rewrite-digests --reason "rebuilt by an external tool"

--rewrite-digests is not repair — see CLI.

Atomic writes

medh5.create and medh5.amend build a temporary file beside the target and os.replace it into position on a clean exit. A file appears complete or not at all; an exception aborts and leaves nothing behind.

amend is copy-on-write, and it copies through every object it does not understand — including ones written by a future minor version. Amending never silently drops what this reader cannot read.

The sampling index

$ medh5 index build cohort/*.medh5 --max-coords 4096

Per voxel annotation, per class: a bounded sample of foreground coordinates, the exact voxel count, and a tight bounding box. That makes foreground patch sampling O(1) in the volume instead of a scan, and gives dataset stats its class counts for a few hundred bytes instead of a decompression pass.

A reader loads the class table and counts once per open, and a draw reads the one coordinate it picked — O(1) in the class count too: 0.10 ms at 63 classes. The index's datasets are stored with the file's label codec, and the coarse occupancy map is chunked one class per chunk — the unit a reader asks about.

An index carries the digest of the annotation it derives from. When they disagree the index is stale, readers must ignore it, and the validator raises W905:

$ medh5 fix cohort/*.medh5 --rebuild-index

A stale index is not a file error. It is a cache that needs rebuilding, and the format says so rather than making the file invalid.

Collections

One .medh5c shard holding many samples, for filesystems that dislike a million small files:

$ medh5 pack cohort/*.medh5 -o shard.medh5c
$ medh5 ls shard.medh5c
$ medh5 unpack shard.medh5c -o restored/
import medh5

medh5.pack(paths, "shard.medh5c")
with medh5.open_collection("shard.medh5c") as c:
    c["case_0001"].images["CT"].read()      # an ordinary Sample

A .medh5c holds many sample roots under samples/<key>. Each member is a sample root, so every reader, validator and loader works on it unchanged.

Packing is a container operation: chunks move as raw bytes, nothing is decompressed, and content_id is preserved. Unpacking reproduces the original files chunk for chunk — there is a test that compares them at that level, because comparing through the value API would decompress and recompress and prove nothing.

sample ⊂ collection is strict containment: a collection is not a sample and does not pretend to be one. medh5 validate dispatches on the kind.

Reading it without medh5

The point of plain HDF5:

import h5py, json

with h5py.File("case_0001.medh5") as f:
    doc = json.loads(f["meta"][()])
    doc["identity"]["subject_id"]
    dict(f["grids"]["ct_tp0"].attrs)       # spacing, origin, direction
    f["images"]["CT_tp0"][10:20]           # needs hdf5plugin for blosc2

import hdf5plugin before opening if the file uses a blosc2 profile; --profile portable avoids that requirement entirely.