15. Appendices¶
Appendix A: JSON Schema Definitions¶
The authoritative schema for Zarr Vectors root, level, and array metadata is
the LinkML model at
zarr_vectors-py/schema/zarr_vectors.linkml.yaml.
Tools may derive JSON Schema, Pydantic models, SQLAlchemy classes,
or OWL ontologies from that file using the standard gen-* LinkML
generators.
The schema covers:
Root metadata fields (
RootMetadata).Per-level metadata (
LevelMetadata).Per-array
.zattrsshapes perzv_arraydiscriminator.Enumerations:
LinksConvention,ObjectIndexConvention,CrossChunkStrategy,CrossLevelStorage,GeometryType,Encoding,CoarseningMethod.
Appendix B: Zarr Implementation Details¶
Zarr version: v3 only.
Per-array layout: each per-spatial-chunk array is a single vlen-bytes array over the level’s chunk grid, one cell per chunk, cell files at
c/<i>/<j>/<k>; the non-spatial arrays are single ragged or dense arrays at their logical path. Internal record framing (fragment-index, manifest-block, link-record streams) lives in project byte layouts inside each cell’s payload.Cell addressing: chunk at absolute coord
c→ cellc - chunk_grid_origin; occupied cells are listed in the array’snonempty_chunksattribute.Sharding: optional, via Zarr v3’s native
sharding_indexedcodec. There is no shard format specific to Zarr Vectors.Group metadata: root, level, and per-array
zarr.jsoncarry Zarr Vectors top-level keys (zarr_vectors,zarr_vectors_level,zv_array) alongside the Zarr v3 standardnode_type/zarr_formatfields.Multi-store hosting: stores work on any Zarr v3 store backend (LocalStore, FsspecStore over S3 / GCS / Azure Blob, in-memory, icechunk transactional stores, …).
Appendix C: Coordinate Reference Systems¶
Zarr Vectors reuses OME-Zarr RFC 4 axes (
name,type,unit) and RFC 5coordinateTransformations(scale,translation) for per-level coordinate transforms.The optional
crsdict on root metadata is opaque to the format — it carries an EPSG code, WKT string, or transform pipeline as defined by the writer’s CRS conventions.Per-axis units use UDUNITS-2 names (
"micrometer","second", …). The format does not stamp placeholder unit strings — unknown units must be omitted.
Appendix D: Compression Codec Reference¶
Default per-array codec pipelines (from
zarr_vectors.encoding.compression):
Array |
Dtype |
Compressor |
Shuffle |
|---|---|---|---|
|
user-declared (float or integer; see §7.1) |
Blosc(Zstd, clevel=5) |
BYTE-SHUFFLE |
|
user-declared |
Blosc(Zstd, clevel=5) |
BYTE-SHUFFLE |
|
user-declared |
Blosc(Zstd, clevel=5) |
BYTE-SHUFFLE |
|
opaque |
none — opaque bytes (see §11.4) |
— |
|
opaque |
none — opaque bytes (see §11.4) |
— |
|
user-declared integer (width chosen to fit |
Blosc(Zstd, clevel=5) |
BITSHUFFLE |
|
user-declared |
Blosc(Zstd, clevel=5) |
BYTE-SHUFFLE |
|
vlen-bytes (opaque manifest blob — §7.6) |
Blosc(Zstd, clevel=5) |
BYTE-SHUFFLE |
|
user-declared |
Blosc(Zstd, clevel=5) |
BYTE-SHUFFLE |
|
vlen-bytes (per-group |
Blosc(Zstd, clevel=5) |
BYTE-SHUFFLE |
|
user-declared |
Blosc(Zstd, clevel=5) |
BYTE-SHUFFLE |
Any of these may additionally be wrapped in sharding_indexed by
supplying a shard_shape; sharding composes with the codec, it does
not replace it. Note that this table is the recommended pipeline —
the default is no compressor at all (see
§11.3).
Mesh stores may use Draco-encoded vertex+face co-encoding instead of
the raw Blosc pipeline; the per-array .zattrs.encoding = "draco"
field flags this. Draco output is already compressed, so the Zarr
codec pipeline is left empty (or wrapped in a minimal pass-through
codec).
The fragment-index byte layout (§7.3) and manifest-block stream
(§7.6) are not Zarr codecs — they are project-internal record
framings carried as the raw bytes inside single-chunk uint8
arrays.
Appendix E: Downsampling Algorithms¶
The reference implementation provides one coarsening strategy:
Per-object coarsening (
coarsening_method = "per_object"): each level’s vertices are produced by reducing each surviving object’s fragments to one metavertex per coarse-bin, with attribute reduction following the per-attribute convention (mean / mode / sum, declared in the array metadata). Object identity is preserved (preserves_object_ids = true); dropped objects leave empty manifest slots.
Future strategies (mesh edge-collapse decimation, skeleton path
simplification, streamline point reduction) plug into the same
coarsening_method slot; see zarr_vectors.multiresolution.
Appendix F: Query Patterns¶
Common access patterns and which arrays they touch:
Pattern |
Arrays read |
|---|---|
Bounding-box at one level |
|
Per-bin sub-chunk filter |
|
Single object reconstruction |
|
All objects in a group |
|
Per-object attribute query |
|
Enumerate occupied chunks |
the array’s |
Pyramid drill-down (fine → coarse) |
|
Links leaving chunk |
|
All links incident on chunk |
one cell read per offsets segment under |
Appendix G: Performance Considerations¶
Chunk size: larger chunks amortise per-chunk overhead at the cost of larger minimum-read units. Coarser pyramid levels may grow
chunk_shapeindependently of level 0.Bin grid: enabling a per-chunk bin grid (
base_bin_shape) lets point-cloud queries narrow to individual bins without decoding the whole chunk — important when `vertex_count_per_chunk100k`.
Fragment vs object granularity: fragments are the unit of pyramid coarsening and re-use. When objects are very small (one fragment apiece, like nuclei), fragments and objects correspond 1-to-1; when objects are very large (a neuron with 10⁵ vertices), fragments naturally chunk the object’s vertices.
Per-link attributes:
link_attributes/<name>/<delta>/<chunk>read cost scales with the number of link rows in the cell oflinks/<delta>/<offsets>— the same fragment index that partitions vertices partitions these attributes.
Appendix H: Extensibility¶
Custom geometry types: writers may add a new value to
geometry_typesoutside the canonical set (point_cloud,line,polyline,streamline,skeleton,graph,mesh), at the cost of dropping conformance-level checks for that geometry.Custom metadata: arbitrary keys may be added under the
zarr_vectorsorzarr_vectors_levelnamespaces; readers should ignore unknown keys.Capability tokens: stores advertise optional features via
format_capabilities. Readers that don’t recognize a token must either treat the corresponding feature as absent or refuse to open the store.Version evolution: a hard-break version bump is the project’s only versioning mechanism; no shim layer ships with the implementation. Appendix J records what each bump changed.
Appendix I: References¶
TRX format specification: https://tee-ar-ex.github.io/trx-python/stable/trx_specifications.html
Zarr v3 specification: https://zarr-specs.readthedocs.io/
OME-Zarr (NGFF) specification: https://ngff.openmicroscopy.org/
LinkML: https://linkml.io/
Neuroglancer precomputed mesh format: https://github.com/google/neuroglancer/blob/master/src/datasource/precomputed/meshes.md
Neuroglancer precomputed annotation format: https://github.com/google/neuroglancer/blob/master/src/datasource/precomputed/annotations.md (see Appendix K for the mapping to Zarr Vectors).
Appendix J: Change Log¶
This appendix is the only place in the specification that describes prior versions of the format. Sections 1-14 and the other appendices describe the current version and nothing else.
Zarr Vectors is versioned per-feature, not per-release: every entry below describes a breaking on-disk change made under a single version bump.
Migration¶
There is no in-place migration utility between any two Zarr Vectors versions. Every bump changed the on-disk record layout in a way that breaks readers built for the previous version; stores must be rewritten from source.
The 0.7 → 0.8 step was briefly an exception — it shipped a one-shot helper that repartitioned the monolithic cross-chunk-link blob into the sharded per-tuple layout. That helper is obsolete, and 0.8 → 0.9 has no replacement, because two independent breaks compose across it:
Every per-spatial-chunk array changed physical form — from a group of single-chunk sub-arrays, one per spatial chunk, to a single vlen-bytes array over the chunk grid (§5.2). Nothing in a 0.8 store sits at a path a 0.9 reader looks at.
Cross-chunk records were keyed by the canonical-sorted tuple of endpoint chunks, and are now keyed by the source chunk with the relationship carried in the path. Re-deriving a source chunk from a sorted tuple is only possible where the record’s
perm_idxsurvived, and the two families’ cell sets do not correspond.
A reader MUST reject a store it cannot identify rather than read it on a best-effort basis (§13.2).
Version-at-a-glance¶
Version |
Headline change |
New capability tokens |
Hard break? |
Migration |
|---|---|---|---|---|
0.9 |
Single vlen array per per-chunk array (cells over the chunk grid); |
— (retires |
Yes |
Rewrite |
0.8.1 |
Flat single-array layout for the dense / ragged arrays; |
— |
Yes |
Rewrite |
0.8 |
Partitioned |
|
Yes |
In-place helper (obsoleted by 0.9) |
0.7 |
Per-level |
— |
Yes |
Rewrite |
0.6 |
Fragment-index byte layout for |
|
Yes |
Rewrite |
0.5 |
NGFF axis alignment; |
— |
Yes |
Rewrite |
0.4.1 |
Bare-integer level group names ( |
— |
Yes (path-only) |
Rewrite |
0.4 |
|
|
Yes |
Rewrite |
Cumulative capability set as of v0.9¶
A v0.9 store may carry any subset of:
multiscale_links(anydelta ≠ 0link array present; first available in 0.4. Since 0.9 it says nothing about chunk-spanning records, only level-spanning ones)fragment_index(mandatory in 0.6+)shared_fragments(per-chunk fragments referenced by multiple object manifests; renamed fromshared_vertex_groupsin 0.6)preserved_object_ids(per-object pyramid retained the parent’s OID space; available since 0.4)
Tokens are open-set; readers must tolerate unknown values.
Per-version detail¶
0.9.0 — two independent on-disk breaks, shipped together.
One array per logical array. Every per-spatial-chunk array —
vertices,vertex_fragments,link_fragments,links/<delta>/<offsets>,vertex_attributes/<name>,fragment_attributes/<name>,link_attributes/<name>/<delta>/<offsets>— becomes one Zarr v3 vlen-bytes array whose shape is the level’s chunk grid, with one cell per spatial chunk and its file at<array>/c/<i>/<j>/<k>. This replaces the layout in which every spatial chunk was its own single-chunkuint8sub-array under a per-array group. A chunk at absolute coordclands in cellc - origin, whereorigin = floor(min_corner / chunk_shape)is stored as the array’schunk_grid_originattribute (absent ⇒ zero origin), which lets data with negative coordinates map onto a 0-indexed array. Non-empty cells are tracked innonempty_chunksfor O(1) enumeration. An optionalshard_shapewraps the cells in Zarr v3’ssharding_indexedcodec. See §5.2.One connectivity family.
cross_chunk_linksandcross_chunk_link_attributesare merged intolinksandlink_attributes. Connectivity is one family: an intra-chunk link is just a link whose relative chunk offset is zero.links/<delta>becomes a group whose children are one rank-D vlen array per relative-offset segment (links/<delta>/<offsets>, cell = the record’s source chunk,vi_klocal tosrc + o_k), andlink_attributes/<name>/<delta>/<offsets>mirrors it exactly — same offsets, same cells, same row order. Endpoint tuples are no longer part of any path, so both families are ordinary chunk-grid arrays and shard like everything else. Thepartitioned_cross_chunk_linkscapability token is retired andmultiscale_linksnarrows to mean only “delta ≠ 0arrays are present”. See §10.6.Migration: rewrite from source (see Migration above).
0.8.1 — flat single-array layout for the dense and ragged arrays. Removes the
group-with-a-data-child pattern used by every non-spatial array (object_attributes,group_attributes,groups, the parametric blobs) and writes each logical array as a single standard Zarr v3 array at its logical path. Sparse coverage is expressed through the array’s ownfill_value— NaN for floats, a sentinel for integers — instead of a siblingpresent_maskchild array, so absence is in-band and cannot desynchronize from the data. Ragged arrays move to the vlen-bytes codec, matching whatobject_index/manifestsalready used. Custom per-array attributes move from the parent group’sattributesblock into the array’s own. Migration: rewrite.0.8.0 — sharded cross-chunk-link layout (§10.6). The single monolithic
cross_chunk_links/<delta>/dataint64 blob (and its parallelcross_chunk_link_attributes/<name>/<delta>/data) is replaced by K-separated sharded vlen-bytes zarr arrays — onekKsub-array per distinct K (1 ≤ K ≤ link_width) undercross_chunk_links/<delta>/. EachkKarray is N-D with shape(Cx,…) * K(one cell per sorted-unique chunk tuple), inner chunks of(1,)*(sid_ndim*K), outer shards of(4,)*(sid_ndim*K)by default, and the zarr v3sharding_indexed+vlen_bytescodecs. Each record’s encoding drops fromlink_width * (sid_ndim + 1) * 8bytes (chunk coords baked into the payload) to9 * link_widthbytes (chunk identity recovered from the cell coord’s K segments via a per-endpointuint8chunk-index in the record). Canonicalization rule:delta=0, link_width=2records MUST emitci = [0, 1](one orientation per undirected edge). Newlayout = "sharded_v1"discriminator stamped on everycross_chunk_links/<delta>/parent group.zattrs; newpartitioned_cross_chunk_linkscapability token, coupled withmultiscale_links(any store with across_chunk_links/<delta>/group MUST carry both).num_linksis no longer at the group level — per-cell counts are derivable from cell byte length (len(bytes) / (9 * link_width)). Migration: in-place helper that regroups records by their sorted unique chunks, rewrites the leaves, and bumpszv_version. That helper is obsolete — see Migration above.0.7.0 — per-level
chunk_shapeoverride onLevelMetadata(§6.6).RootMetadata.chunk_shaperemains the level-0 default; pyramid levels may carry a positive-integer-multiple chunk-shape override so coarser levels can grow chunks the way OME-Zarr image pyramids do via voxel-size scaling. Invariant: per-levelchunk_shapemust be a positive integer multiple of rootchunk_shapealong every axis (nested chunk grids). Cross-pyramid-level link arrays carried both endpoints’ chunk coords inline (still true under 0.8 via the kN array cell path); the per-axis multiplier is exposed bychunk_scale_factor(root, level). Migration: rewrite (no helper).0.6.0 — fragment-index byte layout (§7.3).
vertex_group_offsetsis replaced byvertex_fragments— a v1 byte layout with header, range bitmap, range table, and explicit- CSR row indices, supporting vertex re-use across fragments. Inline self-describing link blobs atdelta = 0are split into a flatlinks/0/<chunk>payload plus a siblinglink_fragments/<chunk>group with the same v1 fragment-index layout.object_index/dataadopts the manifest-block encoding (modes 0 / 1 / 2 — single / range / explicit) with chunk-local fragment references. Theshared_vertex_groupscapability token is renamed toshared_fragments(the sharing primitive is now a fragment, not a contiguous-byte vertex group); the newfragment_indexcapability is mandatory for 0.6+ stores. Migration: rewrite.0.5.0 — NGFF alignment and format simplification. Several on-disk simplifications shipped under the 0.5 series without separate version bumps (consumers should pin to a specific point release):
renamed
format_version→zv_versionto disambiguate from Zarr v3’szarr_formatfield;moved axes to NGFF
multiscales[0].axes(no longer duplicated underspatial_index_dims);dropped per-array dtype duplication (read from
zarr.jsondata_typedirectly);removed
vertex_counts/,metanode_children/,cross_chunk_faces/,attributes/<name>/<key>_offsets, andobject_index/pending/;replaced the
(K, 2)paired vertex-offset layout with a flat(K,)int64 vertex-offsets array (itself superseded by the fragment index in 0.6);cross-chunk face identity moves to
cross_chunk_links/<delta>/withlink_width = 3. Migration: rewrite.
0.4.1 — bare-integer resolution-level group names (
0/,1/, …,N/). Renamed fromresolution_0/,resolution_1/, … to mirror OME-Zarr’s level naming. Path-level break only; on-disk array layouts identical to 0.4. Migration: rename level groups in place.0.4 — multiscale-link arrays. Introduced the
<delta>sub-folder layout forlinks/,cross_chunk_links/,link_attributes/, andcross_chunk_link_attributes/and thecross_level_depth/cross_level_storagewriter knobs. This is the version where cross-pyramid-level links became expressible. Newmultiscale_linkscapability token marks stores with anydelta ≠ 0array. Migration: rewrite.
Appendix K: Mapping from Neuroglancer Precomputed Annotations¶
The Neuroglancer precomputed annotation format stores small geometric primitives — points, lines, axis-aligned bounding boxes, and ellipsoids — with per-annotation properties, per-segment relationships, and a multi-resolution random-subsample spatial index. It serves a similar purpose to Zarr Vectors but with narrower geometry semantics and a different multi-resolution model. This appendix maps the two layouts so authors can pick the right target and converters can translate between them.
K.1 Conceptual mapping¶
Neuroglancer Precomputed Annotations |
Zarr Vectors equivalent |
|---|---|
|
|
|
|
|
NGFF |
|
|
|
|
|
|
|
Pyramid levels |
|
Per-level effective |
|
Implicit via vertex count / bin_shape; pyramid coarsening controlled by |
Sharded vs unsharded |
Zarr v3 chunk-key encoding handles both transparently; sharding is a backend concern, not a schema choice |
Random subsampling between levels |
|
K.2 Mapping the four geometry primitives¶
Precomputed annotations are zero-dimensional primitives, each carrying
a fixed positional record. Zarr Vectors expresses them via the
geometry_types list plus the appropriate per-record layout:
Precomputed |
Positional record |
Zarr Vectors representation |
|---|---|---|
|
1 vector |
|
|
2 vectors (endpoint A, B) |
|
|
2 vectors (min, max) |
Two options. (a) |
|
2 vectors (center, radii) |
|
|
uint32 count + N vectors |
|
Multi-geometry stores are natural in Zarr Vectors (just list every
geometry in geometry_types); the precomputed format restricts a
single store to one annotation_type and would require co-locating
several stores to mix kinds.
K.3 Properties and relationships¶
Per-annotation properties in precomputed correspond directly to per-vertex (or per-object) attributes:
Numeric properties (
uint8,int8, …,float32) → typedvertex_attributes/<name>/<chunk>orobject_attributes/<name>. Zarr Vectors uses the array’sdtypedirectly; no per-property enum table is needed at the schema level.rgb/rgba→ either an(N, 3)/(N, 4)uint8 attribute, or three / four channels of a multi-channel attribute with declaredchannel_names.enum_values/enum_labels→ carried aschannel_names/ per-attribute side metadata. Zarr Vectors does not currently reserve a top-level enum-mapping slot; writers stamp the mapping as a JSON dict in the attribute’s.zattrs.
Relationships (each annotation linked to a list of segment IDs) map onto Zarr Vectors groups:
The precomputed
relationships[<rel_name>]becomes a per-relationshipgroupsarray (or, if multiple relationships, onegroups-like structure per relationship name — typically expressed today by rechunking by relationship; see §8).Each group corresponds to one segment id. Its membership list is the annotation IDs (= object IDs in Zarr Vectors).
The inverse — annotation → list of segments — is then the group-membership matrix (which annotations belong to which groups); in Zarr Vectors this is recovered by iterating
groups, the same operation that powers “show all annotations on segment X” in precomputed.
K.4 The spatial index¶
Both formats use a multi-resolution spatial grid to keep query cost bounded, but the selection rule differs:
Precomputed: each level has a per-cell
limit. A cell with more annotations than the limit gets a random subsample at this level; the rest propagate to finer children. Cell coverage is controlled bygrid_shape×chunk_size, anchored atlower_bound. Levels coarsen by integer division of cells.Zarr Vectors: each level has a fixed
bin_shape(and optionally an overriddenchunk_shape). Coarsening is per-object: each surviving object’s vertices are aggregated into metavertices at the coarser bin grid.object_sparsity ∈ (0, 1]in level metadata records what fraction of objects survived.
Equivalents:
Concept |
Precomputed |
Zarr Vectors |
|---|---|---|
Grid cells per axis at level L |
|
|
Cell size at level L |
|
|
LOD selection knob |
|
|
Cross-level identity |
annotation ID is shared across levels |
OID-preserving pyramid ( |
Drillable parent→child mapping |
implicit (random subsample) |
optional |
Precomputed’s “drill the visible cells until you’ve returned at most
limit annotations per cell” maps to Zarr Vectors’ “ask each level
for the object set; stop when object_sparsity * source_count fits
your budget.” The precomputed model is simpler and gets random
sampling for free; the Zarr Vectors model is more general (handles
extended objects, not just points) at the cost of an explicit
coarsener.
K.5 Practical conversion notes¶
Single annotation type → single Zarr Vectors store is a straightforward 1:1 transcode. Pick the geometry mapping from §K.2; emit one fragment per annotation; populate
object_index/in dense OID order from the originalby_idlisting.Properties transcode field-for-field; preserve the original uint8 / int8 / … types rather than upcasting.
Relationships transcode to groups keyed by segment id. In practice, very large segment-id spaces (10⁸+ segments) suggest using
object_attributes/segment_idrather thangroupsif most segments touch zero or one annotation.Spatial index does NOT transcode 1:1. The pyramid is rebuilt using the per-object coarsener; the precomputed
limitrule has no direct Zarr Vectors equivalent. A reasonable default isreduction_factor = 8,object_sparsity ≈ limit / max_cell_countper level, andcross_level_storage = "implicit"if downstream readers need to drill from coarse to fine.Sharded vs unsharded is invisible to the schema mapping — Zarr Vectors writes per-chunk blobs into a Zarr v3 store regardless of the backend’s sharding.
K.6 What Zarr Vectors adds over precomputed annotations¶
Connected geometries (polylines with branches, skeletons, meshes) live in the same store, not just zero-dimensional primitives.
link_width = 1, 2, 3, ...and the fragment-index partitioning carry the connectivity.Fragment / vertex re-use: explicit fragment indices let multiple objects reference the same per-chunk vertex rows (
shared_fragments), saving storage on overlapping geometries (streamline bundles, mesh rims).Per-link attributes: edges / faces carry their own attribute arrays (
link_attributes/<name>/<delta>/<chunk>), not just the per-annotation properties precomputed supports.Cross-pyramid-level links (§9.6) make the pyramid drillable — given a coarse metavertex, walk back to the fine-level vertices that produced it. Precomputed’s random-subsample model loses this mapping at construction time.
Pluggable coarsening:
coarsening_method = "per_object"is the default, but mesh decimation and streamline point-reduction slot into the same level structure.
K.7 What precomputed annotations preserve that Zarr Vectors does not (yet)¶
Sharded back-end as a first-class schema concern. Zarr Vectors delegates sharding to the underlying Zarr v3 store; it does not expose
shardingblocks per index.Per-property enum tables as a typed schema field (
enum_values/enum_labels). Zarr Vectors carries this as free-form.zattrsmetadata.The “limit” LOD heuristic — automatic, occupancy-driven downsampling without writing a coarsening pass. In Zarr Vectors the writer must pick
bin_ratio/chunk_scale_factor/object_sparsitydeliberately.