Dataset format#
A GVL dataset is a directory written by gvl.write and read
by gvl.Dataset.open. This page is the authoritative
description of its on-disk layout.
Directory layout#
dataset_dir/
├── metadata.json # the Metadata schema (below)
├── input_regions.arrow # original BED regions + region-index map
├── genotypes/ # present iff variants were provided to gvl.write
│ ├── offsets.npy # per (region, sample, ploidy) offsets into variant_idxs.npy; absent when sourced from .svar2
│ ├── svar_meta.json # shape + dtype of offsets.npy — present iff source was .svar
│ ├── variant_idxs.npy # variant indices; absent when sourced from .svar or .svar2
│ ├── dosages.npy # optional, absent when sourced from .svar or .svar2
│ ├── variants.arrow # variant table; absent when sourced from .svar or .svar2
│ └── svar2_ranges/ # present iff source was .svar2 — see "svar2_ranges layout" below
└── intervals/ # or annot_intervals/ when annotated; present iff tracks given
When the dataset was built from an .svar, the heavy per-variant arrays (variant_idxs.npy,
dosages.npy, index.arrow) are not duplicated into the dataset. Instead the dataset
records a back-reference to the source .svar in metadata.json (see svar_link below).
Likewise, a dataset built from an .svar2 records a back-reference (svar2_link, below)
and caches the var-key window for each (region, sample, ploidy) that holds a variant,
under genotypes/svar2_ranges/ — the bulk variant data stays in the .svar2 store. See
“genotypes/svar2_ranges/ layout” below for the on-disk size of this cache; it scales with
variants observed inside regions rather than with regions x samples, so it is small at
cohort scale.
metadata.json schema#
metadata.json is the serialization of genvarloader._dataset._write.Metadata:
Field |
Type |
Notes |
|---|---|---|
|
|
Sample identifiers, sorted. |
|
|
Contig names used to interpret BED coords. |
|
|
Number of regions (after jitter padding). |
|
|
Ploidy when the dataset has genotypes. |
|
|
Maximum coordinate jitter (defaults to 0). |
|
|
Package version that wrote this dataset. Drives format dispatch. |
|
|
Back-reference to a source |
|
|
Back-reference to a source |
|
|
Bounded content fingerprint of |
SvarLink:
Field |
Type |
Notes |
|---|---|---|
|
|
POSIX path from |
|
|
Original absolute path; used as a fallback. |
|
|
Integrity check (see below). |
SvarFingerprint:
Field |
Type |
Notes |
|---|---|---|
|
|
Row count of the svar’s |
|
|
Byte size of the svar’s |
Svar2Link (mirrors SvarLink for a .svar2 source):
Field |
Type |
Notes |
|---|---|---|
|
|
POSIX path from |
|
|
Original absolute path; used as a fallback. |
|
|
Integrity check (see below). |
Svar2Fingerprint:
Field |
Type |
Notes |
|---|---|---|
|
|
Count of the |
|
|
Summed byte size of those data files. |
.svar2 has no variant_idxs.npy/index.arrow analogue exposed cheaply, so its fingerprint
keys on file count + total byte size of the store’s data files rather than a variant count.
genotypes/svar2_ranges/ layout#
Written only when the dataset’s variant source is a .svar2 store. R = number of regions,
S = number of the dataset’s selected samples (not necessarily the full .svar2 cohort),
P = ploidy. A region-CSR (compressed sparse row) table stores only the
(region, sample, ploid) windows that hold a variant:
File |
Shape |
Notes |
|---|---|---|
|
|
CSR row pointer: region |
|
|
|
|
|
A 24-byte record per non-empty cell: |
|
|
Per-region (sample-independent) range into the dense SNP store. |
|
|
Per-region (sample-independent) range into the dense indel store. |
|
|
Maps the dataset’s selected-sample slot to the |
|
— |
Records |
region_ptr.npy, cell_id.npy, and cell_vk.npy are raw, headerless tofile dumps despite
the .npy extension, like dense_snp_range.npy and dense_indel_range.npy. Only
sample_cols.npy is a real .npy file (written with np.save).
Each non-empty cell costs 28 bytes (24 for cell_vk + 4 for cell_id), so size scales with
the number of variants observed inside regions rather than with regions x samples: a dense
(R, S, P, 2) int64 cache — 32 bytes for every cell, including empty ones — would be 128 GB
for the All of Us chr22 grid (R = 3,734, S = 535,662, P = 2); the sparse layout is 504 MB
there, at a realized fill of 0.45%. Genome-wide, the same comparison is 6.93 TB dense against
27.3 GB sparse. gvl.write logs realized fill after the first contig and projects the final
on-disk size from it, warning when the filesystem reports too little free space.
At read time, Dataset.__getitem__ looks up each queried (region, sample, ploid) with a
bounded binary search within that region’s region_ptr block (no search at all when the block
is fully occupied, i.e. every sample/ploid combination in the region holds a variant) to build
the flat per-query inputs for the read-bound Rust kernels — no interval-search tree and no
dense-union rebuild happen per read, unlike the .svar path.
A producer must emit cell_id in strictly ascending order within each region’s block, since the
binary search assumes it. This is not verified on read — doing so would be an O(N) scan over the
whole table, 27 GB genome-wide — so a producer that violates it degrades to silent misses (the
search lands on a wrong-but-plausible position, and the confirmation step then reports “not
found”) rather than an error.
SVAR resolution at open time#
When opening a dataset whose metadata.svar_link is non-null,
Dataset.open resolves the svar in this order:
Caller-provided
svar=...argument.svar_link.relative_pathresolved against the dataset directory.svar_link.absolute_path.A unique
*.svardirectory next to the dataset.
If none match, a FileNotFoundError is raised naming the expected .svar basename. After
resolution, the fingerprint is verified; a mismatch raises ValueError and lists both
expected and observed values.
.svar2 resolution at open time#
When opening a dataset whose metadata.svar2_link is non-null,
Dataset.open resolves the .svar2 store in the same order
as .svar:
Caller-provided
svar2=...argument.svar2_link.relative_pathresolved against the dataset directory.svar2_link.absolute_path.A unique
*.svar2directory next to the dataset.
If none match, a FileNotFoundError is raised naming the expected .svar2 basename and
suggesting svar2=. After resolution, the fingerprint (Svar2Fingerprint, above) is verified;
a mismatch raises ValueError and lists both expected and observed values.
.svar2 variants ALT convention#
For a pure deletion (e.g. VCF GTA>G), decoding with_seqs("variants") yields different raw
ALT bytes depending on the backing store: .svar reports the VCF anchor base (b"G"), while
.svar2 reports the atomized empty ALT (b"") — a genoray .svar2 format convention, not a
bug. Both stores consume the ALT identically when reconstructing haplotype sequence, so
with_seqs("haplotypes") / with_seqs("annotated") output is byte-identical between the two
backends; only RaggedVariants.alt differs, and only for pure-deletion records. The same holds
for with_seqs("variant-windows"): ref_window is byte-identical between the backends, while the
alt/alt_window fields differ only for pure-deletion records (the same empty-vs-anchor ALT).
.svar2 Phase-1 unsupported combinations#
A .svar2-backed dataset supports all four output modes (haplotypes, variants,
variant-windows, and haplotype-realigned tracks), unphased_union, and
var_fields-selected store INFO/FORMAT fields (on both "variants" and "variant-windows").
Haplotype and variants output also support splicing, var_filter="exonic", and
automatic reverse-complementation for negative-strand regions. Spliced
RaggedVariants are returned as complete (transcript, sample, phase, ~variants)
cells.
The following combinations are Phase-1 scope and raise NotImplementedError (or, for
extend_to_length, at write time) instead of silently mis-computing:
Splicing with
variant-windowsor haplotype-realigned tracks.var_filter="exonic"withvariant-windowsor haplotype-realigned tracks.min_af/max_affiltering (.svaronly; see “Should I use.svaror.svar2” in the FAQ).annotatedhaplotypes (with_seqs("annotated")).VarWindowOpt(ref="allele")(bare-allele REF mode; REF alleles aren’t stored in.svar2).Reverse-complement with haplotype-realigned tracks.
Fixed-length (integer
output_length) haplotype-realigned track output.variants/variant-windowsoutput on a dataset written withmax_jitter>0or read withjitter>0(the read-bound decode does not right-clip to the post-jitter window).gvl.write(..., extend_to_length=False)for a.svar2variant source.
(Multi-contig FlankSample track fills are now supported and byte-identical to the .svar
backend — issue #267.)
See the genvarloader skill’s .svar2 section for the full narrative and var_fields semantics.
Format changelog#
Version |
Change |
|---|---|
|
Variant coordinates stored 0-based. |
|
Variant coordinates switched to 1-based. |
|
|
|
|
|
|
|
|
Upgrading legacy datasets. A dataset written before
0.25.0that was built from an.svarwill still open (with aDeprecationWarning). Rungenvarloader.migrate_svar_link(path)to convert the symlink layout to the new metadata layout in place.Shrinking a
.svar2dataset’s range cache.gvl.migratedoes not convert the0.42.xdense range-cache layout to0.43.0’s sparse one — it only handles the 1.x -> 2.0 track array-of-structs to struct-of-arrays migration. To shrink an existing.svar2-backed dataset, re-rungvl.writeagainst the samebed/variants/samples. GVL<= 0.42.1cannot open a0.43.0-or-later.svar2dataset: it fails with a bareKeyError: 'vk_snp_range'— loud, not silently wrong, but unhelpful; upgrade GVL to open it instead.