datalad-fuse
Open files of DataLad datasets and git-annex repositories without downloading them first.
When you clone a DataLad dataset or a
git-annex repository, you get all file
names and the full history, but the content of the large (“annexed”) files
stays where it was published: for example in the S3 storage of DANDI or OpenNeuro, or on a
web server. datalad get downloads such files in full before you can open
them.
datalad-fuse lets you open them right away instead. It looks up the URLs
that git-annex has recorded for a file and, using
fsspec, fetches only the parts of
the file that are actually read.
Is it a good fit?
datalad-fuse pays off when you need a fraction of large files:
some arrays or tables out of NWB/HDF5 files, without the raw data;
the headers of many imaging files;
small metadata or sidecar files spread over a large dataset;
a quick look into a dataset that would not fit on your disk.
When you need whole files, for example the voxel data of compressed NIfTI
(.nii.gz) images, which must be decompressed from the start, datalad
get is usually faster, and keeps the files for later.
There are two ways to use it:
A FUSE mount (
datalad fusefs, Linux with FUSE): the dataset appears as a read-only directory tree, so any program can open its files by name:h5ls, MATLAB, FSL,nwbinspector, shell tools, …From Python (
DatasetAdapter,FsspecAdapter; no FUSE needed): you get Python file objects, which libraries such as h5py, PyNWB or pandas can read from.
Both need git-annex (see Installation).
A quick look
Get a dataset, here a dandiset from the DANDI Archive, which is small to clone (about 2 MB) although its files add up to 1.86 GB:
$ datalad clone https://github.com/dandisets/000582
(OpenNeuro datasets are cloned the same way, e.g. datalad clone
https://github.com/OpenNeuroDatasets/ds000001.)
Mount it. datalad fusefs keeps running until the dataset is unmounted:
$ mkdir mnt
$ datalad fusefs -d 000582 --foreground mnt
Then, in another terminal, use any tool on its files, e.g. h5ls from the
HDF5 tools, and unmount when done:
$ h5ls mnt/sub-10073/sub-10073_ses-17010302_behavior+ecephys.nwb
acquisition Group
analysis Group
...
units Group
$ fusermount -u mnt
Or open files directly from Python:
from contextlib import closing
import h5py
import pynwb
from datalad_fuse.fsspec import DatasetAdapter
# caching=False: keep fetched data in memory only
with closing(DatasetAdapter("000582", caching=False)) as dsa:
with dsa.open("sub-10073/sub-10073_ses-17010302_behavior+ecephys.nwb") as f:
with h5py.File(f, "r") as h5, pynwb.NWBHDF5IO(file=h5) as io:
nwbfile = io.read()
print(nwbfile.units.to_dataframe())
The Tutorial: streaming NWB data from DANDI walks through both in more detail.
Getting started
Development