datalad-fuse

Open files of DataLad datasets and git-annex repositories without downloading them first.

When you clone a DataLad dataset or a git-annex repository, you get all file names and the full history, but the content of the large (“annexed”) files stays where it was published: for example in the S3 storage of DANDI or OpenNeuro, or on a web server. datalad get downloads such files in full before you can open them.

datalad-fuse lets you open them right away instead. It looks up the URLs that git-annex has recorded for a file and, using fsspec, fetches only the parts of the file that are actually read.

Is it a good fit?

datalad-fuse pays off when you need a fraction of large files:

  • some arrays or tables out of NWB/HDF5 files, without the raw data;

  • the headers of many imaging files;

  • small metadata or sidecar files spread over a large dataset;

  • a quick look into a dataset that would not fit on your disk.

When you need whole files, for example the voxel data of compressed NIfTI (.nii.gz) images, which must be decompressed from the start, datalad get is usually faster, and keeps the files for later.

There are two ways to use it:

  • A FUSE mount (datalad fusefs, Linux with FUSE): the dataset appears as a read-only directory tree, so any program can open its files by name: h5ls, MATLAB, FSL, nwbinspector, shell tools, …

  • From Python (DatasetAdapter, FsspecAdapter; no FUSE needed): you get Python file objects, which libraries such as h5py, PyNWB or pandas can read from.

Both need git-annex (see Installation).

A quick look

Get a dataset, here a dandiset from the DANDI Archive, which is small to clone (about 2 MB) although its files add up to 1.86 GB:

$ datalad clone https://github.com/dandisets/000582

(OpenNeuro datasets are cloned the same way, e.g. datalad clone https://github.com/OpenNeuroDatasets/ds000001.)

Mount it. datalad fusefs keeps running until the dataset is unmounted:

$ mkdir mnt
$ datalad fusefs -d 000582 --foreground mnt

Then, in another terminal, use any tool on its files, e.g. h5ls from the HDF5 tools, and unmount when done:

$ h5ls mnt/sub-10073/sub-10073_ses-17010302_behavior+ecephys.nwb
acquisition              Group
analysis                 Group
...
units                    Group
$ fusermount -u mnt

Or open files directly from Python:

from contextlib import closing

import h5py
import pynwb

from datalad_fuse.fsspec import DatasetAdapter

# caching=False: keep fetched data in memory only
with closing(DatasetAdapter("000582", caching=False)) as dsa:
    with dsa.open("sub-10073/sub-10073_ses-17010302_behavior+ecephys.nwb") as f:
        with h5py.File(f, "r") as h5, pynwb.NWBHDF5IO(file=h5) as io:
            nwbfile = io.read()
            print(nwbfile.units.to_dataframe())

The Tutorial: streaming NWB data from DANDI walks through both in more detail.