Python usage

From Python, files of a dataset can be opened without any FUSE mount, as file objects that many libraries accept in place of a file name. This page covers:

  • DatasetAdapter: open files of a single dataset;

  • FsspecAdapter: open files across a dataset and its subdatasets;

  • the datalad commands of datalad-fuse, called from Python;

  • mounting a dataset with FUSE from Python.

The examples use the dandiset cloned in the Tutorial: streaming NWB data from DANDI, and are run from the directory containing it:

nwb_path = "sub-10073/sub-10073_ses-17010302_behavior+ecephys.nwb"

Opening files of a dataset

DatasetAdapter takes the path to a dataset (or any git-annex repository), and opens files by their path relative to the dataset’s top directory:

from contextlib import closing

from datalad_fuse.fsspec import DatasetAdapter

with closing(DatasetAdapter("000582", caching=False)) as dsa:
    with dsa.open(nwb_path) as f:  # binary mode, like open(..., "rb")
        print(f.read(8))  # b'\x89HDF\r\n\x1a\n'
    with dsa.open("dandiset.yaml", "rt") as f:  # text mode
        print(f.readline())

Important

Many libraries (h5py, PyNWB, zarr, …) read data lazily, only when you access them. The file object must stay open until all data you need have been read, so do the reading inside the with blocks. Reading after the file was closed fails, with h5py with an error mentioning “identifier is not of specified type”.

open() returns a seekable, read-only file object:

  • for files read from disk (not annexed, or with content present), a regular Python file object;

  • for files read from a URL, an fsspec file object, which fetches data as it is read.

Text mode ("r" or "rt") accepts encoding (default "utf-8") and errors arguments, as the built-in open() does. Writing is not supported.

The caching argument is required: True keeps fetched data in an on-disk cache inside the dataset, False only buffers them in memory while a file is open (see Caching). With caching=True, dsa.clear() removes the dataset’s cache.

The adapter starts git annex processes to answer its queries; close(), called by contextlib.closing() above, stops them.

Working interactively

When exploring data interactively, e.g. in Jupyter, with blocks are impractical. Open everything step by step instead, and close it in reverse order when you are done:

import h5py
import pynwb

dsa = DatasetAdapter("000582", caching=True)
f = dsa.open(nwb_path)
h5 = h5py.File(f, "r")
io = pynwb.NWBHDF5IO(file=h5)
nwbfile = io.read()

# ... explore nwbfile in further cells ...

io.close()
h5.close()
f.close()
dsa.close()

Passing file objects to other libraries

Anything that reads from a Python file object can read from these. For example:

import json

import h5py
import pandas as pd

with closing(DatasetAdapter("path/to/dataset", caching=True)) as dsa:
    # HDF5, and formats based on it such as NWB
    with dsa.open("data/recording.h5") as f, h5py.File(f, "r") as h5:
        signal = h5["signal"][:1000]
    # tabular text data
    with dsa.open("participants.tsv", "rt") as f:
        participants = pd.read_csv(f, sep="\t")
    # JSON
    with dsa.open("dataset_description.json", "rt") as f:
        description = json.load(f)

Libraries that need a file name rather than a file object cannot use these objects; use a FUSE mount for them (see Mounting from Python).

Inspecting files

get_file_state tells whether a file is annexed and whether its content is present, and returns its git-annex key as an AnnexKey:

with closing(DatasetAdapter("000582", caching=False)) as dsa:
    state, key = dsa.get_file_state(nwb_path)
    print(state)  # FileState.NO_CONTENT
    print(key.size, key.backend)  # 15657857 SHA256E
    for url in dsa.get_urls(str(key)):
        print(url)

which prints the URLs that will be tried, in order:

FileState.NO_CONTENT
15657857 SHA256E
https://api.dandiarchive.org/api/assets/2b9e441b-.../download/
https://dandiarchive.s3.amazonaws.com/blobs/26a/22c/26a22c31-...?versionId=...

The possible states are described in How it works. AnnexKey can also parse and format keys on its own (see Python API reference).

Datasets with subdatasets

FsspecAdapter works on a dataset together with its installed subdatasets: for each path, it finds the (sub)dataset that contains it and uses a DatasetAdapter for that dataset. It is a context manager:

from pathlib import Path

from datalad_fuse.fsspec import FsspecAdapter

root = Path("path/to/superdataset").resolve()
with FsspecAdapter(root, caching=False) as fsa:
    path = root / "subdataset" / "data" / "file.nwb"
    print(fsa.get_file_state(path))
    print(fsa.is_under_annex(path))
    with fsa.open(path) as f:
        header = f.read(1024)

Important

Use an absolute root, and absolute paths under it, as above. Paths relative to root, or to the current directory, are not supported.

Besides open(), get_file_state() and is_under_annex(), it offers get_commit_datetime() (the date of the last commit of the dataset containing a path) and resolve_dataset() (the DatasetAdapter and relative path used for a path).

DataLad commands

The commands of datalad-fuse are available as functions in datalad.api and as methods of datalad.api.Dataset, with the same options as on the command line (see Command-line usage):

from datalad.api import Dataset

ds = Dataset("000582")
res = ds.fsspec_head(nwb_path, bytes=8, result_renderer="disabled")
print(res[0]["data"])  # b'\x89HDF\r\n\x1a\n'
ds.fsspec_cache_clear(recursive=True)

Like all DataLad commands, they return result records (dictionaries); fsspec_head puts the fetched bytes into the data field of its result.

Mounting from Python

datalad.api.fusefs (or Dataset.fusefs) mounts a dataset, and does not return until it is unmounted. To work with a mount from a Python program, run datalad fusefs as a separate process instead:

import os
import subprocess
import time

os.makedirs("mnt", exist_ok=True)
mount = subprocess.Popen(["datalad", "fusefs", "-d", "000582", "--foreground", "mnt"])
while mount.poll() is None and not os.path.ismount("mnt"):
    time.sleep(0.1)  # wait until the mount is ready
try:
    # any code or tool can now open files under mnt/
    with open(os.path.join("mnt", nwb_path), "rb") as f:
        print(f.read(8))
finally:
    subprocess.run(["fusermount", "-u", "mnt"], check=True)
    mount.wait()

To pass other FUSE mount options, mount the file system class DataLadFUSE directly with mfusepy; keyword arguments of FUSE() become mount options. DataLadFUSE needs the absolute path of the dataset, without symbolic links, as returned by os.path.realpath():

import os

from mfusepy import FUSE

from datalad_fuse.fuse_ import DataLadFUSE

FUSE(
    DataLadFUSE(os.path.realpath("000582"), caching=True),
    "mnt",
    foreground=True,
    ro=True,
)