Python usage
From Python, files of a dataset can be opened without any FUSE mount, as file objects that many libraries accept in place of a file name. This page covers:
DatasetAdapter: open files of a single dataset;FsspecAdapter: open files across a dataset and its subdatasets;the
dataladcommands ofdatalad-fuse, called from Python;mounting a dataset with FUSE from Python.
The examples use the dandiset cloned in the Tutorial: streaming NWB data from DANDI, and are run from the directory containing it:
nwb_path = "sub-10073/sub-10073_ses-17010302_behavior+ecephys.nwb"
Opening files of a dataset
DatasetAdapter takes the path to a dataset (or any
git-annex repository), and opens files by their path relative to the
dataset’s top directory:
from contextlib import closing
from datalad_fuse.fsspec import DatasetAdapter
with closing(DatasetAdapter("000582", caching=False)) as dsa:
with dsa.open(nwb_path) as f: # binary mode, like open(..., "rb")
print(f.read(8)) # b'\x89HDF\r\n\x1a\n'
with dsa.open("dandiset.yaml", "rt") as f: # text mode
print(f.readline())
Important
Many libraries (h5py, PyNWB, zarr, …) read data lazily, only when you
access them. The file object must stay open until all data you need have
been read, so do the reading inside the with blocks. Reading after the
file was closed fails, with h5py with an error mentioning “identifier is
not of specified type”.
open() returns a seekable, read-only file object:
for files read from disk (not annexed, or with content present), a regular Python file object;
for files read from a URL, an fsspec file object, which fetches data as it is read.
Text mode ("r" or "rt") accepts encoding (default "utf-8") and
errors arguments, as the built-in open() does. Writing is not
supported.
The caching argument is required: True keeps fetched data in an
on-disk cache inside the dataset, False only buffers them in memory while
a file is open (see Caching). With caching=True, dsa.clear()
removes the dataset’s cache.
The adapter starts git annex processes to answer its queries;
close(), called by contextlib.closing() above, stops them.
Working interactively
When exploring data interactively, e.g. in Jupyter, with blocks are
impractical. Open everything step by step instead, and close it in reverse
order when you are done:
import h5py
import pynwb
dsa = DatasetAdapter("000582", caching=True)
f = dsa.open(nwb_path)
h5 = h5py.File(f, "r")
io = pynwb.NWBHDF5IO(file=h5)
nwbfile = io.read()
# ... explore nwbfile in further cells ...
io.close()
h5.close()
f.close()
dsa.close()
Passing file objects to other libraries
Anything that reads from a Python file object can read from these. For example:
import json
import h5py
import pandas as pd
with closing(DatasetAdapter("path/to/dataset", caching=True)) as dsa:
# HDF5, and formats based on it such as NWB
with dsa.open("data/recording.h5") as f, h5py.File(f, "r") as h5:
signal = h5["signal"][:1000]
# tabular text data
with dsa.open("participants.tsv", "rt") as f:
participants = pd.read_csv(f, sep="\t")
# JSON
with dsa.open("dataset_description.json", "rt") as f:
description = json.load(f)
Libraries that need a file name rather than a file object cannot use these objects; use a FUSE mount for them (see Mounting from Python).
Inspecting files
get_file_state tells whether a file is
annexed and whether its content is present, and returns its git-annex key as
an AnnexKey:
with closing(DatasetAdapter("000582", caching=False)) as dsa:
state, key = dsa.get_file_state(nwb_path)
print(state) # FileState.NO_CONTENT
print(key.size, key.backend) # 15657857 SHA256E
for url in dsa.get_urls(str(key)):
print(url)
which prints the URLs that will be tried, in order:
FileState.NO_CONTENT
15657857 SHA256E
https://api.dandiarchive.org/api/assets/2b9e441b-.../download/
https://dandiarchive.s3.amazonaws.com/blobs/26a/22c/26a22c31-...?versionId=...
The possible states are described in How it works.
AnnexKey can also parse and format keys on its own (see
Python API reference).
Datasets with subdatasets
FsspecAdapter works on a dataset together with its
installed subdatasets: for each path, it finds the (sub)dataset that contains
it and uses a DatasetAdapter for that dataset. It is a
context manager:
from pathlib import Path
from datalad_fuse.fsspec import FsspecAdapter
root = Path("path/to/superdataset").resolve()
with FsspecAdapter(root, caching=False) as fsa:
path = root / "subdataset" / "data" / "file.nwb"
print(fsa.get_file_state(path))
print(fsa.is_under_annex(path))
with fsa.open(path) as f:
header = f.read(1024)
Important
Use an absolute root, and absolute paths under it, as above. Paths
relative to root, or to the current directory, are not supported.
Besides open(), get_file_state() and is_under_annex(), it offers
get_commit_datetime() (the date of the last commit of the dataset
containing a path) and resolve_dataset() (the
DatasetAdapter and relative path used for a path).
DataLad commands
The commands of datalad-fuse are available as functions in
datalad.api and as methods of datalad.api.Dataset, with the same
options as on the command line (see Command-line usage):
from datalad.api import Dataset
ds = Dataset("000582")
res = ds.fsspec_head(nwb_path, bytes=8, result_renderer="disabled")
print(res[0]["data"]) # b'\x89HDF\r\n\x1a\n'
ds.fsspec_cache_clear(recursive=True)
Like all DataLad commands, they return result records (dictionaries);
fsspec_head puts the fetched bytes into the data field of its result.
Mounting from Python
datalad.api.fusefs (or Dataset.fusefs) mounts a dataset, and does not
return until it is unmounted. To work with a mount from a Python program,
run datalad fusefs as a separate process instead:
import os
import subprocess
import time
os.makedirs("mnt", exist_ok=True)
mount = subprocess.Popen(["datalad", "fusefs", "-d", "000582", "--foreground", "mnt"])
while mount.poll() is None and not os.path.ismount("mnt"):
time.sleep(0.1) # wait until the mount is ready
try:
# any code or tool can now open files under mnt/
with open(os.path.join("mnt", nwb_path), "rb") as f:
print(f.read(8))
finally:
subprocess.run(["fusermount", "-u", "mnt"], check=True)
mount.wait()
To pass other FUSE mount options, mount the file system class
DataLadFUSE directly with mfusepy; keyword arguments of
FUSE() become mount options. DataLadFUSE needs the absolute path of
the dataset, without symbolic links, as returned by os.path.realpath():
import os
from mfusepy import FUSE
from datalad_fuse.fuse_ import DataLadFUSE
FUSE(
DataLadFUSE(os.path.realpath("000582"), caching=True),
"mnt",
foreground=True,
ro=True,
)