Command-line usage
datalad-fuse adds three commands to DataLad:
|
Mount a dataset with FUSE, so that any program can read its files. |
|
Print the first lines or bytes of a file. |
|
Remove the on-disk cache of fetched data. |
This page shows how to use them; Command line reference lists all of their options. Each command is also available from Python, see DataLad commands.
Mounting a dataset: datalad fusefs
$ mkdir mnt
$ datalad fusefs -d path/to/dataset --foreground mnt
mounts the dataset at path/to/dataset on the existing, empty directory
mnt (the mount point). Without -d, the dataset containing the
current directory is mounted. Create the mount point outside of the dataset:
a mount point inside it would show up in the mount itself, and as an
untracked directory in datalad status.
The --foreground (-f) option is currently required: the command keeps
running for as long as the dataset is mounted. Run it in a terminal of its
own, or in a tmux/screen session. When running it in the background
of a shell (appending &), wait until the mount is ready before using it,
e.g. with until mountpoint -q mnt; do sleep 0.1; done.
To unmount, press Ctrl-C in the terminal where datalad fusefs runs
(for a background job, bring it to the foreground with fg first), or run:
$ fusermount -u mnt
What you see in the mount
The same directory tree as in the dataset’s working tree.
Annexed files appear as regular files, whether or not their content is present. For a file whose content is not present, the size comes from its annex key and the modification time is the date of the dataset’s last commit; reading it fetches the needed parts from a remote URL (see How it works). A file whose content is present shows the size, permissions and modification time of that content.
The
.gitdirectory is hidden (see--mode-transparentbelow).Installed subdatasets are included; uninstalled ones are empty directories.
Files cannot be written, created or deleted (see Read-only access for a few exceptions).
Options
--caching ondiskKeep fetched data in an on-disk cache for reuse, also by later mounts (see Caching). The default,
none, only buffers data in memory while a file is open.--allow-otherLet other users access the mount; by default, only the user who mounted it can. This requires the line
user_allow_otherin/etc/fuse.conf.--mode-transparentShow the
.gitdirectories. Annexed files whose content is not present then appear as the symlinks they are in the dataset, and the targets of these symlinks under.git/annex/objects/can be read, with their content fetched as needed. Files under.gitcan also be written to.
Clearing the cache on exit
The configuration option datalad.fusefs.cache-clear makes datalad
fusefs remove on-disk caches when it exits:
visitedClear the caches of the (sub)datasets that were accessed in the mount. This only has an effect if the mount used
--caching ondisk.recursiveClear the caches of the mounted dataset and all its installed subdatasets.
Set it like any DataLad or git configuration option, permanently (e.g.
git config --global datalad.fusefs.cache-clear visited) or for a single
call:
$ datalad -c datalad.fusefs.cache-clear=visited fusefs -d ds --foreground --caching ondisk mnt
Peeking into a file: datalad fsspec-head
Prints the first lines (10 by default) or bytes of a file to standard output, fetching only what is needed. It is a quick way to check that the content of a file can be reached, or to look at the header of a file:
$ datalad fsspec-head -d path/to/dataset -n 5 data/participants.tsv
$ datalad fsspec-head -d path/to/dataset -c 8 data/recording.nwb | od -c
Note
Relative paths are interpreted relative to the top directory of the dataset, not to the current directory.
The output is the raw content of the file, without any result rendering, so
it can be piped into other tools. --caching ondisk stores the fetched
data in the dataset’s cache.
Clearing the cache: datalad fsspec-cache-clear
Removes the on-disk cache of a dataset (.git/datalad/cache/fsspec/):
$ datalad fsspec-cache-clear -d path/to/dataset
Add -r (--recursive) to also clear the caches of all installed
subdatasets.