The ‘rule’ is the key foundational concept in Snakemake, and the bulk of implementing a workflow in Snakemake involves specifying a set of rules. Each rule describes the production of a set of outputs, where an output is defined as a change to a filesystem (typically the creation of a new file). The rule forms a recipe that specifies the inputs (the files on which the production of the output depends), the mechanism (the software that executes to produce the output given the information contained in the rule), the parameters (the configuration settings that affect the operation of the mechanism), and the resources (the computational requirements for executing the rule). Each rule is given a user-defined identifier and is specified using a domain-specific language that resembles YAML and incorporates Python.
For our example project, we will need to write a rule for each of the three processing steps (motion correction, temporal averaging, and coregistration). The rules are contained within separate files, each with a .smk extension, within the rules directory described previously. The contents of these rules are shown in Figs. 3 and 4; below, we will work through the key aspects of the construction of these rules.
Fig. 3
The rule (A) and script (B) for the motion-correction step in the example workflow
Fig. 4
The rules for the temporal averaging (A) and coregistration (B) steps in the example workflow
OutputsThe outputs of a rule are listed as items within the output directive. Each output is specified as a string describing a path on a filesystem that is affected by the execution of the rule, and can be given a custom name that can be used to reference the path. The paths are typically specified as relative locations beginning with results/, which are then converted to absolute paths by Snakemake (using the current working directory as the base for the absolute paths, by default) when executed. These strings can either be specified literally or via the return value of a Python function.
The dependence of paths on specific details (e.g., the inclusion of a specific subject number in the name of a directory or file) is alleviated in Snakemake by the use of wildcards. A wildcard is an identifier within braces that is contained within the path string (e.g., ""), which acts as a placeholder in the path specification and allows rules to generalise across different values of the wildcard. Conceptually, this can be thought of as the rule indicating that its execution produces output paths that follow a particular pattern; alternatively, and more in the spirit of Snakemake, it can be thought of as a given path being able to be produced by the rule if the wildcards are set to particular values.
For our example workflow, the output paths for each of the rules contain wildcards for the subject identifier (sub_num). Note that the rule for motion correction (Fig. 3A) has an additional wildcard for the functional imaging task type (task). The inclusion of this additional wildcard is because the motion correction operates independently across the two functional imaging runs (corresponding to the two task types) per participant, whereas the other two rules are subsequent to aggregation over functional imaging runs and thus only vary in participant.
Neuroimaging processing steps can sometimes produce large numbers of output files (for example, Freesurfer’s recon-all command). Although listing each output path is recommended and is best practice, if the output files reside within a common parent directory and no other rules produce outputs within that directory, an alternative is to specify that directory as the rule’s output. When using this approach, the directory path is wrapped within the directory Snakemake helper function.
A processing step can also produce intermediate or extraneous files that are unnecessary for the workflow and do not need to be saved in the filesystem after the job has completed. Such files can be manually removed as part of the rule’s processing. However, a simpler approach is to use the shadow rule directive, which causes the rule to be executed in an isolated temporary location and only the listed output files copied to their destination after successful completion of the rule. The shadow directive is used in the temporal averaging and coregistration rules in the example workflow (Fig. 4) to discard ephemeral intermediate files. Using shadow is also useful for commands that create temporary files with fixed names during their processing, particularly when such files are created in the working directory; without shadow, parallel execution of rules can cause such temporary files to be read from and written to by multiple jobs—potentially with chaotic consequences.
InputsInputs are specified within the input directive and have similar characteristics to the outputs described above. A particularly useful approach to specifying inputs within a workflow is as a direct reference to the output of an upstream (parent) rule; indeed, it is a useful practice to aim to only ever specify the direct path for a given file at one location within the entire workflow. In our example workflow, the rules describe their inputs in terms of the output of a parent rule where possible.
Inputs can also be specified as Python functions that are evaluated when Snakemake constructs its representation of the workflow. Such input functions have a wildcards parameter, which has as its argument a Python object with wildcard variables as attributes. The function is required to return a single string, a list of strings, or a dictionary (with identifiers as keys and paths as values). For simple functions, anonymous (lambda) functions are often appropriate (see the rule shown in Fig. 4A).
Input functions are particularly useful for rules that need to aggregate over runs or participants. For example, the temporal averaging step in our example workflow requires inputs to comprise both functional imaging runs for a given participant. We thus only have one output wildcard: the participant identifier (sub_num). We can use an input function to return a two-item list of strings that describes the required input paths, given a particular value of sub_num (Fig. 4A). The Snakemake helper function collect is useful for inserting such values into strings containing wildcards.
MechanismA Snakemake rule can be considered as a contract that a certain set of output files will be produced from a certain set of input files. The mechanism of the rule describes how this computational process is implemented. If the process is simple, it can be described using a shell directive in which the value contains the command string to execute. For example, the rule for temporal averaging in our example workflow is executed via the sequential execution of two AFNI commands (Fig. 4A). When specifying the shell command, attributes of the rule such as input and output can be referenced within the command to be replaced with the relevant value by Snakemake upon execution.
For any process that is more complicated or difficult to express in a shell command, it is preferable to delegate the description of the mechanism to an external script. Such scripts are typically saved within the scripts workflow sub-directory and can be written in programming languages such as Python, R, Julia, or Rust—or even Jupyter notebooks. The script is specified using a script directive in which the value is a path to the script file, relative to the location of the rule definition. During execution, the script has access to a variable that exposes the information contained within the calling rule. Although sufficiently simple that the shell directive would have been appropriate, we illustrate the use of the script directive in our example workflow in the motion correction step (Fig. 3).
The commands that are executed as part of a rule can often produce textual output that is useful for debugging and quality control. Snakemake provides the log directive to allow a path to be specified for capturing such output streams. This path can be used as the target for redirecting standard output and/or standard error either in the shell directive, as in the temporal averaging and coregistration rules, or within the processing script, as in the motion correction rule.
It is also common to require the use of external scripts for languages and applications that are not directly supported by Snakemake (e.g., a Matlab script). Here, a useful approach is to save the script inside the resources workflow sub-directory and register this path as an input item for the rule. Then, the contents of the script can be passed to the Matlab executable’s standard input within a shell command or an external script in a language like Python. By using environment variables, rule parameters can be interpreted from within the Matlab script (using Matlab’s getenv function).
An optional, but highly recommended, Snakemake capability is to execute jobs in containers. A container combines software and an operating system into a single unit (called an image) that can be executed by a container platform, with Docker being perhaps the most well-known (see Wiebels and Moreau, 2021, for a tutorial on containers in a research context). Snakemake uses Apptainer as its container platform, which is particularly suited to academic and high-performance computing. The advantage of using containers is the increase in reproducibility that is provided by the lack of a need to install the required software (often with a specific version requirement) and the ability to run the software in an isolated and ephemeral environment. Containers are specified using the container directive in the rule, with the value being any URL that is supported by Apptainer as the source of a container image.
Container images for common neuroimaging software can be obtained from the Neurodesk project (Renton et al. 2024). The available software can be viewed on the Neurodesk Github repository,Footnote 7 which then links to images stored in the Docker package registry on Github. Such URLs can then be provided as the value for the container directive; for example, each rule in our example workflow specifies a URL from the Neurodesk project for AFNI. For software that is not explicitly packaged in a container by Neurodesk, custom containers can be created by using the visual interface in the Neurodesk interactive container builderFootnote 8 or by creating Apptainer definitions (see Nüst et al., 2020) and building a container manually.Footnote 9 Our accompanying website gives further details on the use of containers and provides an example of custom container creation.Footnote 10
ParametersThe commands that are executed within rules often have additional parameters that modify the operation of the command. Such parameters can sometimes be hard-coded within the rule’s mechanism (via the shell or script directive described above); however, to provide additional visibility and flexibility of such parameters, they can also be specifically named in the rule by using the params directive. For example, the motion correction step in the example workflow requires the ‘base’ image (the motion correction target) to be specified. By including the base item within the params directive, the rule’s mechanism then has access to params.base when building the command (Fig. 3B).
Parameters that apply across multiple rules or where it is desirable to be easily user-modified can also be specified within a global configuration. This configuration can be read from a file (in YAML or JSON format) at a location specified by the configfile directive. Alternatively, configuration settings can be provided at runtime using the –config parameter in the Snakemake launcher. Rules can then access the values of configured keys using standard Python dictionary notation on a variable named config.
ResourcesSpecifying the computational requirements within a rule can assist with appropriate job scheduling and runtime behaviour. Perhaps the most useful such requirement is for RAM usage, which is specified by the mem key within an item in the resources directive of the rule. For example, specifying that executing a job from the rule is likely to use 5 GB of RAM allows Snakemake to schedule parallel jobs such that the RAM budget of the computational infrastructure is not exceeded. Similarly, the runtime key can be used to indicate the maximum time that a job would take to complete; this is particularly useful for cluster-based execution where such information typically needs to be provided to the cluster scheduler. Other standard resources that can be provided are disk (the required amount of disk space), gpu (the number of GPUs required), and tmpdir (the location for temporary storage).
Workflows are typically run on computing platforms with multiple ‘cores’, which provides the capacity for parallel processing both within and between jobs. The threads directive in a rule sets the number of cores that are allocated to a job that is executed from the rule; the default is one thread. Increasing the thread count can be useful to speed up the execution of programs that utilise multi-threading, such as some of the AFNI programsFootnote 11; for example, we set the thread count in the coregistration rule to two (Fig. 4B) to allow programs called by align_epi_anat.py to use two cores. In addition to such within-job parallelisation, Snakemake also provides between-job parallelisation by executing multiple jobs simultaneously when the total number of threads across running jobs does not exceed the number of available cores; for example, a four-core computing environment would allow Snakemake to execute two coregistration jobs concurrently.
Comments (0)