Repository navigation
Support offline workflows #47
Description
Activity
For caching remote inputs, I found snakemake-storage-plugin-cached-http, which allows skipping remote file checks. It does not support general URLs yet, but I was able to test a pre-release that worked with general URLs without internet connection as long as the files existed in the cache with envvar
SNAKEMAKE_STORAGE_CACHED_HTTP_SKIP_REMOTE_CHECKS=True.I think we'd still want to keep the remote checks when there's internet connection so we could either add a check for network connection to toggle on
skip_remote_checksin the storage settings or direct users to provide the envvar when running without internet.This does not support AWS S3, so for that to work, we'd still have to contribute to the existing plugin, create a cached-s3 plugin, or maybe try to add the skip remote checks feature to the general interface in snakemake-interface-storage-plugins.
Remote file access via snakemake (involving snakemake, snakemake_interface_storage_plugins, and the specific storage plugin, e.g. snakemake_storage_plugin_s3) involves two steps
exists()→exists_in_storage()which probes the remote (s3)retrieve_from_storage()which performs the cache validation, and optionally fetches & saves the file
--not-retrieve-storageskips (2) but not (1) which is why it requires us to be online.nextstrain/measles#147 shows two ways to solve this by modifying our shared
path_or_url()rather than snakemake itself.Approach 1 (the first commit) uses
NEXTSTRAIN_ALLOW_OFFLINE=1to use the cached version if it exists. This works well but has a big flaw: if we're online and we have a cached copy then we'll never revalidate the cached copy.Approach 2 (the second commit) checks to see if we're online by probing the resource (e.g. HeadBucket API call). I think this is exactly the behaviour we want, although we may want to add some ENV variable / config overrides, e.g.
NEXTSTRAIN_REVALIDATE_REMOTE_INPUTS.Thanks for prototyping this in nextstrain/measles#147! It's really helpful to see the Snakemake
AnnotatedStringvalues 🙏I think (1) would be more intuitive by flipping the envvar to the opposite, e.g.
NEXTSTRAIN_USE_REMOTE_INPUTS:online cache exists envvar NEXTSTRAIN_USE_REMOTE_INPUTSbehavior true false true fetch true false false error true true true revalidate true true false use local false false true error false false false error false true true error false true false use local If we're online, have a cached, and set
NEXTSTRAIN_USE_REMOTE_INPUTS=false, I think it's a feature that we can skip downloading the remote files even if they are "newer".Yeah - more useful for power users, but how does this work for a generic
nextstrain run measlesworkflow where we want it to Just Work? I.e. if online then fetch new data, if offline then use the cache.Hmm, this can probably a more generic network check (StackOverflow) rather than a provider specific check at the start of the workflow:
if "NEXTSTRAIN_USE_REMOTE_INPUTS" in os.environ: use_remote_inputs = bool(os.environ["NEXTSTRAIN_USE_REMOTE_INPUTS"]) else: use_remote_inputs = is_online()
Just seeing this discussion from the link in https://github.com/nextstrain/nextstrain-desktop/issues/15 (somehow I wasn't watching this repo for notifications... fixed).
+1 for the behavior in Jover's table with the default being "use remote inputs". Can we read the setting from
configinstead of env var?Reacted by Jover LeePushed up nextstrain/measles@b238682 as an example of using both
config.use_remote_filesand network check to determine workflow behavior.Adding support for local cache in nextstrain/shared#87 and updated nextstrain/measles#147 to test it.
Design pathogen workflows to support running the core workflow offline.