๐ช Hooks and Presets#
Fetchez is designed to be highly extendable. Instead of just downloading files, you can build automated pipelines that process data on the fly.
Processing Hooks#
Fetchez includes a powerful Hook System that allows you to chain actions together. Hooks run in a pipeline, meaning the output of one hook (e.g., unzipping a file) becomes the input for the next (e.g., streaming and processing it).
There are four stages in the Hook lifecycle:
PRE/MANIFEST Stage: (
prestage) Runs before any data is downloaded (e.g., filtering URLs, masking regions).FILE Stage: Runs on each individual file as it is downloaded (e.g., unzipping, converting formats, or piping to stdout).
STREAM Stage: Runs after the file stage where we can stream the fetched data through
readers. (e.g. processing chunked data streams).POST/COLLECTION Stage: (
poststage) Runs after all files are downloaded (e.g., merging grids, calculating checksums).
Each hook defines itโs default stage, which can be changed at any time (though wonโt always work as expected outside of their desred stage, so be careful changing them).
Common Built-in Hooks:#
unzip: Automatically extracts.zipor.gzfiles.pipe: Prints the final absolute path to stdout (useful for piping to GDAL/PDAL).audit: Generates a JSON manifest of everything downloaded and processed.exec: Run a shell command on a file (uses โ{file}โ formatter).
Example (CLI):#
# Download data.zip
# Extract data.tif (via unzip hook)
# Print /path/to/data.tif (via pipe hook)
fetchez run charts --hook unzip --hook pipe
# warp the copernicus files right when they're downloaded
fetchez run -R loc:denver copernicus --pipe | xargs gdalwarp -t_srs EPSG:3857
# build a vrt of the fetched files
gdalbuildvrt cop_merged.vrt $(fetchez run -R -105/-104/39/40 copernicus --pipe)
Pipeline Presets (Macros)#
Tired of typing the same chain of hooks every time? Presets allow you to define reusable workflow macros.
Instead of running this long command:
fetchez run copernicus --hook checksum:algo=sha256 --hook enrich --hook audit:file=log.json
You can define a preset and simply run:
fetchez run copernicus --audit-full
How to create a Preset:#
Presets are simply YAML files that live in your ~/.fetchez/presets/ directory or are provided as an extension or by the community. fetchez automatically scans this folder and the PresetRegistry and turns any valid YAML file into a valid hook.
Create a file: ~/.fetchez/hooks/presets/audit_full.yaml
Define your workflow:
name: audit-full
description: Generate SHA256 hashes, enrichment, and a full JSON audit logs.
hooks:
- name: checksum
args:
algo: sha256
- name: enrich
- name: audit
args:
file: audit_full.json
Run it: Your new preset automatically appears in fetchez as a valid hook!
fetchez run charts --hook audit-full
Extending Hooks and Presets (Plugins and Extensions)#
Fetchez is generic. If you are building a custom tool and want to create your own processing hooks and presets, you can register your own hooks and presets either in your project or in the .fetchez configuration directory and they will be discoverable with the fetchez.registry.HookRegistry and fetchez.registry.PresetRegistry
In your project, make a directory called โhooksโ and/or โhooks/presetsโ; add any python hooks and YAML presets to the appropriate directory and register them with fetchez in your pyproject.toml:
[project.entry-points."fetchez.hooks"]
my_project_hooks = "my_project.hooks"
[project.entry-points."fetchez.hooks.presets"]
my_project_presetes = "my_project.hooks.presets"