Repository navigation
Storage4Climate Beginners Guide
- Get a VSC-account.
- Request access to relevant project pages and groups
- (Optional, but recommended) Follow the Linux-intro training, look at the training schedule.
- (Optional, but recommended) Follow the HPC-intro training, look at the training schedule.
- Follow the Tutorial.
<Add a short summary sentence here explaining the purpose of Storage4Climate (e.g., shared infrastructure for climate data storage and compute within the consortium).>
The Storage4Climate consortium consists of multiple research groups across Belgian institutions.
| Account Number | Group | Institute | PI | Representative |
|---|---|---|---|---|
| 2022_200 | Storage4Climate consortium | - | - | - |
| 2022_201 | BCLIMATE | VUB | Wim Thiery | Shannon de Roos |
| 2022_202 | RCS | KU Leuven | Nicole van Lipzig | Inne Vanderkelen |
| 2022_203 | RSDA | KU Leuven | Gabrielle De Lannoy | Jonas Mortelmans |
| 2022_204 | H_CEL | UGent | Diego Miralles | Damián Insua-Costa |
| 2022_205 | RMIB-UGENT | RMI & UGent | Kwinten van Weverberg | Kobe Vandelanotte |
| 2022_206 | BGLACIER | VUB | Harry Zekollari | Paul Muñoz |
(Describe the Tier-1 compute environment here — e.g., cluster usage, job submission, modules, etc.)
-
project_input/→ Input datasets shared within the project -
project_output/→ Results and processed outputs -
external/→ External datasets (e.g., downloaded or third-party data)
- Use
project_input/for read-only shared input data - Store intermediate and final results in
project_output/ - Use
external/for non-consortium datasets
(Add quotas, cleanup policies, or best practices if relevant.)
VSC provides Tier-1 Data for active research data. ManGO is the web portal for browsing the data and iRODS is the underlying technology that makes it happen.
There is a clear directory structure that mimics the structure on the cluster's project scratch. [TODO]: explain this a bit in more detail.
There are multiple ways to access and transfer your data to and from Tier-1 Data. The choice of the right tool depends very much on your use-case.
- Globus: Globus is arguably the easiest to use and provides robust transfer, sharing and search capabilities for cluster storage and external storage. Use the Globus Web App or learn to programmatically use the Globus CLI or Globus Python SDK. It is usually fast, very robust, easy to use, but lacks the possibility of handling meta-data.
- iRODS clients: Globus is a high level tool. Closer to the underlying iRODS machinery are the iRODS specific clients. They are usually the fastest option, can handle meta-data, but are less robust and usually more difficult to use.
- Mounting: Probably the easiest way to access Tier-1 Data from the cluster is by mounting it (see below for instruction on how). Your data will become accessible through the Linux file tree. No need to adapt your workflow when working on data that is on Tier-1 Data. But beware, when needing the entire contents of a file, it is much slower than the above options. When needing only small parts of a file (e.g., meta-data, one variable out of many in a NetCDF file, ...) it can be very fast though and even much faster than the above options.
To install:
cd ~/.local/bin
ln -s /readonly/dodrio/scratch/projects/2022_200/software/prg/RDM/irodsfs/irodsfs
ln -s /readonly/dodrio/scratch/projects/2022_200/software/prg/RDM/irodsfs-pool/irodsfs-pool
ln -s /readonly/dodrio/scratch/projects/2022_200/software/src/HPC/tools/utils/irods-fsmntMake sure ~/.local/bin is in your $PATH.
Usage:
- Log in to ManGO from the cluster by following the instructions on How to connect. This is needed only once on one node, and will make sure you are connected on all nodes for some time (persistent across separate login sessions).
- Start mount: run
irods-fsmnt start. Tier-1 Data will now be accessible through/tmp/$USER/mnt/irods. - Stop mount: run
irods-fsmnt stop. - The mount runs with some default options, see help menu to see how to override the defaults: run
irods-fsmnt. - The mount will, by default, run with caching enabled, meaning that if you access the same part of a file multiple times, it will only be downloaded once. In some cases, though, you might not want that (there is some overhead). You can turn it off by adding the
--no-pooloption when starting the mount:irods-fsmnt start --no-pool.
Note
On a login node, the mount will stay active and is shared across separate login sessions.
-
🔧 Dashboard for setting up tunnels
-
📚 Tier-1 Documentation (Hortense)
-
💬 Storage4Climate Teams
-
💻 Storage4Climate GitHub Repository