Skip to content

Storage4Climate Beginners Guide

Inne Vanderkelen edited this page Jun 9, 2026 · 21 revisions

1. Getting Started

2. Storage4Climate: What it is and what it is not

<Add a short summary sentence here explaining the purpose of Storage4Climate (e.g., shared infrastructure for climate data storage and compute within the consortium).>

2.1 Consortium Partners

The Storage4Climate consortium consists of multiple research groups across Belgian institutions.

Account Number Group Institute PI Representative
2022_200 Storage4Climate consortium - - -
2022_201 BCLIMATE VUB Wim Thiery Shannon de Roos
2022_202 RCS KU Leuven Nicole van Lipzig Inne Vanderkelen
2022_203 RSDA KU Leuven Gabrielle De Lannoy Jonas Mortelmans
2022_204 H_CEL UGent Diego Miralles Damián Insua-Costa
2022_205 RMIB-UGENT RMI & UGent Kwinten van Weverberg Kobe Vandelanotte
2022_206 BGLACIER VUB Harry Zekollari Paul Muñoz

2.2 Tier-1 Compute

(Describe the Tier-1 compute environment here — e.g., cluster usage, job submission, modules, etc.)

2.3 Cluster storage: SCRATCH Overview & File Structure

Directory Structure

  • project_input/ → Input datasets shared within the project
  • project_output/ → Results and processed outputs
  • external/ → External datasets (e.g., downloaded or third-party data)

Usage Guidelines

  • Use project_input/ for read-only shared input data
  • Store intermediate and final results in project_output/
  • Use external/ for non-consortium datasets

(Add quotas, cleanup policies, or best practices if relevant.)

2.4 External Storage: Tier-1 Data aka iRODS aka ManGO

Purpose

VSC provides Tier-1 Data for active research data. ManGO is the web portal for browsing the data and iRODS is the underlying technology that makes it happen.

Directory Structure

There is a clear directory structure that mimics the structure on the cluster's project scratch. [TODO]: explain this a bit in more detail.

Access and Transfers

There are multiple ways to access and transfer your data to and from Tier-1 Data. The choice of the right tool depends very much on your use-case.

  • Globus: Globus is arguably the easiest to use and provides robust transfer, sharing and search capabilities for cluster storage and external storage. Use the Globus Web App or learn to programmatically use the Globus CLI or Globus Python SDK. It is usually fast, very robust, easy to use, but lacks the possibility of handling meta-data.
  • iRODS clients: Globus is a high level tool. Closer to the underlying iRODS machinery are the iRODS specific clients. They are usually the fastest option, can handle meta-data, but are less robust and usually more difficult to use.
  • Mounting: Probably the easiest way to access Tier-1 Data from the cluster is by mounting it (see below for instruction on how). Your data will become accessible through the Linux file tree. No need to adapt your workflow when working on data that is on Tier-1 Data. But beware, when needing the entire contents of a file, it is much slower than the above options. When needing only small parts of a file (e.g., meta-data, one variable out of many in a NetCDF file, ...) it can be very fast though and even much faster than the above options.

Mounting (read-only)

To install:

cd ~/.local/bin
ln -s /readonly/dodrio/scratch/projects/2022_200/software/prg/RDM/irodsfs/irodsfs
ln -s /readonly/dodrio/scratch/projects/2022_200/software/prg/RDM/irodsfs-pool/irodsfs-pool
ln -s /readonly/dodrio/scratch/projects/2022_200/software/src/HPC/tools/utils/irods-fsmnt

Make sure ~/.local/bin is in your $PATH.

Usage:

  • Log in to ManGO from the cluster by following the instructions on How to connect. This is needed only once on one node, and will make sure you are connected on all nodes for some time (persistent across separate login sessions).
  • Start mount: run irods-fsmnt start. Tier-1 Data will now be accessible through /tmp/$USER/mnt/irods.
  • Stop mount: run irods-fsmnt stop.
  • The mount runs with some default options, see help menu to see how to override the defaults: run irods-fsmnt.
  • The mount will, by default, run with caching enabled, meaning that if you access the same part of a file multiple times, it will only be downloaded once. In some cases, though, you might not want that (there is some overhead). You can turn it off by adding the --no-pool option when starting the mount: irods-fsmnt start --no-pool.

Note

On a login node, the mount will stay active and is shared across separate login sessions.

3. Platforms & Useful Links