Skip to content
zy0114-sudoPublic

About

SAGEGly: Using multimodal information to predict protein-glycan binding sites with the Graph Sample and Aggregate Networks Framework

Resources

Stars

4 stars

Watchers

0 watching

Forks

Repository files navigation

SAGEGly

SAGEGly: Using multimodal information to predict proteinglycan binding sites with the Graph Sample and Aggregate Networks Framework

SAGEGly Pipeline:

SAGEGly Framework

Performance of the SAGEGly model:

image

Comparison with SOTA Methods:

image

Data download:

We provide data, please download from the link: https://zenodo.org/records/20037375

Glycan_binding.zip : 10.9GB

PPI.zip : 15.1GB

Dependencies

biopython==1.83
fair-esm===2.0.0
matplotlib==3.7.5
numpy==1.24.1
pandas==2.0.3
python==3.8.19
scikit-learn==1.3.2
scipy==1.10.1
torch-cluster==1.6.3+pt24cu118
torch-geometric==2.6.1
torch-scatter== 2.1.2+pt24cu118
torch-sparse== 0.6.18+pt24cu118
torch-spline-conv==1.2.2+pt24cu118
torchaudio==2.4.1+cu118
torchvision==0.19.1+cu118
freesasa==2.0.3.post7

1 PPI Data Preparation

Preprocessing the embeddings of the PPI prediction task data.

The training and test embedding data used for model training can be generated using the following notebooks:

/PPI/train_embedding.ipynb
/PPI/test_embedding.ipynb

The coordinates, sequences, solvent-accessible surface area, and ESM parameters of GBPs and glycoproteins are available from Zenodo:https://zenodo.org/records/20037375. These files are required to run the embedding scripts.

2 PPI Data Training

Training the PPI prediction model.

Run the following command:

python /PPI/train.py

The trained model parameters are saved in the following file:

/ppi_train.pt

3 Geometric Filtering

Filtering candidate binding sites based on geometric features.

Run the following command:

/Geometric_filtering/geometric_filtering.ipynb

4 Glycan-Binding Data Preparation

Preprocessing the embeddings of the glycan-binding prediction task data.

The training and test embedding data used for model training can be generated using the following notebooks:

/Glycan_binding/sugar_embedding_train.ipynb
/Glycan_binding/sugar_embedding_test.ipynb

The glycan information, coordinates, sequences, solvent-accessible surface area, and ESM parameters of GBPs and glycoproteins are available from Zenodo:https://zenodo.org/records/20037375. These files are required to run the embedding scripts.

In addition, the processed embedding data are also available from Zenodo:https://zenodo.org/records/20037375 in the following directories:

/glycan binding/train_npz
/glycan binding/test_npz

5 Glycan-Binding Data Training

You can training the glycan-binding prediction model with the npz files.

Run the following command:

python /Glycan_binding/train_glycan.py

The trained model parameters are available in the following file:

/glycan_binding.pt

6 Visualization

Visualization process.

The visualization can be performed using the following notebook:

/Visualization/Visualization.ipynb

The model parameters used for visualization are available in the following file:

/glycan_binding.pt

Reference and cite content

SAGEGly: Using multimodal information to predict protein-glycan binding sites with the Graph Sample and Aggregate Networks Framework.

About

SAGEGly: Using multimodal information to predict protein-glycan binding sites with the Graph Sample and Aggregate Networks Framework

Resources

Stars

4 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages