SAGEGly: Using multimodal information to predict proteinglycan binding sites with the Graph Sample and Aggregate Networks Framework
We provide data, please download from the link: https://zenodo.org/records/20037375
Glycan_binding.zip : 10.9GB
PPI.zip : 15.1GB
biopython==1.83
fair-esm===2.0.0
matplotlib==3.7.5
numpy==1.24.1
pandas==2.0.3
python==3.8.19
scikit-learn==1.3.2
scipy==1.10.1
torch-cluster==1.6.3+pt24cu118
torch-geometric==2.6.1
torch-scatter== 2.1.2+pt24cu118
torch-sparse== 0.6.18+pt24cu118
torch-spline-conv==1.2.2+pt24cu118
torchaudio==2.4.1+cu118
torchvision==0.19.1+cu118
freesasa==2.0.3.post7
Preprocessing the embeddings of the PPI prediction task data.
The training and test embedding data used for model training can be generated using the following notebooks:
/PPI/train_embedding.ipynb
/PPI/test_embedding.ipynbThe coordinates, sequences, solvent-accessible surface area, and ESM parameters of GBPs and glycoproteins are available from Zenodo:https://zenodo.org/records/20037375. These files are required to run the embedding scripts.
Training the PPI prediction model.
Run the following command:
python /PPI/train.pyThe trained model parameters are saved in the following file:
/ppi_train.ptFiltering candidate binding sites based on geometric features.
Run the following command:
/Geometric_filtering/geometric_filtering.ipynbPreprocessing the embeddings of the glycan-binding prediction task data.
The training and test embedding data used for model training can be generated using the following notebooks:
/Glycan_binding/sugar_embedding_train.ipynb
/Glycan_binding/sugar_embedding_test.ipynbThe glycan information, coordinates, sequences, solvent-accessible surface area, and ESM parameters of GBPs and glycoproteins are available from Zenodo:https://zenodo.org/records/20037375. These files are required to run the embedding scripts.
In addition, the processed embedding data are also available from Zenodo:https://zenodo.org/records/20037375 in the following directories:
/glycan binding/train_npz
/glycan binding/test_npzYou can training the glycan-binding prediction model with the npz files.
Run the following command:
python /Glycan_binding/train_glycan.pyThe trained model parameters are available in the following file:
/glycan_binding.ptVisualization process.
The visualization can be performed using the following notebook:
/Visualization/Visualization.ipynbThe model parameters used for visualization are available in the following file:
/glycan_binding.ptSAGEGly: Using multimodal information to predict protein-glycan binding sites with the Graph Sample and Aggregate Networks Framework.