Hello! I saw your recent preprint, "Scalable estimation of microbial co-occurrence networks with Variational Autoencoders" and I'm hopeful your method may solve my issues, but I wanted to touch base to see if you think this method is appropriate for my use cases/scale of problem I'm hoping to address.
Use cases
-
Detecting contaminant pairs that co-occur in microbial genomes: We built a tool to detect and remove contamination from genomes and metagenome assembled genomes. We're running this tool in the ~350k genomes in GTDBrs202. We detect contamination at the order level in about 15% of genomes. We want to use order-level lineages detected in each sample to determine if any contaminants co-occur more than would be expected by chance.
-
Detecting contaminant pairs that co-occur in "isolate" RNAseq data sets: We're creating a compendia of bacterial and archaeal isolate RNA seq data. There are ~60k isolate data sets on the SRA. As part of this, we're looking for contamination in these isolates, and we generally find that there is some (usually 2-10 species detected in each sample). We want to know if any species co-occur (e.g. is Faecalibacterium prausnitzii likely to be contaminated with it's friend Roseburia inulinivorans?
Questions
- Is this method appropriate for these use cases? If not, have you encountered something else that might work?
- What would be the training data? In both cases, I'm identifying lineages present in a sample using GTDB rs202 as the reference.
- Can this approach scale to 60k and 350k samples?
I'd appreciate any insights/feedback you'd be willing to give!
Hello! I saw your recent preprint, "Scalable estimation of microbial co-occurrence networks with Variational Autoencoders" and I'm hopeful your method may solve my issues, but I wanted to touch base to see if you think this method is appropriate for my use cases/scale of problem I'm hoping to address.
Use cases
Detecting contaminant pairs that co-occur in microbial genomes: We built a tool to detect and remove contamination from genomes and metagenome assembled genomes. We're running this tool in the ~350k genomes in GTDBrs202. We detect contamination at the order level in about 15% of genomes. We want to use order-level lineages detected in each sample to determine if any contaminants co-occur more than would be expected by chance.
Detecting contaminant pairs that co-occur in "isolate" RNAseq data sets: We're creating a compendia of bacterial and archaeal isolate RNA seq data. There are ~60k isolate data sets on the SRA. As part of this, we're looking for contamination in these isolates, and we generally find that there is some (usually 2-10 species detected in each sample). We want to know if any species co-occur (e.g. is Faecalibacterium prausnitzii likely to be contaminated with it's friend Roseburia inulinivorans?
Questions
I'd appreciate any insights/feedback you'd be willing to give!