Add SDRF annotations for 19 heart-tissue proteomics datasets - #61
Add SDRF annotations for 19 heart-tissue proteomics datasets#61enriquea wants to merge 2 commits into
Conversation
Human cardiac tissue and cardiac-disease datasets from PRIDE Archive, covering coronary/ischaemic disease, congenital heart disease, valvular disease, cardiomyopathy, heart failure (HFpEF/HFrEF) and cardiac ageing, plus four animal cardiac-disease models. 916 sample rows total. Every row maps to a raw file that exists in the corresponding PRIDE deposit; all ontology terms resolved against OLS4. All files pass `parse_sdrf validate-sdrf --use_ols_cache_only`. Annotated with sdrf-skills (github.com/bigbio/sdrf-skills).
Qodo reviews are paused for this user.Troubleshooting steps vary by plan Learn more → On a Teams plan? Using GitHub Enterprise Server, GitLab Self-Managed, or Bitbucket Data Center? |
|
Important Review skippedReview was skipped due to path filters ⛔ Files ignored due to path filters (19)
CodeRabbit blocks several paths by default. You can override this behavior by explicitly including those paths in the path filters. For example, including ⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Pro Plus Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
The review gate flagged 10 collisions. The deposited pooled runs are two distinct preparations that were wrongly merged under one source name: unfractionated repeat injections (Pooled_1_SHALLOW, Pooled_2..9) and offline high-pH RP fractions (Frac1..16, four re-injected). Split into Pooled_reference_unfractionated (fraction 1, technical replicates 1-9) and Pooled_reference_hpH_RP (fraction N, technical replicate 1-2), so every row has a unique (source name, fraction identifier, technical replicate) coordinate.
Add SDRF annotations for 19 heart-tissue proteomics datasets
A coherent batch of cardiac annotations: human heart tissue across coronary/ischaemic,
congenital, valvular, cardiomyopathy and heart-failure phenotypes, plus four animal
cardiac-disease models. 916 sample rows, all passing
parse_sdrf validate-sdrf.Datasets
Templates:
ms-proteomicsv1.1.0 +humanv1.1.0 (orvertebratesv1.1.0 for theanimal models), with
clinical-metadatav1.0.0 anddia-acquisitionv1.1.0 where used.Validation
sdrf-pipelinesbuilt from GitHubmain(0.1.6), run on these exact paths.comment[data file]value was cross-checked against thelive PRIDE file list (
/projects/{acc}/files/all); 916/916 rows match a deposited run.an exact search for
coronary artery diseasein MONDO returns a susceptibility locusrather than the disorder, and PRIDE's DIA term is
PRIDE:0000450.PXD025096 deposit
.mzXML/.mzMLconversions alongside the.rawfiles, and only the.rawacquisitions are annotated.Sample-metadata provenance
Per-sample assignments come from the deposited run names and, where available, from
open-access supplementary tables:
lists every valve by SKU, tissue type (Normal / pCAVS / AVI) and age. The 20 SKUs flagged for
proteomics match the 20 donors in the run names exactly (DB17 is marked not-included and has
no raw file). Ages are annotated per donor.
PMC12132713.
DatasetS1(DDA) is keyed on theexact IDs used in PXD060431's run names, and
DatasetS13(DIA) on the biopsy IDs embedded inPXD045677's run names, giving per-patient age, sex and ancestry. Note that "HFpEF" inside the
PXD045677 mzML names is the study name, not the group — group comes from the biopsy-ID
lookup, which is how the five non-failing DIA controls are correctly identified.
Where a grouping variable was not recoverable it was left out rather than guessed, and each
case is documented in the file header notes:
aortic valve disorderis assigned to all rows; the bicuspid/tricuspid split is not in thefile names.
characteristics[sex]isnot available.not availablefor disease rather than an inferred group.Points for reviewer confirmation
CMV*in the run names is read as control mitral valve, based on the exact20 + 7 = 27 split matching the stated n=27 and the diseased-vs-control design. Worth a check
against the source publication.
The explicit protocol statement was used and the conflict noted in each file's header:
PXD039662 (PRIDE "Q Exactive HF" vs protocol "Exploris 480") and PXD079292 (PRIDE
"Q Exactive HF" vs protocol "timsTOF HT" with DIA-NN).
matches the DDA arm most closely and DDA is annotated, but the run names do not state the mode.
and phosphoproteome accessions, because the submitter's data-processing protocol declares it for
both. Annotated as declared rather than second-guessed.
Annotation method
Produced with sdrf-skills (agent-assisted), then
validated with
sdrf-pipelines. Candidate selection swept 1,379 PRIDE cardiac projects andexcluded datasets that fail the "heart tissue" criterion on inspection — iPSC-CM cultures,
plasma/serum cohorts, epicardial adipose tissue and pericardial fluid — as well as TMT
experiments deposited without a channel-to-sample key.
Two data-quality observations for the maintainers, unrelated to these files:
Homo sapiens, but its data-processing protocol searches a mouse database and the description
says "murine heart". Excluded here for that reason.
python -m tools checkin sdrf-skills reports valid accessions as hallucinated — it flagsNT=Homo sapiens;AC=NCBITaxon:9606while the same repo'spython -m tools verify NCBITaxon:9606resolves it correctly.