Restore 1,173 datasets deleted by PR #57 (keep batch-5) - #59
Conversation
Qodo reviews are paused for this user.Troubleshooting steps vary by plan Learn more → On a Teams plan? Using GitHub Enterprise Server, GitLab Self-Managed, or Bitbucket Data Center? |
📝 WalkthroughWalkthroughAdded DDA and DIA SDRF files for PXD063467. The annotations cover 96 AC16 and HCM samples with control and three drug-treatment conditions, concentration levels, replicates, exposure metadata, acquisition details, and data files. ChangesPXD063467 SDRF annotations
Estimated code review effort: 2 (Simple) | ~10 minutes Possibly related PRs
Suggested reviewers: 🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Actionable comments posted: 2
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@datasets/PXD063467/PXD063467_dda.sdrf`:
- Line 91: Correct the shared HCM42 low-dose source label by changing
PXD063467_HCM42_Empa_low5 to PXD063467_HCM42_Empa_low6 in
datasets/PXD063467/PXD063467_dda.sdrf lines 91-91 and
datasets/PXD063467/PXD063467_dia.sdrf lines 91-91, leaving the remaining SDRF
fields unchanged.
- Around line 1-2: Include both datasets/PXD063467/PXD063467_dda.sdrf lines 1-2
and datasets/PXD063467/PXD063467_dia.sdrf lines 1-2 in CI validation by either
renaming both files to use the .sdrf.tsv suffix or updating
.github/workflows/validate-sdrf.yml to match .sdrf files.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
| source name characteristics[organism] characteristics[organism part] characteristics[disease] characteristics[cell line] characteristics[cellosaurus accession] characteristics[cellosaurus name] characteristics[biorepository] characteristics[age] characteristics[sex] characteristics[compound] characteristics[biological replicate] characteristics[compound] characteristics[exposure duration] assay name technology type comment[fraction identifier] comment[technical replicate] comment[proteomics data acquisition method] comment[instrument] comment[cleavage agent details] comment[label] comment[reduction reagent] comment[alkylation reagent] comment[modification parameters] comment[modification parameters] comment[modification parameters] comment[data file] comment[sdrf version] comment[sdrf annotation tool] | ||
| PXD063467_1_Control1 Homo sapiens heart normal AC16 CVCL_4U18 AC16 [Human hybrid cardiomyocyte] Merck not available not available not applicable 1 not applicable not applicable 1_AC16_DDA_GA1_1_3985 proteomic profiling by mass spectrometry 1 1 data-dependent acquisition "NT=timsTOF HT;AC=MS:1003404" "NT=Trypsin;AC=MS:1001251" label free sample "NT=TCEP;AC=PRIDE:0000609" "NT=NEM;AC=PRIDE:0000606" "NT=Oxidation;MT=Variable;TA=M;AC=Unimod:35" "NT=Nethylmaleimide;MT=Variable;TA=C;AC=Unimod:108" "NT=NEM:2H(5);MT=Variable;TA=C;AC=Unimod:776" 1_AC16_DDA_GA1_1_3985.d v1.1.0 manual curation |
There was a problem hiding this comment.
🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win
Make both files part of the CI validation set.
The supplied workflow matches only .sdrf.tsv, but both changed paths end in .sdrf. Their records can therefore merge without parse_sdrf validate-sdrf.
datasets/PXD063467/PXD063467_dda.sdrf#L1-L2: rename the file to.sdrf.tsv, or update.github/workflows/validate-sdrf.ymlto include.sdrf.datasets/PXD063467/PXD063467_dia.sdrf#L1-L2: rename the file to.sdrf.tsv, or update.github/workflows/validate-sdrf.ymlto include.sdrf.
📍 Affects 2 files
datasets/PXD063467/PXD063467_dda.sdrf#L1-L2(this comment)datasets/PXD063467/PXD063467_dia.sdrf#L1-L2
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@datasets/PXD063467/PXD063467_dda.sdrf` around lines 1 - 2, Include both
datasets/PXD063467/PXD063467_dda.sdrf lines 1-2 and
datasets/PXD063467/PXD063467_dia.sdrf lines 1-2 in CI validation by either
renaming both files to use the .sdrf.tsv suffix or updating
.github/workflows/validate-sdrf.yml to match .sdrf files.
| PXD063467_HCM39_Empa_low3 Homo sapiens heart normal HCM not available not available PromoCell not available not available empagliflozin 3 10 nM 24 hour HCM39_DDA_GA4_1_4337 proteomic profiling by mass spectrometry 1 1 data-dependent acquisition "NT=timsTOF HT;AC=MS:1003404" "NT=Trypsin;AC=MS:1001251" label free sample "NT=TCEP;AC=PRIDE:0000609" "NT=NEM;AC=PRIDE:0000606" "NT=Oxidation;MT=Variable;TA=M;AC=Unimod:35" "NT=Nethylmaleimide;MT=Variable;TA=C;AC=Unimod:108" "NT=NEM:2H(5);MT=Variable;TA=C;AC=Unimod:776" HCM39_DDA_GA4_1_4337.d v1.1.0 manual curation | ||
| PXD063467_HCM40_Empa_low4 Homo sapiens heart normal HCM not available not available PromoCell not available not available empagliflozin 4 10 nM 24 hour HCM40_DDA_GA1_1_4240 proteomic profiling by mass spectrometry 1 1 data-dependent acquisition "NT=timsTOF HT;AC=MS:1003404" "NT=Trypsin;AC=MS:1001251" label free sample "NT=TCEP;AC=PRIDE:0000609" "NT=NEM;AC=PRIDE:0000606" "NT=Oxidation;MT=Variable;TA=M;AC=Unimod:35" "NT=Nethylmaleimide;MT=Variable;TA=C;AC=Unimod:108" "NT=NEM:2H(5);MT=Variable;TA=C;AC=Unimod:776" HCM40_DDA_GA1_1_4240.d v1.1.0 manual curation | ||
| PXD063467_HCM41_Empa_low5 Homo sapiens heart normal HCM not available not available PromoCell not available not available empagliflozin 5 10 nM 24 hour HCM41_DDA_GA5_1_4244 proteomic profiling by mass spectrometry 1 1 data-dependent acquisition "NT=timsTOF HT;AC=MS:1003404" "NT=Trypsin;AC=MS:1001251" label free sample "NT=TCEP;AC=PRIDE:0000609" "NT=NEM;AC=PRIDE:0000606" "NT=Oxidation;MT=Variable;TA=M;AC=Unimod:35" "NT=Nethylmaleimide;MT=Variable;TA=C;AC=Unimod:108" "NT=NEM:2H(5);MT=Variable;TA=C;AC=Unimod:776" HCM41_DDA_GA5_1_4244.d v1.1.0 manual curation | ||
| PXD063467_HCM42_Empa_low5 Homo sapiens heart normal HCM not available not available PromoCell not available not available empagliflozin 6 10 nM 24 hour HCM42_DDA_GA7_1_4246 proteomic profiling by mass spectrometry 1 1 data-dependent acquisition "NT=timsTOF HT;AC=MS:1003404" "NT=Trypsin;AC=MS:1001251" label free sample "NT=TCEP;AC=PRIDE:0000609" "NT=NEM;AC=PRIDE:0000606" "NT=Oxidation;MT=Variable;TA=M;AC=Unimod:35" "NT=Nethylmaleimide;MT=Variable;TA=C;AC=Unimod:108" "NT=NEM:2H(5);MT=Variable;TA=C;AC=Unimod:776" HCM42_DDA_GA7_1_4246.d v1.1.0 manual curation |
There was a problem hiding this comment.
🗄️ Data Integrity & Integration | 🟡 Minor | ⚡ Quick win
Correct the shared HCM42 low-dose label.
Both records identify biological replicate 6 as HCM42_Empa_low5. The source name should use HCM42_Empa_low6.
datasets/PXD063467/PXD063467_dda.sdrf#L91-L91: changePXD063467_HCM42_Empa_low5toPXD063467_HCM42_Empa_low6.datasets/PXD063467/PXD063467_dia.sdrf#L91-L91: changePXD063467_HCM42_Empa_low5toPXD063467_HCM42_Empa_low6.
📍 Affects 2 files
datasets/PXD063467/PXD063467_dda.sdrf#L91-L91(this comment)datasets/PXD063467/PXD063467_dia.sdrf#L91-L91
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@datasets/PXD063467/PXD063467_dda.sdrf` at line 91, Correct the shared HCM42
low-dose source label by changing PXD063467_HCM42_Empa_low5 to
PXD063467_HCM42_Empa_low6 in datasets/PXD063467/PXD063467_dda.sdrf lines 91-91
and datasets/PXD063467/PXD063467_dia.sdrf lines 91-91, leaving the remaining
SDRF fields unchanged.
Restore datasets deleted by PR #57
PR #57 was merged with a branch whose
datasets/tree had been reduced to onlyits 10 new files (a bad
git rm -r datasets/during a conflict fix), so mergingit deleted 1,173 previously-annotated datasets from
main.This PR:
Net effect vs current main: 1,173 datasets restored, 0 deletions.
Summary by CodeRabbit