Skip to content

Restore 1,173 datasets deleted by PR #57 (keep batch-5) - #59

Merged
ypriverol merged 2 commits into
mainfrom
revert/restore-datasets-pr57
Jul 31, 2026
Merged

Restore 1,173 datasets deleted by PR #57 (keep batch-5)#59
ypriverol merged 2 commits into
mainfrom
revert/restore-datasets-pr57

Conversation

@ypriverol

@ypriverol ypriverol commented Jul 31, 2026

Copy link
Copy Markdown
Member

Restore datasets deleted by PR #57

PR #57 was merged with a branch whose datasets/ tree had been reduced to only
its 10 new files (a bad git rm -r datasets/ during a conflict fix), so merging
it deleted 1,173 previously-annotated datasets from main.

This PR:

Net effect vs current main: 1,173 datasets restored, 0 deletions.

Summary by CodeRabbit

  • New Features
    • Added comprehensive sample metadata for dataset PXD063467, covering 96 proteomics samples.
    • Included DDA and DIA acquisition annotations for AC16 and HCM cardiomyocyte controls and treatments.
    • Documented canagliflozin, dapagliflozin, and empagliflozin exposures at low and high concentrations.
    • Added experimental details including replicates, instruments, sample preparation, modifications, data files, and SDRF versions.

@qodo-code-review

Copy link
Copy Markdown

Qodo reviews are paused for this user.

Troubleshooting steps vary by plan Learn more →

On a Teams plan?
Reviews resume once this user has a paid seat and their Git account is linked in Qodo.
Link Git account →

Using GitHub Enterprise Server, GitLab Self-Managed, or Bitbucket Data Center?
These require an Enterprise plan - Contact us
Contact us →

@coderabbitai

coderabbitai Bot commented Jul 31, 2026

Copy link
Copy Markdown

Review Change Stack

📝 Walkthrough

Walkthrough

Added DDA and DIA SDRF files for PXD063467. The annotations cover 96 AC16 and HCM samples with control and three drug-treatment conditions, concentration levels, replicates, exposure metadata, acquisition details, and data files.

Changes

PXD063467 SDRF annotations

Layer / File(s) Summary
DDA sample annotations
datasets/PXD063467/PXD063467_dda.sdrf
Added 96 annotated DDA records and 98 blank tab-delimited rows.
AC16 DIA annotations
datasets/PXD063467/PXD063467_dia.sdrf
Added AC16 control, canagliflozin, dapagliflozin, and empagliflozin records with concentration, replicate, exposure, acquisition, preparation, and file metadata.
HCM DIA annotations
datasets/PXD063467/PXD063467_dia.sdrf
Added HCM control and drug-treatment records. One canagliflozin filename contains DDA, and one empagliflozin record uses SDRF version v1.1.1.

Estimated code review effort: 2 (Simple) | ~10 minutes

Possibly related PRs

Suggested reviewers: nithujohn

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly summarizes the primary change: restoring 1,173 datasets deleted by PR #57 while retaining batch-5 datasets.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch revert/restore-datasets-pr57

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@ypriverol
ypriverol requested a review from nithujohn July 31, 2026 09:52

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@datasets/PXD063467/PXD063467_dda.sdrf`:
- Line 91: Correct the shared HCM42 low-dose source label by changing
PXD063467_HCM42_Empa_low5 to PXD063467_HCM42_Empa_low6 in
datasets/PXD063467/PXD063467_dda.sdrf lines 91-91 and
datasets/PXD063467/PXD063467_dia.sdrf lines 91-91, leaving the remaining SDRF
fields unchanged.
- Around line 1-2: Include both datasets/PXD063467/PXD063467_dda.sdrf lines 1-2
and datasets/PXD063467/PXD063467_dia.sdrf lines 1-2 in CI validation by either
renaming both files to use the .sdrf.tsv suffix or updating
.github/workflows/validate-sdrf.yml to match .sdrf files.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

Comment on lines +1 to +2
source name characteristics[organism] characteristics[organism part] characteristics[disease] characteristics[cell line] characteristics[cellosaurus accession] characteristics[cellosaurus name] characteristics[biorepository] characteristics[age] characteristics[sex] characteristics[compound] characteristics[biological replicate] characteristics[compound] characteristics[exposure duration] assay name technology type comment[fraction identifier] comment[technical replicate] comment[proteomics data acquisition method] comment[instrument] comment[cleavage agent details] comment[label] comment[reduction reagent] comment[alkylation reagent] comment[modification parameters] comment[modification parameters] comment[modification parameters] comment[data file] comment[sdrf version] comment[sdrf annotation tool]
PXD063467_1_Control1 Homo sapiens heart normal AC16 CVCL_4U18 AC16 [Human hybrid cardiomyocyte] Merck not available not available not applicable 1 not applicable not applicable 1_AC16_DDA_GA1_1_3985 proteomic profiling by mass spectrometry 1 1 data-dependent acquisition "NT=timsTOF HT;AC=MS:1003404" "NT=Trypsin;AC=MS:1001251" label free sample "NT=TCEP;AC=PRIDE:0000609" "NT=NEM;AC=PRIDE:0000606" "NT=Oxidation;MT=Variable;TA=M;AC=Unimod:35" "NT=Nethylmaleimide;MT=Variable;TA=C;AC=Unimod:108" "NT=NEM:2H(5);MT=Variable;TA=C;AC=Unimod:776" 1_AC16_DDA_GA1_1_3985.d v1.1.0 manual curation

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win

Make both files part of the CI validation set.

The supplied workflow matches only .sdrf.tsv, but both changed paths end in .sdrf. Their records can therefore merge without parse_sdrf validate-sdrf.

  • datasets/PXD063467/PXD063467_dda.sdrf#L1-L2: rename the file to .sdrf.tsv, or update .github/workflows/validate-sdrf.yml to include .sdrf.
  • datasets/PXD063467/PXD063467_dia.sdrf#L1-L2: rename the file to .sdrf.tsv, or update .github/workflows/validate-sdrf.yml to include .sdrf.
📍 Affects 2 files
  • datasets/PXD063467/PXD063467_dda.sdrf#L1-L2 (this comment)
  • datasets/PXD063467/PXD063467_dia.sdrf#L1-L2
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@datasets/PXD063467/PXD063467_dda.sdrf` around lines 1 - 2, Include both
datasets/PXD063467/PXD063467_dda.sdrf lines 1-2 and
datasets/PXD063467/PXD063467_dia.sdrf lines 1-2 in CI validation by either
renaming both files to use the .sdrf.tsv suffix or updating
.github/workflows/validate-sdrf.yml to match .sdrf files.

PXD063467_HCM39_Empa_low3 Homo sapiens heart normal HCM not available not available PromoCell not available not available empagliflozin 3 10 nM 24 hour HCM39_DDA_GA4_1_4337 proteomic profiling by mass spectrometry 1 1 data-dependent acquisition "NT=timsTOF HT;AC=MS:1003404" "NT=Trypsin;AC=MS:1001251" label free sample "NT=TCEP;AC=PRIDE:0000609" "NT=NEM;AC=PRIDE:0000606" "NT=Oxidation;MT=Variable;TA=M;AC=Unimod:35" "NT=Nethylmaleimide;MT=Variable;TA=C;AC=Unimod:108" "NT=NEM:2H(5);MT=Variable;TA=C;AC=Unimod:776" HCM39_DDA_GA4_1_4337.d v1.1.0 manual curation
PXD063467_HCM40_Empa_low4 Homo sapiens heart normal HCM not available not available PromoCell not available not available empagliflozin 4 10 nM 24 hour HCM40_DDA_GA1_1_4240 proteomic profiling by mass spectrometry 1 1 data-dependent acquisition "NT=timsTOF HT;AC=MS:1003404" "NT=Trypsin;AC=MS:1001251" label free sample "NT=TCEP;AC=PRIDE:0000609" "NT=NEM;AC=PRIDE:0000606" "NT=Oxidation;MT=Variable;TA=M;AC=Unimod:35" "NT=Nethylmaleimide;MT=Variable;TA=C;AC=Unimod:108" "NT=NEM:2H(5);MT=Variable;TA=C;AC=Unimod:776" HCM40_DDA_GA1_1_4240.d v1.1.0 manual curation
PXD063467_HCM41_Empa_low5 Homo sapiens heart normal HCM not available not available PromoCell not available not available empagliflozin 5 10 nM 24 hour HCM41_DDA_GA5_1_4244 proteomic profiling by mass spectrometry 1 1 data-dependent acquisition "NT=timsTOF HT;AC=MS:1003404" "NT=Trypsin;AC=MS:1001251" label free sample "NT=TCEP;AC=PRIDE:0000609" "NT=NEM;AC=PRIDE:0000606" "NT=Oxidation;MT=Variable;TA=M;AC=Unimod:35" "NT=Nethylmaleimide;MT=Variable;TA=C;AC=Unimod:108" "NT=NEM:2H(5);MT=Variable;TA=C;AC=Unimod:776" HCM41_DDA_GA5_1_4244.d v1.1.0 manual curation
PXD063467_HCM42_Empa_low5 Homo sapiens heart normal HCM not available not available PromoCell not available not available empagliflozin 6 10 nM 24 hour HCM42_DDA_GA7_1_4246 proteomic profiling by mass spectrometry 1 1 data-dependent acquisition "NT=timsTOF HT;AC=MS:1003404" "NT=Trypsin;AC=MS:1001251" label free sample "NT=TCEP;AC=PRIDE:0000609" "NT=NEM;AC=PRIDE:0000606" "NT=Oxidation;MT=Variable;TA=M;AC=Unimod:35" "NT=Nethylmaleimide;MT=Variable;TA=C;AC=Unimod:108" "NT=NEM:2H(5);MT=Variable;TA=C;AC=Unimod:776" HCM42_DDA_GA7_1_4246.d v1.1.0 manual curation

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🗄️ Data Integrity & Integration | 🟡 Minor | ⚡ Quick win

Correct the shared HCM42 low-dose label.

Both records identify biological replicate 6 as HCM42_Empa_low5. The source name should use HCM42_Empa_low6.

  • datasets/PXD063467/PXD063467_dda.sdrf#L91-L91: change PXD063467_HCM42_Empa_low5 to PXD063467_HCM42_Empa_low6.
  • datasets/PXD063467/PXD063467_dia.sdrf#L91-L91: change PXD063467_HCM42_Empa_low5 to PXD063467_HCM42_Empa_low6.
📍 Affects 2 files
  • datasets/PXD063467/PXD063467_dda.sdrf#L91-L91 (this comment)
  • datasets/PXD063467/PXD063467_dia.sdrf#L91-L91
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@datasets/PXD063467/PXD063467_dda.sdrf` at line 91, Correct the shared HCM42
low-dose source label by changing PXD063467_HCM42_Empa_low5 to
PXD063467_HCM42_Empa_low6 in datasets/PXD063467/PXD063467_dda.sdrf lines 91-91
and datasets/PXD063467/PXD063467_dia.sdrf lines 91-91, leaving the remaining
SDRF fields unchanged.

@ypriverol
ypriverol merged commit 813d6ac into main Jul 31, 2026
2 of 3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants