Skip to content

About

Batch extraction of protein-only targets from protein–nucleic-acid PDB/mmCIF complexes.

Resources

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

3 Commits

Folders and files

Repository files navigation

MolStrctExtractor

MolStrctExtractor is a lightweight Python tool for extracting protein-only target structures from protein–nucleic-acid complexes stored as PDB or mmCIF files.

It is designed for benchmark preparation and structure-based protein/RNA modeling workflows where the input structures contain RNA, DNA, proteins, ligands, ions, water, or repeated crystallographic copies, but the downstream task requires protein-only PDB files.

Features

  • Batch-processes all PDB/mmCIF structures directly under one folder
  • Identifies protein and nucleic-acid polymer chains
  • Uses protein–nucleic-acid contacts to identify the relevant biological assembly
  • Retains one representative of every distinct protein entity in that assembly
  • Preserves interacting copies of the same protein when they form a substantial protein–protein interface
  • Removes RNA, DNA, water, ions, and non-protein ligands
  • Supports multiple independent protein targets from one complex
  • Supports explicit author-chain overrides for unusual complexes
  • Generates an auditable JSON report containing:
    • detected chains and polymer types
    • entity and biological-assembly assignments
    • protein–nucleic-acid contact counts
    • protein–protein interface counts
    • selected chains and selection reasons
    • output verification statistics
  • Refuses to overwrite existing outputs unless explicitly requested

Requirements

  • Python 3.10 or newer
  • Biopython 1.84 or newer

Install the dependency with:

pip install biopython

Basic usage

python MolStrctExtractor.py ./complexes -o ./extracted-proteins

If --output-dir is omitted, results are written to:

INPUT_DIR/extracted-proteins

For an input file named:

8TG4.cif

the default output name is:

8tg4protein.pdb

An extraction_report.json file is also generated in the output directory.

Automatic selection logic

For each structure, MolStrctExtractor:

  1. Reads polymer and biological-assembly metadata from mmCIF when available.
  2. Detects protein and nucleic-acid chains.
  3. Counts residue-level contacts using a configurable heavy-atom distance cutoff.
  4. Identifies the protein chain with the strongest nucleic-acid interface.
  5. Selects the biological assembly containing that chain.
  6. Retains one representative of every distinct protein entity in the selected assembly.
  7. Retains multiple copies of the same protein entity only when they form a sufficiently large protein–protein interface.
  8. Writes only amino-acid residues from the selected chains.
  9. Re-parses the output PDB to verify its chains and residue types.

This distinction is important for complexes containing several different protein components. For example, an antibody heavy chain, antibody light chain, and an RNA-binding protein are retained if they belong to the selected assembly, even when only one of them directly contacts RNA.

Explicit chain overrides

Automatic biological-target selection cannot resolve every unusual or experimentally engineered complex. Overrides can be supplied as JSON using author chain IDs.

One target containing multiple chains

{
  "6CF2": ["A", "B", "F"]
}

This produces one protein structure containing chains A, B, and F.

Multiple independent targets from one complex

{
  "1URN": [
    ["A"],
    ["B"]
  ]
}

This produces:

1urnprotein-a.pdb
1urnprotein-b.pdb

Run with an override file:

python MolStrctExtractor.py ./complexes \
  -o ./extracted-proteins \
  --overrides extraction_overrides.json

Command-line options

usage: MolStrctExtractor.py [-h]
                            [-o OUTPUT_DIR]
                            [--overrides OVERRIDES]
                            [--contact-cutoff CONTACT_CUTOFF]
                            [--min-interface-pairs MIN_INTERFACE_PAIRS]
                            [--keep-all-proteins]
                            [--overwrite]
                            input_dir

Important options

  • --overrides PATH
    JSON file defining explicit author-chain selections.

  • --contact-cutoff FLOAT
    Heavy-atom distance cutoff for residue-level contacts. Default: 5.0 Å.

  • --min-interface-pairs INTEGER
    Minimum number of protein residue pairs required to retain interacting copies of the same protein entity. Default: 10.

  • --keep-all-proteins
    Keeps every detected protein chain without automatic assembly selection or symmetry reduction.

  • --overwrite
    Replaces outputs generated by an earlier run.

Input and output formats

Supported input formats:

  • .cif
  • .mmcif
  • .pdb
  • .ent

Output format:

  • Protein-only .pdb

The program processes structure files directly inside the input folder. Directory traversal is not recursive.

Author and label chain IDs

mmCIF files distinguish between label chain IDs and author chain IDs.

MolStrctExtractor:

  • uses mmCIF label IDs internally for entity and biological-assembly metadata
  • writes author chain IDs to the output PDB
  • expects author chain IDs in override files

Because the legacy PDB format supports only one-character chain IDs, structures with multi-character author chain IDs currently require conversion or code adaptation.

Missing residues and chain discontinuities

MolStrctExtractor extracts experimentally modeled coordinates only.

It does not:

  • reconstruct missing residues
  • repair unresolved loops
  • join discontinuous coordinate fragments
  • predict missing RNA or protein structure
  • expand symmetry-generated coordinates that are absent from the input model

Therefore, discontinuities present in the source structure remain present in the extracted PDB.

Audit report

The generated extraction_report.json records the complete selection process for each output, including:

{
  "input_file": "complexes/6CF2.cif",
  "output_file": "extracted-proteins/6cf2protein.pdb",
  "selection": {
    "selected_chain_ids": ["A", "B", "F"],
    "primary_chain_id": "F",
    "method": "auto",
    "reason": "...",
    "warnings": []
  }
}

Reviewing this report is recommended before using automatically extracted structures in a benchmark.

Testing

Run the test suite with:

python -m unittest discover -v

The tests cover:

  • nucleic-acid contact ranking
  • biological-assembly filtering
  • retention of distinct protein entities
  • symmetry-copy reduction
  • retention of interacting same-entity subunits
  • exclusion of water and ligands
  • explicit chain overrides
  • multiple independent targets
  • output naming

Limitations

Automatic selection is a heuristic and cannot determine biological intent in every deposited structure.

Manual review or explicit overrides are recommended when:

  • multiple alternative biological assemblies are present
  • several non-equivalent protein–RNA interfaces occur in one structure
  • the deposited assembly contains crystallization chaperones
  • different copies represent scientifically relevant conformational states
  • biological-assembly metadata are absent or unreliable
  • the input is an old PDB file without detailed entity metadata

mmCIF is preferred over PDB because it provides explicit polymer, entity, author-chain, label-chain, and biological-assembly annotations.

License

MIT License

About

Batch extraction of protein-only targets from protein–nucleic-acid PDB/mmCIF complexes.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages