MolStrctExtractor is a lightweight Python tool for extracting protein-only target structures from protein–nucleic-acid complexes stored as PDB or mmCIF files.
It is designed for benchmark preparation and structure-based protein/RNA modeling workflows where the input structures contain RNA, DNA, proteins, ligands, ions, water, or repeated crystallographic copies, but the downstream task requires protein-only PDB files.
- Batch-processes all PDB/mmCIF structures directly under one folder
- Identifies protein and nucleic-acid polymer chains
- Uses protein–nucleic-acid contacts to identify the relevant biological assembly
- Retains one representative of every distinct protein entity in that assembly
- Preserves interacting copies of the same protein when they form a substantial protein–protein interface
- Removes RNA, DNA, water, ions, and non-protein ligands
- Supports multiple independent protein targets from one complex
- Supports explicit author-chain overrides for unusual complexes
- Generates an auditable JSON report containing:
- detected chains and polymer types
- entity and biological-assembly assignments
- protein–nucleic-acid contact counts
- protein–protein interface counts
- selected chains and selection reasons
- output verification statistics
- Refuses to overwrite existing outputs unless explicitly requested
- Python 3.10 or newer
- Biopython 1.84 or newer
Install the dependency with:
pip install biopythonpython MolStrctExtractor.py ./complexes -o ./extracted-proteinsIf --output-dir is omitted, results are written to:
INPUT_DIR/extracted-proteins
For an input file named:
8TG4.cif
the default output name is:
8tg4protein.pdb
An extraction_report.json file is also generated in the output directory.
For each structure, MolStrctExtractor:
- Reads polymer and biological-assembly metadata from mmCIF when available.
- Detects protein and nucleic-acid chains.
- Counts residue-level contacts using a configurable heavy-atom distance cutoff.
- Identifies the protein chain with the strongest nucleic-acid interface.
- Selects the biological assembly containing that chain.
- Retains one representative of every distinct protein entity in the selected assembly.
- Retains multiple copies of the same protein entity only when they form a sufficiently large protein–protein interface.
- Writes only amino-acid residues from the selected chains.
- Re-parses the output PDB to verify its chains and residue types.
This distinction is important for complexes containing several different protein components. For example, an antibody heavy chain, antibody light chain, and an RNA-binding protein are retained if they belong to the selected assembly, even when only one of them directly contacts RNA.
Automatic biological-target selection cannot resolve every unusual or experimentally engineered complex. Overrides can be supplied as JSON using author chain IDs.
{
"6CF2": ["A", "B", "F"]
}This produces one protein structure containing chains A, B, and F.
{
"1URN": [
["A"],
["B"]
]
}This produces:
1urnprotein-a.pdb
1urnprotein-b.pdb
Run with an override file:
python MolStrctExtractor.py ./complexes \
-o ./extracted-proteins \
--overrides extraction_overrides.jsonusage: MolStrctExtractor.py [-h]
[-o OUTPUT_DIR]
[--overrides OVERRIDES]
[--contact-cutoff CONTACT_CUTOFF]
[--min-interface-pairs MIN_INTERFACE_PAIRS]
[--keep-all-proteins]
[--overwrite]
input_dir
-
--overrides PATH
JSON file defining explicit author-chain selections. -
--contact-cutoff FLOAT
Heavy-atom distance cutoff for residue-level contacts. Default:5.0 Å. -
--min-interface-pairs INTEGER
Minimum number of protein residue pairs required to retain interacting copies of the same protein entity. Default:10. -
--keep-all-proteins
Keeps every detected protein chain without automatic assembly selection or symmetry reduction. -
--overwrite
Replaces outputs generated by an earlier run.
Supported input formats:
.cif.mmcif.pdb.ent
Output format:
- Protein-only
.pdb
The program processes structure files directly inside the input folder. Directory traversal is not recursive.
mmCIF files distinguish between label chain IDs and author chain IDs.
MolStrctExtractor:
- uses mmCIF label IDs internally for entity and biological-assembly metadata
- writes author chain IDs to the output PDB
- expects author chain IDs in override files
Because the legacy PDB format supports only one-character chain IDs, structures with multi-character author chain IDs currently require conversion or code adaptation.
MolStrctExtractor extracts experimentally modeled coordinates only.
It does not:
- reconstruct missing residues
- repair unresolved loops
- join discontinuous coordinate fragments
- predict missing RNA or protein structure
- expand symmetry-generated coordinates that are absent from the input model
Therefore, discontinuities present in the source structure remain present in the extracted PDB.
The generated extraction_report.json records the complete selection process for each output, including:
{
"input_file": "complexes/6CF2.cif",
"output_file": "extracted-proteins/6cf2protein.pdb",
"selection": {
"selected_chain_ids": ["A", "B", "F"],
"primary_chain_id": "F",
"method": "auto",
"reason": "...",
"warnings": []
}
}Reviewing this report is recommended before using automatically extracted structures in a benchmark.
Run the test suite with:
python -m unittest discover -vThe tests cover:
- nucleic-acid contact ranking
- biological-assembly filtering
- retention of distinct protein entities
- symmetry-copy reduction
- retention of interacting same-entity subunits
- exclusion of water and ligands
- explicit chain overrides
- multiple independent targets
- output naming
Automatic selection is a heuristic and cannot determine biological intent in every deposited structure.
Manual review or explicit overrides are recommended when:
- multiple alternative biological assemblies are present
- several non-equivalent protein–RNA interfaces occur in one structure
- the deposited assembly contains crystallization chaperones
- different copies represent scientifically relevant conformational states
- biological-assembly metadata are absent or unreliable
- the input is an old PDB file without detailed entity metadata
mmCIF is preferred over PDB because it provides explicit polymer, entity, author-chain, label-chain, and biological-assembly annotations.
MIT License