Interface to RDKit#
RDKit is a popular cheminformatics package
and thus can be used to supplement Biotite with a variety of functionalities focused
on small molecules, such as conversion from/to textual representations
(e.g. SMILES and InChI) and visualization as structural formulas.
Basically, the biotite.interface.rdkit subpackage provides only two functions:
to_mol() to obtain a rdkit.Chem.rdchem.Mol from an AtomArray
and from_mol() for the reverse direction.
The rest happens within the realm of RDKit.
This tutorial will only give a small glance on how the interface can be used.
For comprehensive documentation refer to the
RDKit documentation.
First example: Depiction as structural formula#
RDKit allows rendering structural formulas using pillow. For a proper structural formula, we need to compute proper 2D coordinates first.
import biotite.interface.rdkit as rdkit_interface
import biotite.structure.info as struc
import rdkit.Chem.AllChem as Chem
from rdkit.Chem.Draw import MolToImage
penicillin = struc.residue("PNN")
mol = rdkit_interface.to_mol(penicillin)
# We do not want to include explicit hydrogen atoms in the structural formula
mol = Chem.RemoveHs(mol)
Chem.Compute2DCoords(mol)
image = MolToImage(mol, size=(600, 400))
display(image)
Second example: Creating a molecule from SMILES#
Although the Chemical Component Dictionary accessible from
biotite.structure.info already provides all compounds found in the PDB,
there are a myriad of compounds out there that are not part of it.
One way to to obtain them as AtomArray is passing a SMILES string to
RDKit to obtain the topology of the molecule and then computing the coordinates.
ERTAPENEM_SMILES = "C[C@@H]1[C@@H]2[C@H](C(=O)N2C(=C1S[C@H]3C[C@H](NC3)C(=O)NC4=CC=CC(=C4)C(=O)O)C(=O)O)[C@@H](C)O"
mol = Chem.MolFromSmiles(ERTAPENEM_SMILES)
# RDKit uses implicit hydrogen atoms by default, but Biotite requires explicit ones
mol = Chem.AddHs(mol)
# Create a 3D conformer
conformer_id = Chem.EmbedMolecule(mol)
Chem.UFFOptimizeMolecule(mol)
ertapenem = rdkit_interface.from_mol(mol, conformer_id)
print(ertapenem)
0 C -3.567 1.554 1.942
0 C -3.312 0.902 0.579
0 C -4.565 0.266 -0.110
0 C -5.832 0.053 0.579
0 C -5.541 -1.418 0.535
0 O -6.133 -2.459 0.923
0 N -4.218 -1.124 -0.061
0 C -2.903 -1.411 0.331
0 C -2.332 -0.264 0.700
0 S -0.648 -0.077 1.307
0 C 0.044 0.432 -0.311
0 C 1.462 0.929 -0.172
0 C 2.070 0.659 -1.544
0 N 1.314 -0.439 -2.163
0 C 0.174 -0.724 -1.289
0 C 3.530 0.313 -1.433
0 O 3.906 -0.885 -1.565
0 N 4.487 1.351 -1.171
0 C 5.843 1.113 -0.755
0 C 6.847 2.009 -1.145
0 C 8.167 1.813 -0.735
0 C 8.496 0.729 0.082
0 C 7.500 -0.169 0.506
0 C 6.173 0.035 0.088
0 C 7.840 -1.310 1.386
0 O 9.025 -1.481 1.782
0 O 6.852 -2.207 1.786
0 C -2.359 -2.777 0.417
0 O -3.036 -3.741 -0.029
0 O -1.119 -3.021 0.998
0 C -7.101 0.489 -0.206
0 C -7.363 -0.280 -1.515
0 O -8.230 0.409 0.626
0 H -3.855 0.794 2.699
0 H -2.644 2.061 2.293
0 H -4.367 2.319 1.854
0 H -2.881 1.667 -0.102
0 H -4.655 0.601 -1.167
0 H -5.861 0.411 1.627
0 H -0.576 1.238 -0.763
0 H 1.994 0.349 0.617
0 H 1.490 2.011 0.084
0 H 1.969 1.580 -2.166
0 H 0.949 -0.114 -3.090
0 H 0.383 -1.670 -0.746
0 H -0.763 -0.852 -1.876
0 H 4.195 2.339 -1.348
0 H 6.607 2.855 -1.777
0 H 8.938 2.505 -1.049
0 H 9.526 0.599 0.388
0 H 5.393 -0.627 0.441
0 H 7.074 -2.989 2.390
0 H -0.750 -3.962 1.061
0 H -6.969 1.560 -0.473
0 H -6.509 -0.190 -2.216
0 H -8.261 0.141 -2.015
0 H -7.554 -1.355 -1.313
0 H -8.342 -0.543 0.884