In ProgressCourse work · Bioinformatics Algorithms

CrypTree: Phylogeny-Guided Analysis of Cryptic Binding Site Conservation and Its Predictive Limits

Jahid Hasan, Dr. Md. Shamsuzzoha Bayzid

Abstract

Cryptic binding sites are hidden pockets in a protein. They stay closed in the protein normally but open only when a drug or another molecule binds with it. This makes them useful to discover drug. It also makes hard to find them as well, because they cannot be seen in the closed structure. Predicting them from sequence is therefore the practical option. Existing methods look at one protein at a time and ignore its evolutionary history. We tested whether evolutionary history helps or not. We used CryptoBench’s Dataset, which holds 1,274 cryptic protein chains which are 373,870 residues, of which 5.4% are cryptic residue. For each chain we collected homologs from UniRef50, built a maximum-likelihood tree, and measured how fast each site evolves. Cryptic residues change about 32% more slowly than other residues. This holds in 115 of 140 families, and again, at 32%, when each of 1,242 chains is analysed on its own. But slow evolution is a weak signal for a single residue. A classifier built on these rates reaches an AUPRC of 0.1375 on 251 of 261 test chains (72,636 residues), against 0.054 for random guessing. The tree-based measure is no better than conservation measured without a tree. The ancestral-state model could be fitted on only 9 of 140 families. Fifteen times more homologs gave fewer labelled chains per family, 2.69 instead of 3.06, not more. CryptoBench is built to be non-redundant, spanning 968 UniProt identifiers, so methods that need related labelled proteins cannot work on it.