Int J Biol Macromol. 2026 Jun 21:153150. doi: 10.1016/j.ijbiomac.2026.153150. Online ahead of print.

ABSTRACT

Cytochrome P450 (CYP450) proteins are a vital enzyme superfamily involved in drug metabolism, detoxification, and biosynthesis. Accurate identification of CYP450 proteins in large-scale proteomic datasets is crucial for advancing pharmacogenomics and industrial biocatalysis. However, current methods are limited by the high sequence diversity and low homology across CYP450 subfamilies, which hampers the detection of distant homologs. To address these challenges, a deep learning framework was developed that integrates multiple feature types, including protein sequence information, semantic embeddings, and evolutionary conservation signals. A comprehensive analysis of over 1000 model combinations demonstrated that optimal combinations of modalities provide complementary information, while suboptimal combinations resulted in weaker performance. The best-performing model, which integrates semantic embeddings from pre-trained protein language models (PLMs), achieved an accuracy exceeding 95% and an area under the receiver operating characteristic curve (auROC) of 0.970 on both training and internal test sets, confirming the efficacy of the multimodal fusion strategy. Multimodal interpretability analyses further elucidated the relative importance of features and their interactions, offering valuable insights into the model’s decision-making process. This approach outperforms traditional machine learning methods, providing a robust and accurate solution for CYP450 protein identification, with implications for enzyme engineering and drug discovery.

PMID:42324002 | DOI:10.1016/j.ijbiomac.2026.153150