A SAR and QSAR study on cyclin dependent kinase 4 inhibitors using machine learning methods

Xiaoyang Pang; Yunyang Zhao; Guo Li; Jianrong Liu; Aixia Yan

doi:10.1039/D2DD00143H

A SAR and QSAR study on cyclin dependent kinase 4 inhibitors using machine learning methods†

Xiaoyang Pang,^a Yunyang Zhao,^a Guo Li,^a Jianrong Liu*^b and Aixia Yan

*^a

Author affiliations

* Corresponding authors

^a State Key Laboratory of Chemical Resource Engineering, Department of Pharmaceutical Engineering, Beijing University of Chemical Technology, Beijing, P. R. China
E-mail: yanax@mail.buct.edu.cn

^b BUCT-Paris Curie Engineer School, Beijing University of Chemical Technology, Beijing 100029, P. R. China
E-mail: liujianrong2017@126.com

Abstract

Cyclin dependent kinase 4 (CDK4) is a promising target for cancer treatment, and developing new effective CDK4 inhibitors is of great significance in anticancer therapy. In this study, we conducted a structure activity relationship (SAR) study on 3018 CDK4 inhibitors. We applied four machine learning methods, which were Multiple Linear Regression (MLR), Random Forest (RF), Support Vector Machine (SVM) and Deep Neural Network (DNN), to develop 18 classification models based on 3018 inhibitors (dataset 1), 18 classification models based on dataset 1 and decoys, and 24 quantitative structure–activity relationship (QSAR) models based on 1427 inhibitors (dataset 2). We obtained some optimal models. Based on dataset 1, Model A2, built by SVM and MACCS fingerprints, has a prediction accuracy (Q) of 92.68% and a Matthews correlation coefficient (MCC) of 0.874 for the test set. Based on dataset 1 and decoys, Model C2, built by SVM and MACCS fingerprints, has a Q of 98.5% and a MCC of 0.937 for the test set. Based on dataset 2, Model F7, built by SVM and MOE descriptors, has a coefficient of determination (R²) of 0.824 and a root mean squared error (RMSE) of 0.534 for the test set. For classification models, it was found that the more samples used for modelling, the more robust the models, and the better the performance of the models. Moreover, we clustered 3018 inhibitors into 12 subsets, and analysed their scaffolds and fragment features. It was found that 2-aminopyrimidine, pyridine, piperazine and cyclopentane were common scaffolds and fragments in highly active inhibitors. This study can provide guidance for the discovery and optimization of CDK4 inhibitor lead compounds.

Supplementary files

Article information

DOI: https://doi.org/10.1039/D2DD00143H
Article type: Paper
Submitted: 18 Dec 2022
Accepted: 29 May 2023
First published: 31 May 2023
This article is Open Access

Download Citation

Digital Discovery, 2023,2, 1026-1041

Permissions

Request permissions

A SAR and QSAR study on cyclin dependent kinase 4 inhibitors using machine learning methods

X. Pang, Y. Zhao, G. Li, J. Liu and A. Yan, Digital Discovery, 2023, 2, 1026 DOI: 10.1039/D2DD00143H

This article is licensed under a Creative Commons Attribution-NonCommercial 3.0 Unported Licence. You can use material from this article in other publications, without requesting further permission from the RSC, provided that the correct acknowledgement is given and it is not used for commercial purposes.

To request permission to reproduce material from this article in a commercial publication, please go to the Copyright Clearance Center request page.

If you are an author contributing to an RSC publication, you do not need to request permission provided correct acknowledgement is given.

If you are the author of this article, you do not need to request permission to reproduce figures and diagrams provided correct acknowledgement is given. If you want to reproduce the whole article in a third-party commercial publication (excluding your thesis/dissertation for which permission is not required) please go to the Copyright Clearance Center request page.

Digital Discovery

A SAR and QSAR study on cyclin dependent kinase 4 inhibitors using machine learning methods†

Abstract

Supplementary files

Article information

Download Citation

Permissions

A SAR and QSAR study on cyclin dependent kinase 4 inhibitors using machine learning methods

Social activity

Search articles by author

Spotlight

Advertisements