DEVELOPMENT OF A MACHINE LEARNING MODEL FOR PREDICTING MOLECULAR PROPERTIES OF DRUG CANDIDATES FOR PERSONALIZED MEDICINE
Keywords:
Machine Learning; Molecular Screening; Drug-Likeness; Molecular Descriptors; Biomol Predictor.Abstract
The challenge of picking molecular candidates that offer attractive biophysochemical properties, including favourable pharmacokinetic and safety characteristics, continues to be one of the primary reasons for candidates being dropped from the race in the drug development process, with a vast majority failing to pass the stage into clinical testing. Using machine learning (ML) for screening molecules to complement experimental screening is both highly efficient and low cost, because it relies on statistical relationships between the molecular descriptors and the pharmacologically relevant endpoints. In the current work, we explained the rationale of design of a free, web-based biochemical screening tool named BioMol Predictor designed by the undergraduates (UG) of Chemistry department, Thal University Bhakkar, that helps predict biochemical factors including drug-likeness, toxicity risk, absorption, synthetic accessibility, bioactivity and environmental-remediation potential of any molecule. To test the feasibility of the descriptor-based classification method that is the basis of tools of this type, we chose a curated set of 71 well characterized and characterized pharmaceutical products, including 47 small molecule orally administered drugs and 24 biodegradable (non-drug) industrial chemicals, macromolecules and biologics (not orally administered). Stratified 5-fold cross validation (CV) approach was used to train the three classifiers (Random Forest, Logistic Regression, and Support Vector Machines) for assessment, and to compare with classical rule of 5 filter – Lipinski Rule. The best performance of the models is given by the Random Forest model (accuracy = 0.90, F1-score = 0.93, ROC-AUC = 0.94), better than the linear baseline, kernel baseline and rule background (accuracy = 0.86). The best descriptors were topological polar surface area and molecular weight. These findings align with the principle of the other ML screening tools based on descriptors (e.g. BioMol Predictor) but suggest that more data, of the magnitude necessary to validate, and other classes of descriptors are necessary for their integration to replace laboratory confirmation. We end our discussion with the use of the tool in teaching and what are some shortcomings and plans for scaling up the tool and applying it more rigorously. Molecular chemistry rarely advances in a linear fashion, requiring instead a process of trial and error to develop increasingly sophisticated models addressing specific objectives.Molecular chemistry does not proceed in a linear way alone, and it is essential to have a process of trial and error, to find models that are becoming more and more complex and that meet increasingly specific objectives.


