An Integrative Approach to Drug Development Using Machine Learning
Despite recent advances in life sciences and technology, the amount of time and money spent in the drug development process remain drastically inflated. Thus, there is a need to rapidly recognize characteristics that will help identify novel therapies. First, we address the increased need for drug repurposing, the approach of identifying new indications for approved or investigational drugs. We present a novel drug repurposing method called Creating A Translational Network for Indication Prediction (CATNIP) which relies solely on biological and chemical drug characteristics to identify disease areas for specific drugs and drug classes. This drug-focused approach could allow our approach to be used for both FDA approved drugs as well as investigational drugs. Our method, trained with 2,576 diverse small molecules, is built using easily interpretable features, such as chemical structure and targets, allowing for probable drug-disease mechanisms to be discovered from the predictions made. The strength of this method’s approach is demonstrated through a repurposing network that can be utilized identify drug class candidate repurposing opportunities. In order to treat many of these indications, a drug compound is orally ingested by a patient. One of the major absorption sites for drugs is the small intestine, and drug properties such as permeability are proven important to maximize treatment efforts. Poor absorption of drug candidates is likely to lead to failure in the drug development process, so we propose an innovative approach to predict the permeability of a drug. The Caco-2 cell model is a standard surrogate for predicting in vitro intestinal permeability. We collected one of the largest experimentally based datasets of Caco-2 values to create a computational model. Using an approach called graph convolutional networks that treats molecules as graphs, we are able to take in a line-notation form molecular structure and successfully make predictions about a drug compound’s permeability. Altogether, this work demonstrates how the integration of diverse datasets can aid in addressing the multitude of challenging problems in the field of drug discovery. Computational approaches such as these, that prioritize applicability and interpretability, have the strong potential to transform and improve upon the drug development pipeline.