1
|
Sheng N, Huang L, Lu Y, Wang H, Yang L, Gao L, Xie X, Fu Y, Wang Y. Data resources and computational methods for lncRNA-disease association prediction. Comput Biol Med 2023; 153:106527. [PMID: 36610216 DOI: 10.1016/j.compbiomed.2022.106527] [Citation(s) in RCA: 5] [Impact Index Per Article: 5.0] [Reference Citation Analysis] [Abstract] [Key Words] [MESH Headings] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Received: 09/20/2022] [Revised: 12/08/2022] [Accepted: 12/31/2022] [Indexed: 01/03/2023]
Abstract
Increasing interest has been attracted in deciphering the potential disease pathogenesis through lncRNA-disease association (LDA) prediction, regarding to the diverse functional roles of lncRNAs in genome regulation. Whilst, computational models and algorithms benefit systematic biology research, even facilitate the classical biological experimental procedures. In this review, we introduce representative diseases associated with lncRNAs, such as cancers, cardiovascular diseases, and neurological diseases. Current publicly available resources related to lncRNAs and diseases have also been included. Furthermore, all of the 64 computational methods for LDA prediction have been divided into 5 groups, including machine learning-based methods, network propagation-based methods, matrix factorization- and completion-based methods, deep learning-based methods, and graph neural network-based methods. The common evaluation methods and metrics in LDA prediction have also been discussed. Finally, the challenges and future trends in LDA prediction have been discussed. Recent advances in LDA prediction approaches have been summarized in the GitHub repository at https://github.com/sheng-n/lncRNA-disease-methods.
Collapse
Affiliation(s)
- Nan Sheng
- Key Laboratory of Symbol Computation and Knowledge Engineering of Ministry of Education, College of Computer Science and Technology, Jilin University, Changchun, China
| | - Lan Huang
- Key Laboratory of Symbol Computation and Knowledge Engineering of Ministry of Education, College of Computer Science and Technology, Jilin University, Changchun, China.
| | - Yuting Lu
- School of Artificial Intelligence, Jilin University, Changchun, China
| | - Hao Wang
- Department of Hepatopancreatobiliary Surgery, Second Affiliated Hospital of Harbin Medical University, Harbin, China
| | - Lili Yang
- Key Laboratory of Symbol Computation and Knowledge Engineering of Ministry of Education, College of Computer Science and Technology, Jilin University, Changchun, China; Department of Obstetrics, The First Hospital of Jilin University, Changchun, China
| | - Ling Gao
- Key Laboratory of Symbol Computation and Knowledge Engineering of Ministry of Education, College of Computer Science and Technology, Jilin University, Changchun, China
| | - Xuping Xie
- Key Laboratory of Symbol Computation and Knowledge Engineering of Ministry of Education, College of Computer Science and Technology, Jilin University, Changchun, China
| | - Yuan Fu
- Institute of Biological, Environmental and Rural Sciences, Aberystwyth University, Aberystwyth, Ceredigion, United Kingdom
| | - Yan Wang
- Key Laboratory of Symbol Computation and Knowledge Engineering of Ministry of Education, College of Computer Science and Technology, Jilin University, Changchun, China; School of Artificial Intelligence, Jilin University, Changchun, China.
| |
Collapse
|
2
|
Xie GB, Chen RB, Lin ZY, Gu GS, Yu JR, Liu ZG, Cui J, Lin LQ, Chen LC. Predicting lncRNA-disease associations based on combining selective similarity matrix fusion and bidirectional linear neighborhood label propagation. Brief Bioinform 2023; 24:6966536. [PMID: 36592062 DOI: 10.1093/bib/bbac595] [Citation(s) in RCA: 7] [Impact Index Per Article: 7.0] [Reference Citation Analysis] [Abstract] [Key Words] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Received: 10/31/2022] [Revised: 11/30/2022] [Accepted: 12/04/2022] [Indexed: 01/03/2023] Open
Abstract
Recent studies have revealed that long noncoding RNAs (lncRNAs) are closely linked to several human diseases, providing new opportunities for their use in detection and therapy. Many graph propagation and similarity fusion approaches can be used for predicting potential lncRNA-disease associations. However, existing similarity fusion approaches suffer from noise and self-similarity loss in the fusion process. To address these problems, a new prediction approach, termed SSMF-BLNP, based on organically combining selective similarity matrix fusion (SSMF) and bidirectional linear neighborhood label propagation (BLNP), is proposed in this paper to predict lncRNA-disease associations. In SSMF, self-similarity networks of lncRNAs and diseases are obtained by selective preprocessing and nonlinear iterative fusion. The fusion process assigns weights to each initial similarity network and introduces a unit matrix that can reduce noise and compensate for the loss of self-similarity. In BLNP, the initial lncRNA-disease associations are employed in both lncRNA and disease directions as label information for linear neighborhood label propagation. The propagation was then performed on the self-similarity network obtained from SSMF to derive the scoring matrix for predicting the relationships between lncRNAs and diseases. Experimental results showed that SSMF-BLNP performed better than seven other state of-the-art approaches. Furthermore, a case study demonstrated up to 100% and 80% accuracy in 10 lncRNAs associated with hepatocellular carcinoma and 10 lncRNAs associated with renal cell carcinoma, respectively. The source code and datasets used in this paper are available at: https://github.com/RuiBingo/SSMF-BLNP.
Collapse
Affiliation(s)
- Guo-Bo Xie
- School of Computer, Guangdong University of Technology, Guangzhou, 510000, China
| | - Rui-Bin Chen
- School of Computer, Guangdong University of Technology, Guangzhou, 510000, China
| | - Zhi-Yi Lin
- School of Computer, Guangdong University of Technology, Guangzhou, 510000, China
| | - Guo-Sheng Gu
- School of Computer, Guangdong University of Technology, Guangzhou, 510000, China
| | - Jun-Rui Yu
- School of Computer, Guangdong University of Technology, Guangzhou, 510000, China
| | - Zhen-Guo Liu
- Department of Thoracic Surgery, The First Affiliated Hospital of Sun Yat-sen University, Guangzhou, 510080, China
| | - Ji Cui
- Department of Gastrointestinal Surgery, The First Affiliated Hospital of Sun Yat-sen University, Guangzhou, 510080, China
| | - Lie-Qing Lin
- Center of Campus Network & Modern Educational Technology, Guangdong University of Technology, Guangzhou, 510000, China
| | - Lang-Cheng Chen
- Center of Campus Network & Modern Educational Technology, Guangdong University of Technology, Guangzhou, 510000, China
| |
Collapse
|
3
|
Hyperbolic matrix factorization improves prediction of drug-target associations. Sci Rep 2023; 13:959. [PMID: 36653463 PMCID: PMC9849222 DOI: 10.1038/s41598-023-27995-5] [Citation(s) in RCA: 2] [Impact Index Per Article: 2.0] [Reference Citation Analysis] [Abstract] [Key Words] [Track Full Text] [Figures] [Journal Information] [Subscribe] [Scholar Register] [Received: 07/24/2022] [Accepted: 01/11/2023] [Indexed: 01/19/2023] Open
Abstract
Past research in computational systems biology has focused more on the development and applications of advanced statistical and numerical optimization techniques and much less on understanding the geometry of the biological space. By representing biological entities as points in a low dimensional Euclidean space, state-of-the-art methods for drug-target interaction (DTI) prediction implicitly assume the flat geometry of the biological space. In contrast, recent theoretical studies suggest that biological systems exhibit tree-like topology with a high degree of clustering. As a consequence, embedding a biological system in a flat space leads to distortion of distances between biological objects. Here, we present a novel matrix factorization methodology for drug-target interaction prediction that uses hyperbolic space as the latent biological space. When benchmarked against classical, Euclidean methods, hyperbolic matrix factorization exhibits superior accuracy while lowering embedding dimension by an order of magnitude. We see this as additional evidence that the hyperbolic geometry underpins large biological networks.
Collapse
|
4
|
Li J, Wang D, Yang Z, Liu M. HEGANLDA: A Computational Model for Predicting Potential Lncrna-Disease Associations Based On Multiple Heterogeneous Networks. IEEE/ACM TRANSACTIONS ON COMPUTATIONAL BIOLOGY AND BIOINFORMATICS 2023; 20:388-398. [PMID: 34932483 DOI: 10.1109/tcbb.2021.3136886] [Citation(s) in RCA: 1] [Impact Index Per Article: 1.0] [Reference Citation Analysis] [Abstract] [MESH Headings] [Track Full Text] [Subscribe] [Scholar Register] [Indexed: 05/04/2023]
Abstract
Long non-coding RNAs (lncRNAs) play vital regulatory roles in many human complex diseases, however, the number of validated lncRNA-disease associations is notable rare so far. How to predict potential lncRNA-disease associations precisely through computational methods remains challenging. In this study, we proposed a novel method, LDVCHN (LncRNA-Disease Vector Calculation Heterogeneous Networks), and also developed the corresponding model, HEGANLDA (Heterogeneous Embedding Generative Adversarial Networks LncRNA-Disease Association), for predicting potential lncRNA-disease associations. In HEGANLDA, the graph embedding algorithm (HeGAN) was introduced for mapping all nodes in the lncRNA-miRNA-disease heterogeneous network into the low-dimensional vectors which severed as the inputs of LDVCHN. HEGANLDA effectively adopted the XGBoost (eXtreme Gradient Boosting) classifier, which was trained by the low-dimensional vectors, to predict potential lncRNA-disease associations. The 10-fold cross-validation method was utilized to evaluate the performance of our model, our model finally achieved an area under the ROC curve of 0.983. According to the experiment results, HEGANLDA outperformed any one of five current state-of-the-art methods. To further evaluate the effectiveness of HEGANLDA in predicting potential lncRNA-disease associations, both case studies and robustness tests were performed and the results confirmed its effectiveness and robustness. The source code and data of HEGANLDA are available at https://github.com/HEGANLDA/HEGANLDA.
Collapse
|
5
|
Li J, Yang X, Guan Y, Pan Z. Prediction of Drug–Target Interaction Using Dual-Network Integrated Logistic Matrix Factorization and Knowledge Graph Embedding. Molecules 2022; 27:molecules27165131. [PMID: 36014371 PMCID: PMC9412517 DOI: 10.3390/molecules27165131] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [MESH Headings] [Track Full Text] [Download PDF] [Figures] [Journal Information] [Subscribe] [Scholar Register] [Received: 06/25/2022] [Revised: 08/02/2022] [Accepted: 08/05/2022] [Indexed: 11/16/2022] Open
Abstract
Nowadays, drug–target interactions (DTIs) prediction is a fundamental part of drug repositioning. However, on the one hand, drug–target interactions prediction models usually consider drugs or targets information, which ignore prior knowledge between drugs and targets. On the other hand, models incorporating priori knowledge cannot make interactions prediction for under-studied drugs and targets. Hence, this article proposes a novel dual-network integrated logistic matrix factorization DTIs prediction scheme (Ro-DNILMF) via a knowledge graph embedding approach. This model adds prior knowledge as input data into the prediction model and inherits the advantages of the DNILMF model, which can predict under-studied drug–target interactions. Firstly, a knowledge graph embedding model based on relational rotation (RotatE) is trained to construct the interaction adjacency matrix and integrate prior knowledge. Secondly, a dual-network integrated logistic matrix factorization prediction model (DNILMF) is used to predict new drugs and targets. Finally, several experiments conducted on the public datasets are used to demonstrate that the proposed method outperforms the single base-line model and some mainstream methods on efficiency.
Collapse
Affiliation(s)
- Jiaxin Li
- College of Computer Science & Technology, Qingdao University, Qingdao 266071, China
| | - Xixin Yang
- College of Computer Science & Technology, Qingdao University, Qingdao 266071, China
- School of Automation, Qingdao University, Qingdao 266017, China
- Correspondence:
| | - Yuanlin Guan
- Key Lab of Industrial Fluid Energy Conservation and Pollution Control, Ministry of Education, Qingdao University of Technology, Qingdao 266520, China
- School of Mechanical & Automotive Engineering, Qingdao University of Technology, Qingdao 266520, China
| | - Zhenkuan Pan
- College of Computer Science & Technology, Qingdao University, Qingdao 266071, China
| |
Collapse
|
6
|
Liu Y, Yang H, Zheng C, Wang K, Yan J, Cao H, Zhang Y. NCP-BiRW: A Hybrid Approach for Predicting Long Noncoding RNA-Disease Associations by Network Consistency Projection and Bi-Random Walk. Front Genet 2022; 13:862272. [PMID: 35495166 PMCID: PMC9043107 DOI: 10.3389/fgene.2022.862272] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Track Full Text] [Download PDF] [Figures] [Journal Information] [Subscribe] [Scholar Register] [Received: 01/25/2022] [Accepted: 03/21/2022] [Indexed: 12/06/2022] Open
Abstract
Long non-coding RNAs (lncRNAs) play significant roles in the disease process. Understanding the pathological mechanisms of lncRNAs during the course of various diseases will help clinicians prevent and treat diseases. With the emergence of high-throughput techniques, many biological experiments have been developed to study lncRNA-disease associations. Because experimental methods are costly, slow, and laborious, a growing number of computational models have emerged. Here, we present a new approach using network consistency projection and bi-random walk (NCP-BiRW) to infer hidden lncRNA-disease associations. First, integrated similarity networks for lncRNAs and diseases were constructed by merging similarity information. Subsequently, network consistency projection was applied to calculate space projection scores for lncRNAs and diseases, which were then introduced into a bi-random walk method for association prediction. To test model performance, we employed 5- and 10-fold cross-validation, with the area under the receiver operating characteristic curve as the evaluation indicator. The computational results showed that our method outperformed the other five advanced algorithms. In addition, the novel method was applied to another dataset in the Mammalian ncRNA-Disease Repository (MNDR) database and showed excellent performance. Finally, case studies were carried out on atherosclerosis and leukemia to confirm the effectiveness of our method in practice. In conclusion, we could infer lncRNA-disease associations using the NCP-BiRW model, which may benefit biomedical studies in the future.
Collapse
Affiliation(s)
- Yanling Liu
- Department of Health Statistics, School of Public Health, Shanxi Medical University, Taiyuan, China
- Department of Mathematics, Changzhi Medical College, Changzhi, China
| | - Hong Yang
- Department of Health Statistics, School of Public Health, Shanxi Medical University, Taiyuan, China
| | - Chu Zheng
- Department of Health Statistics, School of Public Health, Shanxi Medical University, Taiyuan, China
| | - Ke Wang
- Department of Health Statistics, School of Public Health, Shanxi Medical University, Taiyuan, China
| | - Jingjing Yan
- Department of Health Statistics, School of Public Health, Shanxi Medical University, Taiyuan, China
| | - Hongyan Cao
- Department of Health Statistics, School of Public Health, Shanxi Medical University, Taiyuan, China
| | - Yanbo Zhang
- Department of Health Statistics, School of Public Health, Shanxi Medical University, Taiyuan, China
- Shanxi Provincial Key Laboratory of Major Diseases Risk Assessment, Taiyuan, China
- School of Health and Service Management, Shanxi University of Chinese Medicine, Taiyuan, China
- *Correspondence:Yanbo Zhang,
| |
Collapse
|
7
|
Alshahrani M, Almansour A, Alkhaldi A, Thafar MA, Uludag M, Essack M, Hoehndorf R. Combining biomedical knowledge graphs and text to improve predictions for drug-target interactions and drug-indications. PeerJ 2022; 10:e13061. [PMID: 35402106 PMCID: PMC8988936 DOI: 10.7717/peerj.13061] [Citation(s) in RCA: 5] [Impact Index Per Article: 2.5] [Reference Citation Analysis] [Abstract] [Key Words] [Track Full Text] [Download PDF] [Figures] [Journal Information] [Subscribe] [Scholar Register] [Received: 08/28/2020] [Accepted: 02/13/2022] [Indexed: 01/11/2023] Open
Abstract
Biomedical knowledge is represented in structured databases and published in biomedical literature, and different computational approaches have been developed to exploit each type of information in predictive models. However, the information in structured databases and literature is often complementary. We developed a machine learning method that combines information from literature and databases to predict drug targets and indications. To effectively utilize information in published literature, we integrate knowledge graphs and published literature using named entity recognition and normalization before applying a machine learning model that utilizes the combination of graph and literature. We then use supervised machine learning to show the effects of combining features from biomedical knowledge and published literature on the prediction of drug targets and drug indications. We demonstrate that our approach using datasets for drug-target interactions and drug indications is scalable to large graphs and can be used to improve the ranking of targets and indications by exploiting features from either structure or unstructured information alone.
Collapse
Affiliation(s)
- Mona Alshahrani
- National Center for Artificial Intelligence (NCAI), Saudi Data and Artificial Intelligence Authority (SDAIA), Riyadh, Saudi Arabia
| | - Abdullah Almansour
- National Center for Artificial Intelligence (NCAI), Saudi Data and Artificial Intelligence Authority (SDAIA), Riyadh, Saudi Arabia
| | - Asma Alkhaldi
- National Center for Artificial Intelligence (NCAI), Saudi Data and Artificial Intelligence Authority (SDAIA), Riyadh, Saudi Arabia
| | - Maha A. Thafar
- College of Computers and Information Technology, Taif University, Taif, Saudi Arabia,Computer, Electrical and Mathematical Sciences and Engineering Division (CEMSE), Computational Bioscience Research Center (CBRC), King Abdullah University of Science and Technology (KAUST), King Abdullah University of Science and Technology, Thuwal, Saudi Arabia
| | - Mahmut Uludag
- Computer, Electrical and Mathematical Sciences and Engineering Division (CEMSE), Computational Bioscience Research Center (CBRC), King Abdullah University of Science and Technology (KAUST), King Abdullah University of Science and Technology, Thuwal, Saudi Arabia
| | - Magbubah Essack
- Computer, Electrical and Mathematical Sciences and Engineering Division (CEMSE), Computational Bioscience Research Center (CBRC), King Abdullah University of Science and Technology (KAUST), King Abdullah University of Science and Technology, Thuwal, Saudi Arabia
| | - Robert Hoehndorf
- Computer, Electrical and Mathematical Sciences and Engineering Division (CEMSE), Computational Bioscience Research Center (CBRC), King Abdullah University of Science and Technology (KAUST), King Abdullah University of Science and Technology, Thuwal, Saudi Arabia
| |
Collapse
|
8
|
Xuan P, Gong Z, Cui H, Li B, Zhang T. Fully connected autoencoder and convolutional neural network with attention-based method for inferring disease-related lncRNAs. Brief Bioinform 2022; 23:6561435. [PMID: 35362511 DOI: 10.1093/bib/bbac089] [Citation(s) in RCA: 7] [Impact Index Per Article: 3.5] [Reference Citation Analysis] [Abstract] [Key Words] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Received: 12/06/2021] [Revised: 01/17/2022] [Accepted: 02/23/2022] [Indexed: 11/14/2022] Open
Abstract
Since abnormal expression of long noncoding RNAs (lncRNAs) is often closely related to various human diseases, identification of disease-associated lncRNAs is helpful for exploring the complex pathogenesis. Most of recent methods concentrate on exploiting multiple kinds of data related to lncRNAs and diseases for predicting candidate disease-related lncRNAs. These methods, however, failed to deeply integrate the topology information from the meta-paths that are composed of lncRNA, disease and microRNA (miRNA) nodes. We proposed a new method based on fully connected autoencoders and convolutional neural networks, called ACLDA, for inferring potential disease-related lncRNA candidates. A heterogeneous graph that consists of lncRNA, disease and miRNA nodes were firstly constructed to integrate similarities, associations and interactions among them. Fully connected autoencoder-based module was established to extract the low-dimensional features of lncRNA, disease and miRNA nodes in the heterogeneous graph. We designed the attention mechanisms at the node feature level and at the meta-path level to learn more informative features and meta-paths. A module based on convolutional neural networks was constructed to encode the local topologies of lncRNA and disease nodes from multiple meta-path perspectives. The comprehensive experimental results demonstrated ACLDA achieves superior performance than several state-of-the-art prediction methods. Case studies on breast, lung and colon cancers demonstrated that ACLDA is able to discover the potential disease-related lncRNAs.
Collapse
Affiliation(s)
- Ping Xuan
- School of Computer Science and Technology, Heilongjiang University, Harbin 150080, China
| | - Zhe Gong
- School of Computer Science and Technology, Heilongjiang University, Harbin 150080, China
| | - Hui Cui
- Department of Computer Science and Information Technology, La Trobe University, Melbourne 3083, Australia
| | - Bochong Li
- Center for Frontier Medical Engineering, Chiba University, Chiba 2638522, Japan
| | - Tiangang Zhang
- School of Mathematical Science, Heilongjiang University, Harbin 150080, China
| |
Collapse
|
9
|
Wang L, Zhong C. gGATLDA: lncRNA-disease association prediction based on graph-level graph attention network. BMC Bioinformatics 2022; 23:11. [PMID: 34983363 PMCID: PMC8729153 DOI: 10.1186/s12859-021-04548-z] [Citation(s) in RCA: 13] [Impact Index Per Article: 6.5] [Reference Citation Analysis] [Abstract] [Key Words] [Track Full Text] [Download PDF] [Figures] [Journal Information] [Subscribe] [Scholar Register] [Received: 09/18/2021] [Accepted: 12/21/2021] [Indexed: 01/20/2023] Open
Abstract
Background Long non-coding RNAs (lncRNAs) are related to human diseases by regulating gene expression. Identifying lncRNA-disease associations (LDAs) will contribute to diagnose, treatment, and prognosis of diseases. However, the identification of LDAs by the biological experiments is time-consuming, costly and inefficient. Therefore, the development of efficient and high-accuracy computational methods for predicting LDAs is of great significance. Results In this paper, we propose a novel computational method (gGATLDA) to predict LDAs based on graph-level graph attention network. Firstly, we extract the enclosing subgraphs of each lncRNA-disease pair. Secondly, we construct the feature vectors by integrating lncRNA similarity and disease similarity as node attributes in subgraphs. Finally, we train a graph neural network (GNN) model by feeding the subgraphs and feature vectors to it, and use the trained GNN model to predict lncRNA-disease potential association scores. The experimental results show that our method can achieve higher area under the receiver operation characteristic curve (AUC), area under the precision recall curve (AUPR), accuracy and F1-Score than the state-of-the-art methods in five fold cross-validation. Case studies show that our method can effectively identify lncRNAs associated with breast cancer, gastric cancer, prostate cancer, and renal cancer. Conclusion The experimental results indicate that our method is a useful approach for predicting potential LDAs.
Collapse
Affiliation(s)
- Li Wang
- School of Computer Science and Engineering, South China University of Technology, Guangzhou, China.,School of Computer, Electronics and Information, Guangxi University, Nanning, China
| | - Cheng Zhong
- School of Computer, Electronics and Information, Guangxi University, Nanning, China. .,Key Laboratory of Parallel and Distributed Computing in Guangxi Colleges and Universities, Guangxi University, Nanning, China.
| |
Collapse
|
10
|
Wang L, Shang M, Dai Q, He PA. Prediction of lncRNA-disease association based on a Laplace normalized random walk with restart algorithm on heterogeneous networks. BMC Bioinformatics 2022; 23:5. [PMID: 34983367 PMCID: PMC8729064 DOI: 10.1186/s12859-021-04538-1] [Citation(s) in RCA: 14] [Impact Index Per Article: 7.0] [Reference Citation Analysis] [Abstract] [Key Words] [Track Full Text] [Download PDF] [Figures] [Journal Information] [Subscribe] [Scholar Register] [Received: 10/08/2021] [Accepted: 12/15/2021] [Indexed: 12/23/2022] Open
Abstract
BACKGROUND More and more evidence showed that long non-coding RNAs (lncRNAs) play important roles in the development and progression of human sophisticated diseases. Therefore, predicting human lncRNA-disease associations is a challenging and urgently task in bioinformatics to research of human sophisticated diseases. RESULTS In the work, a global network-based computational framework called as LRWRHLDA were proposed which is a universal network-based method. Firstly, four isomorphic networks include lncRNA similarity network, disease similarity network, gene similarity network and miRNA similarity network were constructed. And then, six heterogeneous networks include known lncRNA-disease, lncRNA-gene, lncRNA-miRNA, disease-gene, disease-miRNA, and gene-miRNA associations network were applied to design a multi-layer network. Finally, the Laplace normalized random walk with restart algorithm in this global network is suggested to predict the relationship between lncRNAs and diseases. CONCLUSIONS The ten-fold cross validation is used to evaluate the performance of LRWRHLDA. As a result, LRWRHLDA achieves an AUC of 0.98402, which is higher than other compared methods. Furthermore, LRWRHLDA can predict isolated disease-related lnRNA (isolated lnRNA related disease). The results for colorectal cancer, lung adenocarcinoma, stomach cancer and breast cancer have been verified by other researches. The case studies indicated that our method is effective.
Collapse
Affiliation(s)
- Liugen Wang
- School of Science, Zhejiang Sci-Tech University, Hangzhou, 310018, China
| | - Min Shang
- School of Science, Zhejiang Sci-Tech University, Hangzhou, 310018, China
| | - Qi Dai
- College of Life Science, Zhejiang Sci-Tech University, Hangzhou, 310018, China
| | - Ping-An He
- School of Science, Zhejiang Sci-Tech University, Hangzhou, 310018, China.
| |
Collapse
|
11
|
Xuan P, Zhan L, Cui H, Zhang T, Nakaguchi T, Zhang W. Graph Triple-Attention Network for Disease-related LncRNA Prediction. IEEE J Biomed Health Inform 2021; 26:2839-2849. [PMID: 34813484 DOI: 10.1109/jbhi.2021.3130110] [Citation(s) in RCA: 10] [Impact Index Per Article: 3.3] [Reference Citation Analysis] [Abstract] [MESH Headings] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Indexed: 11/06/2022]
Abstract
Abnormal expressions of long non-coding RNAs (lncRNAs) are associated with various human diseases. Identifying disease-related lncRNAs can help clarify complex disease pathogeneses. The latest methods for lncRNA-disease association prediction rely on diverse data about lncRNAs and diseases. These methods, however, cannot adequately integrate the neighbour topological information of lncRNA and disease nodes. Moreover, more intrinsic features of lncRNA-disease node pairs can be explored to better predict the latent associations between lncRNAs and diseases. We developed a novel method, named GTAN, to predict the association propensities between lncRNAs and diseases. GTAN integrates various information about lncRNAs and diseases, including similarities, associations and interactions among lncRNAs, diseases and miRNAs, and exploits neighbour topology and attribute representations of a pair of lncRNA-disease nodes. We adopted in GTAN a graph neural network architecture with three attention mechanisms and multi-layer convolutional neural networks. First, a neighbour-level self-attention mechanism is constructed to learn the importance of each neighbour for an interested lncRNA or disease node. Second, topology-level attention is proposed to enhance contextual dependencies among multiple local topology representations of the lncRNA or disease node. An attention-enhanced graph neural network framework is then established to learn a topology representation of top-ranked neighbours for a pair of lncRNA-disease nodes. GTAN also has attribute-level attention to distinguish various contributions of attributes of the lncRNA-disease pair. Finally, attribute representation is learned by multi-layer CNN to integrate detailed features and representative features of the pair. Extensive experimental results demonstrated that GTAN outperformed state-of-the-art methods. The improved recall rates also showed GTANs capacity for retrieving more actual lncRNA-disease associations in the top-ranked candidates. The ablation studies confirmed the important contributions of three attention mechanisms. Case studies on lung cancer, prostate cancer and colon cancer further showed GTANs ability in discovering potential lncRNA candidates related to diseases.
Collapse
|
12
|
Xie G, Huang B, Sun Y, Wu C, Han Y. RWSF-BLP: a novel lncRNA-disease association prediction model using random walk-based multi-similarity fusion and bidirectional label propagation. Mol Genet Genomics 2021; 296:473-483. [PMID: 33590345 DOI: 10.1007/s00438-021-01764-3] [Citation(s) in RCA: 11] [Impact Index Per Article: 3.7] [Reference Citation Analysis] [Abstract] [Key Words] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Received: 12/06/2020] [Accepted: 01/28/2021] [Indexed: 12/13/2022]
Abstract
An increasing number of studies and experiments have demonstrated that long noncoding RNAs (lncRNAs) have a massive impact on various biological processes. Predicting potential associations between lncRNAs and diseases not only can improve our understanding of the molecular mechanisms of human diseases but also can facilitate the identification of biomarkers for disease diagnosis, treatment, and prevention. However, identifying such associations through experiments is costly and demanding, thereby prompting researchers to develop computational methods to complement these experiments. In this paper, we constructed a novel model called RWSF-BLP (a novel lncRNA-disease association prediction model using Random Walk-based multi-Similarity Fusion and Bidirectional Label Propagation), which applies an efficient random walk-based multi-similarity fusion (RWSF) method to fuse different similarity matrices and utilizes bidirectional label propagation to predict potential lncRNA-disease associations. Leave-one-out cross-validation (LOOCV) and 5-fold cross-validation (5-fold-CV) were implemented in the evaluation RWSF-BLP performance. Results showed that, RWSF-BLP has reliable AUCs of 0.9086 and 0.9115 ± 0.0044 under the framework of LOOCV and 5-fold-CV and outperformed other four canonical methods. Case studies on lung cancer and leukemia demonstrated that potential lncRNA-disease associations can be predicted through our method. Therefore, our method can accurately infer potential lncRNA-disease associations and may be a good choice in future biomedical research.
Collapse
Affiliation(s)
- Guobo Xie
- School of Computer Science, Guangdong University of Technology, Guangzhou, China
| | - Bin Huang
- School of Computer Science, Guangdong University of Technology, Guangzhou, China
| | - Yuping Sun
- School of Computer Science, Guangdong University of Technology, Guangzhou, China.
| | - Changhai Wu
- School of Computer Science, Guangdong University of Technology, Guangzhou, China
| | - Yuqiong Han
- School of Computer Science, Guangdong University of Technology, Guangzhou, China
| |
Collapse
|
13
|
Wang W, Dai Q, Li F, Xiong Y, Wei DQ. MLCDForest: multi-label classification with deep forest in disease prediction for long non-coding RNAs. Brief Bioinform 2020; 22:5855393. [PMID: 32520339 DOI: 10.1093/bib/bbaa104] [Citation(s) in RCA: 24] [Impact Index Per Article: 6.0] [Reference Citation Analysis] [Abstract] [Key Words] [MESH Headings] [Journal Information] [Subscribe] [Scholar Register] [Received: 03/25/2020] [Revised: 04/28/2020] [Accepted: 05/02/2020] [Indexed: 12/18/2022] Open
Abstract
The long non-coding RNAs (lncRNAs) are subject of intensive recent studies due to its association with various human diseases. It is desirable to build the artificial intelligence-based models for prediction of diseases or tissues based on the lncRNAs data, which will be useful in disease diagnosis and therapy. The accuracy and robustness of existing models based on the machine learning techniques are subject to further improvement. In this study, we propose a deep learning model, called Multi-Label Classifications with Deep Forest, termed MLCDForest, to address multi-label classification on tissue prediction for a given lncRNA, which can be regarded as an implementation of the deep forest model in multi-label classification. The MLCDForest is a sequential multi-label-grained scanning method, which distinguishes from the standard deep forest model. It is proposed to train in sequential of multi-labels with label correlation considered. A systematic comparison using the lncRNA-disease association datasets demonstrates that our method consistently shows superior performance over the state-of-the-art methods in disease prediction. Considering label correlation in the sequential multi-label-grained scanning, our model provides a powerful tool to make multi-label classification and tissue prediction based on given lncRNAs.
Collapse
|
14
|
Sheng N, Cui H, Zhang T, Xuan P. Attentional multi-level representation encoding based on convolutional and variance autoencoders for lncRNA-disease association prediction. Brief Bioinform 2020; 22:5841901. [PMID: 32444875 DOI: 10.1093/bib/bbaa067] [Citation(s) in RCA: 34] [Impact Index Per Article: 8.5] [Reference Citation Analysis] [Abstract] [Key Words] [MESH Headings] [Journal Information] [Subscribe] [Scholar Register] [Received: 01/21/2020] [Revised: 03/30/2020] [Accepted: 03/31/2020] [Indexed: 01/01/2023] Open
Abstract
As the abnormalities of long non-coding RNAs (lncRNAs) are closely related to various human diseases, identifying disease-related lncRNAs is important for understanding the pathogenesis of complex diseases. Most of current data-driven methods for disease-related lncRNA candidate prediction are based on diseases and lncRNAs. Those methods, however, fail to consider the deeply embedded node attributes of lncRNA-disease pairs, which contain multiple relations and representations across lncRNAs, diseases and miRNAs. Moreover, the low-dimensional feature distribution at the pairwise level has not been taken into account. We propose a prediction model, VADLP, to extract, encode and adaptively integrate multi-level representations. Firstly, a triple-layer heterogeneous graph is constructed with weighted inter-layer and intra-layer edges to integrate the similarities and correlations among lncRNAs, diseases and miRNAs. We then define three representations including node attributes, pairwise topology and feature distribution. Node attributes are derived from the graph by an embedding strategy to represent the lncRNA-disease associations, which are inferred via their common lncRNAs, diseases and miRNAs. Pairwise topology is formulated by random walk algorithm and encoded by a convolutional autoencoder to represent the hidden topological structural relations between a pair of lncRNA and disease. The new feature distribution is modeled by a variance autoencoder to reveal the underlying lncRNA-disease relationship. Finally, an attentional representation-level integration module is constructed to adaptively fuse the three representations for lncRNA-disease association prediction. The proposed model is tested over a public dataset with a comprehensive list of evaluations. Our model outperforms six state-of-the-art lncRNA-disease prediction models with statistical significance. The ablation study showed the important contributions of three representations. In particular, the improved recall rates under different top $k$ values demonstrate that our model is powerful in discovering true disease-related lncRNAs in the top-ranked candidates. Case studies of three cancers further proved the capacity of our model to discover potential disease-related lncRNAs.
Collapse
|
15
|
Yan C, Zhang Z, Bao S, Hou P, Zhou M, Xu C, Sun J. Computational Methods and Applications for Identifying Disease-Associated lncRNAs as Potential Biomarkers and Therapeutic Targets. MOLECULAR THERAPY. NUCLEIC ACIDS 2020; 21:156-171. [PMID: 32585624 PMCID: PMC7321789 DOI: 10.1016/j.omtn.2020.05.018] [Citation(s) in RCA: 23] [Impact Index Per Article: 5.8] [Reference Citation Analysis] [Abstract] [Key Words] [Track Full Text] [Download PDF] [Figures] [Subscribe] [Scholar Register] [Received: 03/13/2020] [Revised: 04/06/2020] [Accepted: 05/18/2020] [Indexed: 12/12/2022]
Abstract
Long non-coding RNAs (lncRNAs) have been recognized as critical components of a broad genomic regulatory network and play pivotal roles in physiological and pathological processes. Identification of disease-associated lncRNAs is becoming increasingly crucial for fundamentally improving our understanding of molecular mechanisms of disease and developing novel biomarkers and therapeutic targets. Considering lower efficiency and higher time and labor cost of biological experiments, computer-aided inference of disease-associated RNAs has become a promising avenue for facilitating the study of lncRNA functions and provides complementary value for experimental studies. In this study, we first summarize data and knowledge resources publicly available for the study of lncRNA-disease associations. Then, we present an updated systematic overview of dozens of computational methods and models for inferring lncRNA-disease associations proposed in recent years. Finally, we explore the perspectives and challenges for further studies. Our study provides a guide for biologists and medical scientists to look for dedicated resources and more competent tools for accelerating the unraveling of disease-associated lncRNAs.
Collapse
Affiliation(s)
- Congcong Yan
- School of Biomedical Engineering, School of Ophthalmology & Optometry and Eye Hospital, Wenzhou Medical University, Wenzhou 325027, P.R. China
| | - Zicheng Zhang
- School of Biomedical Engineering, School of Ophthalmology & Optometry and Eye Hospital, Wenzhou Medical University, Wenzhou 325027, P.R. China
| | - Siqi Bao
- School of Biomedical Engineering, School of Ophthalmology & Optometry and Eye Hospital, Wenzhou Medical University, Wenzhou 325027, P.R. China
| | - Ping Hou
- School of Biomedical Engineering, School of Ophthalmology & Optometry and Eye Hospital, Wenzhou Medical University, Wenzhou 325027, P.R. China
| | - Meng Zhou
- School of Biomedical Engineering, School of Ophthalmology & Optometry and Eye Hospital, Wenzhou Medical University, Wenzhou 325027, P.R. China
| | - Chongyong Xu
- Department of Radiology, The Second Affiliated Hospital of Wenzhou Medical University, Wenzhou 325027, P.R. China.
| | - Jie Sun
- School of Biomedical Engineering, School of Ophthalmology & Optometry and Eye Hospital, Wenzhou Medical University, Wenzhou 325027, P.R. China.
| |
Collapse
|