1
|
Florentino BR, Parmezan Bonidia R, Sanches NH, da Rocha UN, de Carvalho AC. BioPrediction-RPI: Democratizing the prediction of interaction between non-coding RNA and protein with end-to-end machine learning. Comput Struct Biotechnol J 2024; 23:2267-2276. [PMID: 38827228 PMCID: PMC11140557 DOI: 10.1016/j.csbj.2024.05.031] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Key Words] [Track Full Text] [Download PDF] [Figures] [Journal Information] [Subscribe] [Scholar Register] [Received: 03/14/2024] [Revised: 05/16/2024] [Accepted: 05/16/2024] [Indexed: 06/04/2024] Open
Abstract
Machine Learning (ML) algorithms have been important tools for the extraction of useful knowledge from biological sequences, particularly in healthcare, agriculture, and the environment. However, the categorical and unstructured nature of these sequences requiring usually additional feature engineering steps, before an ML algorithm can be efficiently applied. The addition of these steps to the ML algorithm creates a processing pipeline, known as end-to-end ML. Despite the excellent results obtained by applying end-to-end ML to biotechnology problems, the performance obtained depends on the expertise of the user in the components of the pipeline. In this work, we propose an end-to-end ML-based framework called BioPrediction-RPI, which can identify implicit interactions between sequences, such as pairs of non-coding RNA and proteins, without the need for specialized expertise in end-to-end ML. This framework applies feature engineering to represent each sequence by structural and topological features. These features are divided into feature groups and used to train partial models, whose partial decisions are combined into a final decision, which, provides insights to the user by giving an interpretability report. In our experiments, the developed framework was competitive when compared with various expert-created models. We assessed BioPrediction-RPI with 12 datasets when it presented equal or better performance than all tools in 40% to 100% of cases, depending on the experiment. Finally, BioPrediction-RPI can fine-tune models based on new data and perform at the same level as ML experts, democratizing end-to-end ML and increasing its access to those working in biological sciences.
Collapse
Affiliation(s)
- Bruno Rafael Florentino
- Institute of Mathematics and Computer Sciences, University of São Paulo, São Carlos, 13566-590, São Paulo, Brazil
| | - Robson Parmezan Bonidia
- Institute of Mathematics and Computer Sciences, University of São Paulo, São Carlos, 13566-590, São Paulo, Brazil
- Department of Computer Science, Federal University of Technology-Paraná (UTFPR), Cornélio Procópio, 86300-000, Paraná, Brazil
| | - Natan Henrique Sanches
- Institute of Mathematics and Computer Sciences, University of São Paulo, São Carlos, 13566-590, São Paulo, Brazil
| | - Ulisses N. da Rocha
- Department of Environmental Microbiology, Helmholtz Centre for Environmental Research-UFZ GmbH, Leipzig, Saxony, Germany
| | - André C.P.L.F. de Carvalho
- Institute of Mathematics and Computer Sciences, University of São Paulo, São Carlos, 13566-590, São Paulo, Brazil
| |
Collapse
|
2
|
Wang XF, Yu CQ, You ZH, Wang Y, Huang L, Qiao Y, Wang L, Li ZW. BEROLECMI: a novel prediction method to infer circRNA-miRNA interaction from the role definition of molecular attributes and biological networks. BMC Bioinformatics 2024; 25:264. [PMID: 39127625 DOI: 10.1186/s12859-024-05891-7] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Key Words] [MESH Headings] [Grants] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Received: 08/10/2023] [Accepted: 08/01/2024] [Indexed: 08/12/2024] Open
Abstract
Circular RNA (CircRNA)-microRNA (miRNA) interaction (CMI) is an important model for the regulation of biological processes by non-coding RNA (ncRNA), which provides a new perspective for the study of human complex diseases. However, the existing CMI prediction models mainly rely on the nearest neighbor structure in the biological network, ignoring the molecular network topology, so it is difficult to improve the prediction performance. In this paper, we proposed a new CMI prediction method, BEROLECMI, which uses molecular sequence attributes, molecular self-similarity, and biological network topology to define the specific role feature representation for molecules to infer the new CMI. BEROLECMI effectively makes up for the lack of network topology in the CMI prediction model and achieves the highest prediction performance in three commonly used data sets. In the case study, 14 of the 15 pairs of unknown CMIs were correctly predicted.
Collapse
Affiliation(s)
- Xin-Fei Wang
- School of Information Engineering, Xijing University, Xi'an, China
| | - Chang-Qing Yu
- School of Information Engineering, Xijing University, Xi'an, China.
| | - Zhu-Hong You
- School of Computer Science, Northwestern Polytechnical University, Xi'an, China.
| | - Yan Wang
- Key Laboratory of Symbol Computation and Knowledge Engineering of Ministry of Education, College of Computer Science and Technology, Jilin University, Changchun, China.
- School of Artificial Intelligence, Jilin University, Changchun, China.
| | - Lan Huang
- Key Laboratory of Symbol Computation and Knowledge Engineering of Ministry of Education, College of Computer Science and Technology, Jilin University, Changchun, China
| | - Yan Qiao
- College of Agriculture and Forestry, Longdong University, Qingyang, China
| | - Lei Wang
- School of Computer Science and Technology, China University of Mining and Technology, Xuzhou, China
- Guangxi Academy of Sciences, Nanning, China
| | - Zheng-Wei Li
- School of Computer Science and Technology, China University of Mining and Technology, Xuzhou, China
| |
Collapse
|
3
|
Diao B, Luo J, Guo Y. A comprehensive survey on deep learning-based identification and predicting the interaction mechanism of long non-coding RNAs. Brief Funct Genomics 2024; 23:314-324. [PMID: 38576205 DOI: 10.1093/bfgp/elae010] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Key Words] [MESH Headings] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Received: 12/06/2023] [Revised: 02/25/2024] [Accepted: 03/14/2024] [Indexed: 04/06/2024] Open
Abstract
Long noncoding RNAs (lncRNAs) have been discovered to be extensively involved in eukaryotic epigenetic, transcriptional, and post-transcriptional regulatory processes with the advancements in sequencing technology and genomics research. Therefore, they play crucial roles in the body's normal physiology and various disease outcomes. Presently, numerous unknown lncRNA sequencing data require exploration. Establishing deep learning-based prediction models for lncRNAs provides valuable insights for researchers, substantially reducing time and costs associated with trial and error and facilitating the disease-relevant lncRNA identification for prognosis analysis and targeted drug development as the era of artificial intelligence progresses. However, most lncRNA-related researchers lack awareness of the latest advancements in deep learning models and model selection and application in functional research on lncRNAs. Thus, we elucidate the concept of deep learning models, explore several prevalent deep learning algorithms and their data preferences, conduct a comprehensive review of recent literature studies with exemplary predictive performance over the past 5 years in conjunction with diverse prediction functions, critically analyze and discuss the merits and limitations of current deep learning models and solutions, while also proposing prospects based on cutting-edge advancements in lncRNA research.
Collapse
Affiliation(s)
- Biyu Diao
- Department of Breast Surgery, The First Affiliated Hospital of Ningbo University, No. 59, Liuting Street, Haishu District, Ningbo 315000, China
| | - Jin Luo
- Department of Breast Surgery, The First Affiliated Hospital of Ningbo University, No. 59, Liuting Street, Haishu District, Ningbo 315000, China
| | - Yu Guo
- Department of Breast Surgery, The First Affiliated Hospital of Ningbo University, No. 59, Liuting Street, Haishu District, Ningbo 315000, China
| |
Collapse
|
4
|
Peng L, Ren M, Huang L, Chen M. GEnDDn: An lncRNA-Disease Association Identification Framework Based on Dual-Net Neural Architecture and Deep Neural Network. Interdiscip Sci 2024; 16:418-438. [PMID: 38733474 DOI: 10.1007/s12539-024-00619-w] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Key Words] [MESH Headings] [Grants] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Received: 11/18/2023] [Revised: 02/02/2024] [Accepted: 02/03/2024] [Indexed: 05/13/2024]
Abstract
Accumulating studies have demonstrated close relationships between long non-coding RNAs (lncRNAs) and diseases. Identification of new lncRNA-disease associations (LDAs) enables us to better understand disease mechanisms and further provides promising insights into cancer targeted therapy and anti-cancer drug design. Here, we present an LDA prediction framework called GEnDDn based on deep learning. GEnDDn mainly comprises two steps: First, features of both lncRNAs and diseases are extracted by combining similarity computation, non-negative matrix factorization, and graph attention auto-encoder, respectively. And each lncRNA-disease pair (LDP) is depicted as a vector based on concatenation operation on the extracted features. Subsequently, unknown LDPs are classified by aggregating dual-net neural architecture and deep neural network. Using six different evaluation metrics, we found that GEnDDn surpassed four competing LDA identification methods (SDLDA, LDNFSGB, IPCARF, LDASR) on the lncRNADisease and MNDR databases under fivefold cross-validation experiments on lncRNAs, diseases, LDPs, and independent lncRNAs and independent diseases, respectively. Ablation experiments further validated the powerful LDA prediction performance of GEnDDn. Furthermore, we utilized GEnDDn to find underlying lncRNAs for lung cancer and breast cancer. The results elucidated that there may be dense linkages between IFNG-AS1 and lung cancer as well as between HIF1A-AS1 and breast cancer. The results require further biomedical experimental verification. GEnDDn is publicly available at https://github.com/plhhnu/GEnDDn.
Collapse
Affiliation(s)
- Lihong Peng
- College of Life Science and Chemistry, Hunan University of Technology, Zhuzhou, 412007, China
| | - Mengnan Ren
- College of Life Science and Chemistry, Hunan University of Technology, Zhuzhou, 412007, China
| | - Liangliang Huang
- College of Life Science and Chemistry, Hunan University of Technology, Zhuzhou, 412007, China
| | - Min Chen
- School of Computer Science, Hunan Institute of Technology, Hengyang, 421002, China.
| |
Collapse
|
5
|
Liu W, Teng Z, Li Z, Chen J. CVGAE: A Self-Supervised Generative Method for Gene Regulatory Network Inference Using Single-Cell RNA Sequencing Data. Interdiscip Sci 2024:10.1007/s12539-024-00633-y. [PMID: 38778003 DOI: 10.1007/s12539-024-00633-y] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Key Words] [Grants] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Received: 11/05/2023] [Revised: 04/07/2024] [Accepted: 04/09/2024] [Indexed: 05/25/2024]
Abstract
Gene regulatory network (GRN) inference based on single-cell RNA sequencing data (scRNAseq) plays a crucial role in understanding the regulatory mechanisms between genes. Various computational methods have been employed for GRN inference, but their performance in terms of network accuracy and model generalization is not satisfactory, and their poor performance is caused by high-dimensional data and network sparsity. In this paper, we propose a self-supervised method for gene regulatory network inference using single-cell RNA sequencing data (CVGAE). CVGAE uses graph neural network for inductive representation learning, which merges gene expression data and observed topology into a low-dimensional vector space. The well-trained vectors will be used to calculate mathematical distance of each gene, and further predict interactions between genes. In overall framework, FastICA is implemented to relief computational complexity caused by high dimensional data, and CVGAE adopts multi-stacked GraphSAGE layers as an encoder and an improved decoder to overcome network sparsity. CVGAE is evaluated on several single cell datasets containing four related ground-truth networks, and the result shows that CVGAE achieve better performance than comparative methods. To validate learning and generalization capabilities, CVGAE is applied in few-shot environment by change the ratio of train set and test set. In condition of few-shot, CVGAE obtains comparable or superior performance.
Collapse
Affiliation(s)
- Wei Liu
- School of Computer Science, Xiangtan University, Xiangtan, 411105, China.
| | - Zhijie Teng
- School of Computer Science, Xiangtan University, Xiangtan, 411105, China
| | - Zejun Li
- School of Computer Science and Engineering, Hunan Institute of Technology, Hengyang, 412002, China
| | - Jing Chen
- School of Electronic and Information Engineering, Suzhou University of Science and Technology, Suzhou, 215009, China.
| |
Collapse
|
6
|
Jiang L, Jia L, Wang Y, Wu Y, Yue J. Adap-BDCM: Adaptive Bilinear Dynamic Cascade Model for Classification Tasks on CNV Datasets. Interdiscip Sci 2024:10.1007/s12539-024-00635-w. [PMID: 38758306 DOI: 10.1007/s12539-024-00635-w] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Key Words] [Grants] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Received: 12/21/2023] [Revised: 04/18/2024] [Accepted: 04/23/2024] [Indexed: 05/18/2024]
Abstract
Copy number variation (CNV) is an essential genetic driving factor of cancer formation and progression, making intelligent classification based on CNV feasible. However, there are a few challenges in the current machine learning and deep learning methods, such as the design of base classifier combination schemes in ensemble methods and the selection of layers of neural networks, which often result in low accuracy. Therefore, an adaptive bilinear dynamic cascade model (Adap-BDCM) is developed to further enhance the accuracy and applicability of these methods for intelligent classification on CNV datasets. In this model, a feature selection module is introduced to mitigate the interference of redundant information, and a bilinear model based on the gated attention mechanism is proposed to extract more beneficial deep fusion features. Furthermore, an adaptive base classifier selection scheme is designed to overcome the difficulty of manually designing base classifier combinations and enhance the applicability of the model. Lastly, a novel feature fusion scheme with an attribute recall submodule is constructed, effectively avoiding getting stuck in local solutions and missing some valuable information. Numerous experiments have demonstrated that our Adap-BDCM model exhibits optimal performance in cancer classification, stage prediction, and recurrence on CNV datasets. This study can assist physicians in making diagnoses faster and better.
Collapse
Affiliation(s)
- Liancheng Jiang
- College of Computer Science and Technology (College of Data Science), Taiyuan University of Technology, Taiyuan, 030600, China
| | - Liye Jia
- College of Computer Science and Technology, Taiyuan Normal University, Taiyuan, 030619, China
| | - Yizhen Wang
- College of Computer Science and Technology (College of Data Science), Taiyuan University of Technology, Taiyuan, 030600, China
| | - Yongfei Wu
- College of Computer Science and Technology (College of Data Science), Taiyuan University of Technology, Taiyuan, 030600, China
| | - Junhong Yue
- College of Computer Science and Technology (College of Data Science), Taiyuan University of Technology, Taiyuan, 030600, China.
| |
Collapse
|
7
|
Zhang Y, Li X. Empowering Graph Neural Networks with Block-Based Dual Adaptive Deep Adjustment for Drug Resistance-Related NcRNA Discovery. J Chem Inf Model 2024; 64:3537-3547. [PMID: 38523272 DOI: 10.1021/acs.jcim.3c01973] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [MESH Headings] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Indexed: 03/26/2024]
Abstract
Drug resistance to chemotherapeutic agents remains a formidable challenge in cancer treatment, significantly impacting treatment efficacy. Extensive research has exposed the intimate involvement of noncoding RNAs (ncRNAs) in conferring resistance to cancer drugs. Understanding the intricate associations between ncRNAs and drug resistance is of pivotal importance in advancing clinical interventions and expediting drug development. However, traditional biological experimental methods are hampered by limitations, such as labor intensiveness, time consumption, and constraints in scalability. Addressing these challenges necessitates the development of efficient computational methods for the accurate prediction of potential ncRNA-drug resistance associations (NDRA). However, most existing predictive models primarily focus on known ncRNA-drug resistance associations, often neglecting the critical aspect of similarity information between ncRNAs and drug resistance. This oversight may hinder the accuracy of characterizing these associations. To overcome the limitations of existing computational models, we proposed B-NDRA, a computational framework designed for the discovery of drug resistance-related ncRNA. Initially, we constructed a heterogeneous graph that integrates ncRNA-drug resistance pairs, leveraging both known associations and similarity fusion information between ncRNAs and drug resistance. Subsequently, we employed an attention mechanism to aggregate local features of graph nodes following a dimensionality reduction of node features. Further, a graph neural network (GNN) facilitated the learning of global node embeddings. Notably, the integration of dual adaptive deep adjustment architectures, encompassing intrablock and interblock methodologies, enabled efficient extraction of global features while balancing local and global features. Finally, B-NDRA employed a multilayer perceptron to predict associations between ncRNAs and drug resistance. Through rigorous 5-fold cross-validation, B-NDRA achieved average AUC, AUPR, Accuracy, Precision, Recall, and F1-score values of 92.2%, 91.9%, 84.88%, 86.9%, 82.37%, and 84.44%, respectively. Furthermore, comparative evaluations were conducted on established models, namely, GAEMDA, GRPAMDA, and LRGCPND. The results, obtained through three distinct 5-fold cross-validation strategies, demonstrated a notable performance improvement across almost all metrics for our B-NDRA. Specific case studies targeting Doxorubicin and Imatinib further validated the practicality of our B-NDRA in discovering potential NDRA. These results confirm the potential of our B-NDRA as a valuable tool in advancing cancer research and therapeutic development. The source code and data set of B-NDRA can be found at https://github.com/XuanLi1145/B-NDRA.
Collapse
Affiliation(s)
- Yi Zhang
- Guilin University of Technology, Guilin 541004, China
- Guangxi Key Laboratory of Embedded Technology and Intelligent System, Guilin University of Technology, Guilin 541004, China
| | - Xuanzhao Li
- Guilin University of Technology, Guilin 541004, China
- Guangxi Key Laboratory of Embedded Technology and Intelligent System, Guilin University of Technology, Guilin 541004, China
| |
Collapse
|
8
|
Li X, Qu W, Yan J, Tan J. RPI-EDLCN: An Ensemble Deep Learning Framework Based on Capsule Network for ncRNA-Protein Interaction Prediction. J Chem Inf Model 2024; 64:2221-2235. [PMID: 37158609 DOI: 10.1021/acs.jcim.3c00377] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [MESH Headings] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Indexed: 05/10/2023]
Abstract
Noncoding RNAs (ncRNAs) play crucial roles in many cellular life activities by interacting with proteins. Identification of ncRNA-protein interactions (ncRPIs) is key to understanding the function of ncRNAs. Although a number of computational methods for predicting ncRPIs have been developed, the problem of predicting ncRPIs remains challenging. It has always been the focus of ncRPIs research to select suitable feature extraction methods and develop a deep learning architecture with better recognition performance. In this work, we proposed an ensemble deep learning framework, RPI-EDLCN, based on a capsule network (CapsuleNet) to predict ncRPIs. In terms of feature input, we extracted the sequence features, secondary structure sequence features, motif information, and physicochemical properties of ncRNA/protein. The sequence and secondary structure sequence features of ncRNA/protein are encoded by the conjoint k-mer method and then input into an ensemble deep learning model based on CapsuleNet by combining the motif information and physicochemical properties. In this model, the encoding features are processed by convolution neural network (CNN), deep neural network (DNN), and stacked autoencoder (SAE). Then the advanced features obtained from the processing are input into the CapsuleNet for further feature learning. Compared with other state-of-the-art methods under 5-fold cross-validation, the performance of RPI-EDLCN is the best, and the accuracy of RPI-EDLCN on RPI1807, RPI2241, and NPInter v2.0 data sets was 93.8%, 88.2%, and 91.9%, respectively. The results of the independent test indicated that RPI-EDLCN can effectively predict potential ncRPIs in different organisms. In addition, RPI-EDLCN successfully predicted hub ncRNAs and proteins in Mus musculus ncRNA-protein networks. Overall, our model can be used as an effective tool to predict ncRPIs and provides some useful guidance for future biological studies.
Collapse
Affiliation(s)
- Xiaoyi Li
- Department of Biomedical Engineering, Faculty of Environment and Life, Beijing University of Technology, Beijing International Science and Technology Cooperation Base for Intelligent Physiological Measurement and Clinical Transformation, Beijing 100124, China
| | - Wenyan Qu
- Department of Biomedical Engineering, Faculty of Environment and Life, Beijing University of Technology, Beijing International Science and Technology Cooperation Base for Intelligent Physiological Measurement and Clinical Transformation, Beijing 100124, China
| | - Jing Yan
- Department of Biomedical Engineering, Faculty of Environment and Life, Beijing University of Technology, Beijing International Science and Technology Cooperation Base for Intelligent Physiological Measurement and Clinical Transformation, Beijing 100124, China
| | - Jianjun Tan
- Department of Biomedical Engineering, Faculty of Environment and Life, Beijing University of Technology, Beijing International Science and Technology Cooperation Base for Intelligent Physiological Measurement and Clinical Transformation, Beijing 100124, China
| |
Collapse
|
9
|
Zhou L, Peng X, Zeng L, Peng L. Finding potential lncRNA-disease associations using a boosting-based ensemble learning model. Front Genet 2024; 15:1356205. [PMID: 38495672 PMCID: PMC10940470 DOI: 10.3389/fgene.2024.1356205] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Key Words] [Grants] [Track Full Text] [Figures] [Journal Information] [Subscribe] [Scholar Register] [Received: 12/18/2023] [Accepted: 02/01/2024] [Indexed: 03/19/2024] Open
Abstract
Introduction: Long non-coding RNAs (lncRNAs) have been in the clinical use as potential prognostic biomarkers of various types of cancer. Identifying associations between lncRNAs and diseases helps capture the potential biomarkers and design efficient therapeutic options for diseases. Wet experiments for identifying these associations are costly and laborious. Methods: We developed LDA-SABC, a novel boosting-based framework for lncRNA-disease association (LDA) prediction. LDA-SABC extracts LDA features based on singular value decomposition (SVD) and classifies lncRNA-disease pairs (LDPs) by incorporating LightGBM and AdaBoost into the convolutional neural network. Results: The LDA-SABC performance was evaluated under five-fold cross validations (CVs) on lncRNAs, diseases, and LDPs. It obviously outperformed four other classical LDA inference methods (SDLDA, LDNFSGB, LDASR, and IPCAF) through precision, recall, accuracy, F1 score, AUC, and AUPR. Based on the accurate LDA prediction performance of LDA-SABC, we used it to find potential lncRNA biomarkers for lung cancer. The results elucidated that 7SK and HULC could have a relationship with non-small-cell lung cancer (NSCLC) and lung adenocarcinoma (LUAD), respectively. Conclusion: We hope that our proposed LDA-SABC method can help improve the LDA identification.
Collapse
Affiliation(s)
- Liqian Zhou
- School of Computer Science, Hunan University of Technology, Zhuzhou, Hunan, China
| | - Xinhuai Peng
- School of Computer Science, Hunan University of Technology, Zhuzhou, Hunan, China
| | - Lijun Zeng
- School of Computer Science, Hunan Institute of Technology, Hengyang, China
| | - Lihong Peng
- School of Computer Science, Hunan University of Technology, Zhuzhou, Hunan, China
| |
Collapse
|
10
|
Zhuang J, Midgley AC, Wei Y, Liu Q, Kong D, Huang X. Machine-Learning-Assisted Nanozyme Design: Lessons from Materials and Engineered Enzymes. ADVANCED MATERIALS (DEERFIELD BEACH, FLA.) 2024; 36:e2210848. [PMID: 36701424 DOI: 10.1002/adma.202210848] [Citation(s) in RCA: 21] [Impact Index Per Article: 21.0] [Reference Citation Analysis] [Abstract] [Key Words] [MESH Headings] [Grants] [Track Full Text] [Subscribe] [Scholar Register] [Received: 11/21/2022] [Revised: 01/03/2023] [Indexed: 05/11/2023]
Abstract
Nanozymes are nanomaterials that exhibit enzyme-like biomimicry. In combination with intrinsic characteristics of nanomaterials, nanozymes have broad applicability in materials science, chemical engineering, bioengineering, biochemistry, and disease theranostics. Recently, the heterogeneity of published results has highlighted the complexity and diversity of nanozymes in terms of consistency of catalytic capacity. Machine learning (ML) shows promising potential for discovering new materials, yet it remains challenging for the design of new nanozymes based on ML approaches. Alternatively, ML is employed to promote optimization of intelligent design and application of catalytic materials and engineered enzymes. Incorporation of the successful ML algorithms used in the intelligent design of catalytic materials and engineered enzymes can concomitantly facilitate the guided development of next-generation nanozymes with desirable properties. Here, recent progress in ML, its utilization in the design of catalytic materials and enzymes, and how emergent ML applications serve as promising strategies to circumvent challenges associated with time-expensive and laborious testing in nanozyme research and development are summarized. The potential applications of successful examples of ML-aided catalytic materials and engineered enzymes in nanozyme design are also highlighted, with special focus on the unified aims in enhancing design and recapitulation of substrate selectivity and catalytic activity.
Collapse
Affiliation(s)
- Jie Zhuang
- School of Medicine, and State, Key Laboratory of Medicinal Chemical Biology, Nankai University, Tianjin, 300071, China
| | - Adam C Midgley
- Key Laboratory of Bioactive Materials for the Ministry of Education, College of Life Sciences, State Key Laboratory of Medicinal Chemical Biology, and Frontiers, Science Center for Cell Responses, Nankai University, Tianjin, 300071, China
| | - Yonghua Wei
- Key Laboratory of Bioactive Materials for the Ministry of Education, College of Life Sciences, State Key Laboratory of Medicinal Chemical Biology, and Frontiers, Science Center for Cell Responses, Nankai University, Tianjin, 300071, China
| | - Qiqi Liu
- Key Laboratory of Bioactive Materials for the Ministry of Education, College of Life Sciences, State Key Laboratory of Medicinal Chemical Biology, and Frontiers, Science Center for Cell Responses, Nankai University, Tianjin, 300071, China
| | - Deling Kong
- Key Laboratory of Bioactive Materials for the Ministry of Education, College of Life Sciences, State Key Laboratory of Medicinal Chemical Biology, and Frontiers, Science Center for Cell Responses, Nankai University, Tianjin, 300071, China
| | - Xinglu Huang
- Key Laboratory of Bioactive Materials for the Ministry of Education, College of Life Sciences, State Key Laboratory of Medicinal Chemical Biology, and Frontiers, Science Center for Cell Responses, Nankai University, Tianjin, 300071, China
| |
Collapse
|
11
|
Peng L, Huang L, Su Q, Tian G, Chen M, Han G. LDA-VGHB: identifying potential lncRNA-disease associations with singular value decomposition, variational graph auto-encoder and heterogeneous Newton boosting machine. Brief Bioinform 2023; 25:bbad466. [PMID: 38127089 PMCID: PMC10734633 DOI: 10.1093/bib/bbad466] [Citation(s) in RCA: 6] [Impact Index Per Article: 6.0] [Reference Citation Analysis] [Abstract] [Key Words] [MESH Headings] [Grants] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Received: 08/05/2023] [Revised: 10/05/2023] [Accepted: 11/25/2023] [Indexed: 12/23/2023] Open
Abstract
Long noncoding RNAs (lncRNAs) participate in various biological processes and have close linkages with diseases. In vivo and in vitro experiments have validated many associations between lncRNAs and diseases. However, biological experiments are time-consuming and expensive. Here, we introduce LDA-VGHB, an lncRNA-disease association (LDA) identification framework, by incorporating feature extraction based on singular value decomposition and variational graph autoencoder and LDA classification based on heterogeneous Newton boosting machine. LDA-VGHB was compared with four classical LDA prediction methods (i.e. SDLDA, LDNFSGB, IPCARF and LDASR) and four popular boosting models (XGBoost, AdaBoost, CatBoost and LightGBM) under 5-fold cross-validations on lncRNAs, diseases, lncRNA-disease pairs and independent lncRNAs and independent diseases, respectively. It greatly outperformed the other methods with its prominent performance under four different cross-validations on the lncRNADisease and MNDR databases. We further investigated potential lncRNAs for lung cancer, breast cancer, colorectal cancer and kidney neoplasms and inferred the top 20 lncRNAs associated with them among all their unobserved lncRNAs. The results showed that most of the predicted top 20 lncRNAs have been verified by biomedical experiments provided by the Lnc2Cancer 3.0, lncRNADisease v2.0 and RNADisease databases as well as publications. We found that HAR1A, KCNQ1DN, ZFAT-AS1 and HAR1B could associate with lung cancer, breast cancer, colorectal cancer and kidney neoplasms, respectively. The results need further biological experimental validation. We foresee that LDA-VGHB was capable of identifying possible lncRNAs for complex diseases. LDA-VGHB is publicly available at https://github.com/plhhnu/LDA-VGHB.
Collapse
Affiliation(s)
- Lihong Peng
- School of Computer Science, Hunan University of Technology, 412007, Hunan, China
- College of Life Sciences and Chemistry, Hunan University of Technology, 412007, Hunan, China
| | - Liangliang Huang
- School of Computer Science, Hunan University of Technology, 412007, Hunan, China
| | - Qiongli Su
- Department of Pharmacy, the Affiliated Zhuzhou Hospital Xiangya Medical College CSU, 412007, Hunan, China
| | - Geng Tian
- Geneis (Beijing) Co. Ltd, China, 100102, Beijing, China
| | - Min Chen
- School of Computer Science, Hunan Institute of Technology, 421002, No. 18 Henghua Road, Zhuhui District, Hengyang, Hunan, China
| | - Guosheng Han
- School of Mathematics and Computational Science, Xiangtan University, 411105, Yuhu District, Xiangtan, Hunan, China
- Hunan Key Laboratory for Computation and Simulation in Science and Engineering, Xiangtan University, 411105, Yuhu District, Xiangtan, Hunan, China
| |
Collapse
|
12
|
Xie W, Chen X, Zheng Z, Wang F, Zhu X, Lin Q, Sun Y, Wong KC. LncRNA-Top: Controlled deep learning approaches for lncRNA gene regulatory relationship annotations across different platforms. iScience 2023; 26:108197. [PMID: 37965148 PMCID: PMC10641498 DOI: 10.1016/j.isci.2023.108197] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Key Words] [Track Full Text] [Download PDF] [Figures] [Journal Information] [Subscribe] [Scholar Register] [Received: 04/24/2023] [Revised: 08/10/2023] [Accepted: 10/10/2023] [Indexed: 11/16/2023] Open
Abstract
By soaking microRNAs (miRNAs), long non-coding RNAs (lncRNAs) have the potential to regulate gene expression. Few methods have been created based on this mechanism to anticipate the lncRNA-gene relationship prediction. Hence, we present lncRNA-Top to forecast potential lncRNA-gene regulation relationships. Specifically, we constructed controlled deep-learning methods using 12417 lncRNAs and 16127 genes. We have provided retrospective and innovative views among negative sampling, random seeds, cross-validation, metrics, and independent datasets. The AUC, AUPR, and our defined precision@k were leveraged to evaluate performance. In-depth case studies demonstrate that 47 out of 100 projected top unknown pairings were recorded in publications, supporting the predictive power. Our additional software can annotate the scores with target candidates. The lncRNA-Top will be a helpful tool to uncover prospective lncRNA targets and better comprehend the regulatory processes of lncRNAs.
Collapse
Affiliation(s)
- Weidun Xie
- Department of Computer Science, City University of Hong Kong, Kowloon Tong, Hong Kong SAR
| | - Xingjian Chen
- Department of Computer Science, City University of Hong Kong, Kowloon Tong, Hong Kong SAR
| | - Zetian Zheng
- Department of Computer Science, City University of Hong Kong, Kowloon Tong, Hong Kong SAR
| | - Fuzhou Wang
- Department of Computer Science, City University of Hong Kong, Kowloon Tong, Hong Kong SAR
| | - Xiaowei Zhu
- Department of Neuroscience, Jockey Club College of Veterinary Medicine and Life Sciences, City University of Hong Kong, Kowloon Tong, Hong Kong SAR
| | - Qiuzhen Lin
- College of Computer Science and Software Engineering, Shenzhen University, Shenzhen, China
| | - Yanni Sun
- Department of Electrical Engineering, City University of Hong Kong, Kowloon Tong, Hong Kong SAR
| | - Ka-Chun Wong
- Department of Computer Science, City University of Hong Kong, Kowloon Tong, Hong Kong SAR
- Shenzhen Research Institute, City University of Hong Kong, Shenzhen, China
- Hong Kong Institute for Data Science, City University of Hong Kong, Kowloon Tong, Hong Kong SAR
| |
Collapse
|
13
|
Su Z, Lu H, Wu Y, Li Z, Duan L. Predicting potential lncRNA biomarkers for lung cancer and neuroblastoma based on an ensemble of a deep neural network and LightGBM. Front Genet 2023; 14:1238095. [PMID: 37655066 PMCID: PMC10466784 DOI: 10.3389/fgene.2023.1238095] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Key Words] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Received: 06/10/2023] [Accepted: 07/19/2023] [Indexed: 09/02/2023] Open
Abstract
Introduction: Lung cancer is one of the most frequent neoplasms worldwide with approximately 2.2 million new cases and 1.8 million deaths each year. The expression levels of programmed death ligand-1 (PDL1) demonstrate a complex association with lung cancer. Neuroblastoma is a high-risk malignant tumor and is mainly involved in childhood patients. Identification of new biomarkers for these two diseases can significantly promote their diagnosis and therapy. However, in vivo experiments to discover potential biomarkers are costly and laborious. Consequently, artificial intelligence technologies, especially machine learning methods, provide a powerful avenue to find new biomarkers for various diseases. Methods: We developed a machine learning-based method named LDAenDL to detect potential long noncoding RNA (lncRNA) biomarkers for lung cancer and neuroblastoma using an ensemble of a deep neural network and LightGBM. LDAenDL first computes the Gaussian kernel similarity and functional similarity of lncRNAs and the Gaussian kernel similarity and semantic similarity of diseases to obtain their similar networks. Next, LDAenDL combines a graph convolutional network, graph attention network, and convolutional neural network to learn the biological features of the lncRNAs and diseases based on their similarity networks. Third, these features are concatenated and fed to an ensemble model composed of a deep neural network and LightGBM to find new lncRNA-disease associations (LDAs). Finally, the proposed LDAenDL method is applied to identify possible lncRNA biomarkers associated with lung cancer and neuroblastoma. Results: The experimental results show that LDAenDL computed the best AUCs of 0.8701, 107 0.8953, and 0.9110 under cross-validation on lncRNAs, diseases, and lncRNA-disease pairs on Dataset 1, respectively, and 0.9490, 0.9157, and 0.9708 on Dataset 2, respectively. Furthermore, AUPRs of 0.8903, 0.9061, and 0.9166 under three cross-validations were obtained on Dataset 1, and 0.9582, 0.9122, and 0.9743 on Dataset 2. The results demonstrate that LDAenDL significantly outperformed the other four classical LDA prediction methods (i.e., SDLDA, LDNFSGB, IPCAF, and LDASR). Case studies demonstrate that CCDC26 and IFNG-AS1 may be new biomarkers of lung cancer, SNHG3 may associate with PDL1 for lung cancer, and HOTAIR and BDNF-AS may be potential biomarkers of neuroblastoma. Conclusion: We hope that the proposed LDAenDL method can help the development of targeted therapies for these two diseases.
Collapse
Affiliation(s)
- Zhenguo Su
- Clinical Lab, Yantai Affiliated Hospital of Binzhou Medical University, Yantai, China
| | - Huihui Lu
- Department of Thoracic Cardiovascular Surgery, Hunan Province Directly Affiliated TCM Hospital, Zhuzhou, China
| | - Yan Wu
- Geneis (Beijing) Co., Ltd., Beijing, China
| | - Zejun Li
- School of Computer Science, Hunan Institute of Technology, Hengyang, China
| | - Lian Duan
- Faculty of Pediatrics, The Chinese PLA General Hospital, Beijing, China
- Department of Pediatric Surgery, The Seventh Medical Center of PLA General Hospital, Beijing, China
- National Engineering Laboratory for Birth Defects Prevention and Control of Key Technology, Beijing, China
- Beijing Key Laboratory of Pediatric Organ Failure, Beijing, China
| |
Collapse
|
14
|
Kim Y, Lee M. Deep Learning Approaches for lncRNA-Mediated Mechanisms: A Comprehensive Review of Recent Developments. Int J Mol Sci 2023; 24:10299. [PMID: 37373445 DOI: 10.3390/ijms241210299] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Key Words] [Grants] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Received: 05/24/2023] [Revised: 06/16/2023] [Accepted: 06/17/2023] [Indexed: 06/29/2023] Open
Abstract
This review paper provides an extensive analysis of the rapidly evolving convergence of deep learning and long non-coding RNAs (lncRNAs). Considering the recent advancements in deep learning and the increasing recognition of lncRNAs as crucial components in various biological processes, this review aims to offer a comprehensive examination of these intertwined research areas. The remarkable progress in deep learning necessitates thoroughly exploring its latest applications in the study of lncRNAs. Therefore, this review provides insights into the growing significance of incorporating deep learning methodologies to unravel the intricate roles of lncRNAs. By scrutinizing the most recent research spanning from 2021 to 2023, this paper provides a comprehensive understanding of how deep learning techniques are employed in investigating lncRNAs, thereby contributing valuable insights to this rapidly evolving field. The review is aimed at researchers and practitioners looking to integrate deep learning advancements into their lncRNA studies.
Collapse
Affiliation(s)
- Yoojoong Kim
- School of Computer Science and Information Engineering, The Catholic University of Korea, Bucheon 14662, Republic of Korea
| | - Minhyeok Lee
- School of Electrical and Electronics Engineering, Chung-Ang University, Seoul 06974, Republic of Korea
| |
Collapse
|
15
|
Wong L, Wang L, You ZH, Yuan CA, Huang YA, Cao MY. GKLOMLI: a link prediction model for inferring miRNA-lncRNA interactions by using Gaussian kernel-based method on network profile and linear optimization algorithm. BMC Bioinformatics 2023; 24:188. [PMID: 37158823 PMCID: PMC10169329 DOI: 10.1186/s12859-023-05309-w] [Citation(s) in RCA: 15] [Impact Index Per Article: 15.0] [Reference Citation Analysis] [Abstract] [Key Words] [MESH Headings] [Grants] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Received: 08/30/2022] [Accepted: 04/27/2023] [Indexed: 05/10/2023] Open
Abstract
BACKGROUND The limited knowledge of miRNA-lncRNA interactions is considered as an obstruction of revealing the regulatory mechanism. Accumulating evidence on Human diseases indicates that the modulation of gene expression has a great relationship with the interactions between miRNAs and lncRNAs. However, such interaction validation via crosslinking-immunoprecipitation and high-throughput sequencing (CLIP-seq) experiments that inevitably costs too much money and time but with unsatisfactory results. Therefore, more and more computational prediction tools have been developed to offer many reliable candidates for a better design of further bio-experiments. METHODS In this work, we proposed a novel link prediction model based on Gaussian kernel-based method and linear optimization algorithm for inferring miRNA-lncRNA interactions (GKLOMLI). Given an observed miRNA-lncRNA interaction network, the Gaussian kernel-based method was employed to output two similarity matrixes of miRNAs and lncRNAs. Based on the integrated matrix combined with similarity matrixes and the observed interaction network, a linear optimization-based link prediction model was trained for inferring miRNA-lncRNA interactions. RESULTS To evaluate the performance of our proposed method, k-fold cross-validation (CV) and leave-one-out CV were implemented, in which each CV experiment was carried out 100 times on a training set generated randomly. The high area under the curves (AUCs) at 0.8623 ± 0.0027 (2-fold CV), 0.9053 ± 0.0017 (5-fold CV), 0.9151 ± 0.0013 (10-fold CV), and 0.9236 (LOO-CV), illustrated the precision and reliability of our proposed method. CONCLUSION GKLOMLI with high performance is anticipated to be used to reveal underlying interactions between miRNA and their target lncRNAs, and deciphers the potential mechanisms of the complex diseases.
Collapse
Affiliation(s)
- Leon Wong
- Guangxi Key Lab of Human-machine Interaction and Intelligent Decision, Guangxi Academy of Sciences, Nanning, 530007, China
- Institute of Machine Learning and Systems Biology, School of Electronics and Information Engineering, Tongji University, 200092, Shanghai, China
| | - Lei Wang
- Guangxi Key Lab of Human-machine Interaction and Intelligent Decision, Guangxi Academy of Sciences, Nanning, 530007, China.
- College of Information Science and Engineering, Zaozhuang University, Zaozhuang, 277160, China.
| | - Zhu-Hong You
- School of Computer Science, Northwestern Polytechnical University, Xi'an, 710139, China.
| | - Chang-An Yuan
- Guangxi Key Lab of Human-machine Interaction and Intelligent Decision, Guangxi Academy of Sciences, Nanning, 530007, China
| | - Yu-An Huang
- School of Computer Science, Northwestern Polytechnical University, Xi'an, 710139, China
| | - Mei-Yuan Cao
- School of Electrical and Electronic Engineering, Guangdong Technology College, Zhaoqing, 526100, China
- Faculty of Information Science and Technology, Universiti Kebangsaan Malaysia, UKM, 43600, Bangi, Selangor, Malaysia
| |
Collapse
|
16
|
Zhao Z, Luo Q, Liu Y, Jiang K, Zhou L, Dai R, Wang H. Multi-level integrative analysis of the roles of lncRNAs and differential mRNAs in the progression of chronic pancreatitis to pancreatic ductal adenocarcinoma. BMC Genomics 2023; 24:101. [PMID: 36879212 PMCID: PMC9990329 DOI: 10.1186/s12864-023-09209-4] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Key Words] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Received: 11/28/2022] [Accepted: 02/27/2023] [Indexed: 03/08/2023] Open
Abstract
BACKGROUND Pancreatic ductal adenocarcinoma (PDAC) is one of the most malignant tumors and approximately 5% of patients with chronic pancreatitis (CP) inevitably develop PDAC. This study aims explore the key gene regulation involved in the progression of CP to PDAC, with a particular emphasis on the function of lncRNAs. RESULTS A total of 103 pancreatic tissue samples collected from 11 to 92 patients with CP and PDAC, respectively, were included in this study. After normalizing and logarithmically converting the original data, differentially expressed lncRNAs (DElncRNAs) and mRNAs (DEGs) in each dataset were selected. To determine the main functional pathways of differential mRNAs, we further annotated DEGs using gene ontology (GO) and analyzed the Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment. In addition, the interaction between lncRNA-miRNA-mRNA was clarified and the protein-protein interaction (PPI) network was constructed to screen for key modules and determine hub genes. Finally, quantitative real-time polymerase chain reaction (qPCR) was used to detect the changes in non-coding RNAs and key mRNAs in the pancreatic tissues of patients with CP and PDAC. In this study, 230 lncRNAs and 17,668 mRNAs were included. There were nine upregulated lncRNAs and 188 downregulated lncRNAs. Furthermore, 2334 upregulated differential mRNAs and 10,341 downregulated differential mRNAs were included in the enrichment analysis. From the KEGG enrichment analysis, cytokine-cytokine receptor interaction, calcium signaling pathway, cAMP signaling pathway, and nicotine addiction exhibited significant differences. Additionally, a total of 52 lncRNAs, 104 miRNAs, and 312 mRNAs were included in the construction of a potential lncRNA-miRNA-mRNA regulatory network. PPI network was established and two of the five central DEGs were created in this module, suggesting that lysophosphatidic acid receptor 1 (LPAR1) and regulator of calcineurin 2 (RCAN2) may play significant roles in the progression from CP to PDAC. Finally, the PCR results suggested that LINC01547/hsa-miR-4694-3p/LPAR1 and LINC00482/hsa-miR-6756-3p/RCAN2 play important roles in the carcinogenesis process of CP. CONCLUSION Two signaling axes critical in the progression of CP to PDAC were screened out. Our findings will be useful for novel insights into the molecular mechanism and potential diagnostic or therapeutic biomarkers for CP and PDAC.
Collapse
Affiliation(s)
- Zhirong Zhao
- Affiliated Hospital of Southwest Jiaotong University, The General Hospital of Western Theater Command, Chengdu, 610031, Sichuan, China.,Pancreatic injury and repair Key laboratory of Sichuan Province, The General Hospital of Western Theater Command, Chengdu, Sichuan, China
| | - Qiang Luo
- Department of Cardiology, Affiliated Hospital of Southwest Jiaotong University, The Third People's Hospital of Chengdu, Chengdu, Sichuan, China
| | - Yi Liu
- School of Medicine, Jianghan University, 430056, Wuhan, Hubei, China
| | - Kexin Jiang
- Affiliated Hospital of Southwest Jiaotong University, The General Hospital of Western Theater Command, Chengdu, 610031, Sichuan, China
| | - Lichen Zhou
- Affiliated Hospital of Southwest Jiaotong University, The General Hospital of Western Theater Command, Chengdu, 610031, Sichuan, China
| | - Ruiwu Dai
- Affiliated Hospital of Southwest Jiaotong University, The General Hospital of Western Theater Command, Chengdu, 610031, Sichuan, China. .,Pancreatic injury and repair Key laboratory of Sichuan Province, The General Hospital of Western Theater Command, Chengdu, Sichuan, China.
| | - Han Wang
- Department of Cardiology, Affiliated Hospital of Southwest Jiaotong University, The Third People's Hospital of Chengdu, Chengdu, Sichuan, China.
| |
Collapse
|
17
|
Fu Y, Si A, Wei X, Lin X, Ma Y, Qiu H, Guo Z, Pan Y, Zhang Y, Kong X, Li S, Shi Y, Wu H. Combining a machine-learning derived 4-lncRNA signature with AFP and TNM stages in predicting early recurrence of hepatocellular carcinoma. BMC Genomics 2023; 24:89. [PMID: 36849926 PMCID: PMC9972730 DOI: 10.1186/s12864-023-09194-8] [Citation(s) in RCA: 4] [Impact Index Per Article: 4.0] [Reference Citation Analysis] [Abstract] [Key Words] [Track Full Text] [Figures] [Journal Information] [Subscribe] [Scholar Register] [Received: 01/09/2023] [Accepted: 02/17/2023] [Indexed: 03/01/2023] Open
Abstract
BACKGROUND Near 70% of hepatocellular carcinoma (HCC) recurrence is early recurrence within 2-year post surgery. Long non-coding RNAs (lncRNAs) are intensively involved in HCC progression and serve as biomarkers for HCC prognosis. The aim of this study is to construct a lncRNA-based signature for predicting HCC early recurrence. METHODS Data of RNA expression and associated clinical information were accessed from The Cancer Genome Atlas Liver Hepatocellular Carcinoma (TCGA-LIHC) database. Recurrence associated differentially expressed lncRNAs (DELncs) were determined by three DEG methods and two survival analyses methods. DELncs involved in the signature were selected by three machine learning methods and multivariate Cox analysis. Additionally, the signature was validated in a cohort of HCC patients from an external source. In order to gain insight into the biological functions of this signature, gene sets enrichment analyses, immune infiltration analyses, as well as immune and drug therapy prediction analyses were conducted. RESULTS A 4-lncRNA signature consisting of AC108463.1, AF131217.1, CMB9-22P13.1, TMCC1-AS1 was constructed. Patients in the high-risk group showed significantly higher early recurrence rate compared to those in the low-risk group. Combination of the signature, AFP and TNM further improved the early HCC recurrence predictive performance. Several molecular pathways and gene sets associated with HCC pathogenesis are enriched in the high-risk group. Antitumor immune cells, such as activated B cell, type 1 T helper cell, natural killer cell and effective memory CD8 T cell are enriched in patients with low-risk HCCs. HCC patients in the low- and high-risk group had differential sensitivities to various antitumor drugs. Finally, predictive performance of this signature was validated in an external cohort of patients with HCC. CONCLUSION Combined with TNM and AFP, the 4-lncRNA signature presents excellent predictability of HCC early recurrence.
Collapse
Affiliation(s)
- Yi Fu
- grid.507037.60000 0004 1764 1277Shanghai Key Laboratory of Molecular Imaging, Zhoupu Hospital, Shanghai University of Medicine and Health Sciences, Shanghai, China ,grid.507037.60000 0004 1764 1277Collaborative Innovation Center for Biomedicines, Shanghai University of Medicine and Health Sciences, Shanghai, China ,grid.507037.60000 0004 1764 1277School of Medical Instruments, Shanghai University of Medicine and Health Sciences, Shanghai, China
| | - Anfeng Si
- grid.41156.370000 0001 2314 964XDepartment of Surgical Oncology, Jinling Hospital, Medical School of Nanjing University, Nanjing, China
| | - Xindong Wei
- grid.412585.f0000 0004 0604 8558Central Laboratory, Department of Liver Diseases, Shuguang Hospital, Shanghai University of Chinese Traditional Medicine, Shanghai, China
| | - Xinjie Lin
- grid.507037.60000 0004 1764 1277Shanghai Key Laboratory of Molecular Imaging, Zhoupu Hospital, Shanghai University of Medicine and Health Sciences, Shanghai, China ,grid.507037.60000 0004 1764 1277Collaborative Innovation Center for Biomedicines, Shanghai University of Medicine and Health Sciences, Shanghai, China
| | - Yujie Ma
- grid.507037.60000 0004 1764 1277Shanghai Key Laboratory of Molecular Imaging, Zhoupu Hospital, Shanghai University of Medicine and Health Sciences, Shanghai, China ,grid.507037.60000 0004 1764 1277Collaborative Innovation Center for Biomedicines, Shanghai University of Medicine and Health Sciences, Shanghai, China
| | - Huimin Qiu
- grid.507037.60000 0004 1764 1277Collaborative Innovation Center for Biomedicines, Shanghai University of Medicine and Health Sciences, Shanghai, China ,grid.267139.80000 0000 9188 055XSchool of Health Science and Engineering, University of Shanghai for Science and Technology, Shanghai, China
| | - Zhinan Guo
- grid.507037.60000 0004 1764 1277Collaborative Innovation Center for Biomedicines, Shanghai University of Medicine and Health Sciences, Shanghai, China ,grid.412543.50000 0001 0033 4148School of Kinesiology, Shanghai University of Sport, Shanghai, China
| | - Yong Pan
- grid.268099.c0000 0001 0348 3990Department of Infectious Disease, Zhoushan Hospital, Wenzhou Medical University, Zhoushan, China
| | - Yiru Zhang
- grid.268099.c0000 0001 0348 3990Department of Infectious Disease, Zhoushan Hospital, Wenzhou Medical University, Zhoushan, China
| | - Xiaoni Kong
- grid.412585.f0000 0004 0604 8558Central Laboratory, Department of Liver Diseases, Shuguang Hospital, Shanghai University of Chinese Traditional Medicine, Shanghai, China
| | - Shibo Li
- Department of Infectious Disease, Zhoushan Hospital, Wenzhou Medical University, Zhoushan, China.
| | - Yanjun Shi
- Abdominal Transplantation Center, General Surgery, School of Medicine, Ruijin Hospital, Shanghai Jiao Tong University, Shanghai, China.
| | - Hailong Wu
- Shanghai Key Laboratory of Molecular Imaging, Zhoupu Hospital, Shanghai University of Medicine and Health Sciences, Shanghai, China. .,Collaborative Innovation Center for Biomedicines, Shanghai University of Medicine and Health Sciences, Shanghai, China. .,School of Health Science and Engineering, University of Shanghai for Science and Technology, Shanghai, China. .,School of Kinesiology, Shanghai University of Sport, Shanghai, China.
| |
Collapse
|
18
|
Li S, Chang M, Tong L, Wang Y, Wang M, Wang F. Screening potential lncRNA biomarkers for breast cancer and colorectal cancer combining random walk and logistic matrix factorization. Front Genet 2023; 13:1023615. [PMID: 36744179 PMCID: PMC9895102 DOI: 10.3389/fgene.2022.1023615] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Key Words] [Track Full Text] [Figures] [Journal Information] [Subscribe] [Scholar Register] [Received: 08/20/2022] [Accepted: 10/10/2022] [Indexed: 01/21/2023] Open
Abstract
Breast cancer and colorectal cancer are two of the most common malignant tumors worldwide. They cause the leading causes of cancer mortality. Many researches have demonstrated that long noncoding RNAs (lncRNAs) have close linkages with the occurrence and development of the two cancers. Therefore, it is essential to design an effective way to identify potential lncRNA biomarkers for them. In this study, we developed a computational method (LDA-RWLMF) by integrating random walk with restart and Logistic Matrix Factorization to investigate the roles of lncRNA biomarkers in the prognosis and diagnosis of the two cancers. We first fuse disease semantic and Gaussian association profile similarities and lncRNA functional and Gaussian association profile similarities. Second, we design a negative selection algorithm to extract negative LncRNA-Disease Associations (LDA) based on random walk. Third, we develop a logistic matrix factorization model to predict possible LDAs. We compare our proposed LDA-RWLMF method with four classical LDA prediction methods, that is, LNCSIM1, LNCSIM2, ILNCSIM, and IDSSIM. The results from 5-fold cross validation on the MNDR dataset show that LDA-RWLMF computes the best AUC value of 0.9312, outperforming the above four LDA prediction methods. Finally, we rank all lncRNA biomarkers for the two cancers after determining the performance of LDA-RWLMF, respectively. We find that 48 and 50 lncRNAs have the highest association scores with breast cancer and colorectal cancer among all lncRNAs known to associate with them on the MNDR dataset, respectively. We predict that lncRNAs HULC and HAR1A could be separately potential biomarkers for breast cancer and colorectal cancer and need to biomedical experimental validation.
Collapse
|
19
|
Peng L, Yang J, Wang M, Zhou L. Editorial: Machine learning-based methods for RNA data analysis—Volume II. Front Genet 2022; 13:1010089. [DOI: 10.3389/fgene.2022.1010089] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Track Full Text] [Journal Information] [Subscribe] [Scholar Register] [Received: 08/02/2022] [Accepted: 09/20/2022] [Indexed: 12/02/2022] Open
|
20
|
Su Q, Tan Q, Liu X, Wu L. Prioritizing potential circRNA biomarkers for bladder cancer and bladder urothelial cancer based on an ensemble model. Front Genet 2022; 13:1001608. [PMID: 36186429 PMCID: PMC9521272 DOI: 10.3389/fgene.2022.1001608] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Track Full Text] [Download PDF] [Figures] [Journal Information] [Subscribe] [Scholar Register] [Received: 07/23/2022] [Accepted: 08/15/2022] [Indexed: 12/03/2022] Open
Abstract
Bladder cancer is the most common cancer of the urinary system. Bladder urothelial cancer accounts for 90% of bladder cancer. These two cancers have high morbidity and mortality rates worldwide. The identification of biomarkers for bladder cancer and bladder urothelial cancer helps in their diagnosis and treatment. circRNAs are considered oncogenes or tumor suppressors in cancers, and they play important roles in the occurrence and development of cancers. In this manuscript, we developed an Ensemble model, CDA-EnRWLRLS, to predict circRNA-Disease Associations (CDA) combining Random Walk with restart and Laplacian Regularized Least Squares, and further screen potential biomarkers for bladder cancer and bladder urothelial cancer. First, we compute disease similarity by combining the semantic similarity and association profile similarity of diseases and circRNA similarity by combining the functional similarity and association profile similarity of circRNAs. Second, we score each circRNA-disease pair by random walk with restart and Laplacian regularized least squares, respectively. Third, circRNA-disease association scores from these models are integrated to obtain the final CDAs by the soft voting approach. Finally, we use CDA-EnRWLRLS to screen potential circRNA biomarkers for bladder cancer and bladder urothelial cancer. CDA-EnRWLRLS is compared to three classical CDA prediction methods (CD-LNLP, DWNN-RLS, and KATZHCDA) and two individual models (CDA-RWR and CDA-LRLS), and obtains better AUC of 0.8654. We predict that circHIPK3 has the highest association with bladder cancer and may be its potential biomarker. In addition, circSMARCA5 has the highest association with bladder urothelial cancer and may be its possible biomarker.
Collapse
|
21
|
Guo Z, Hui Y, Kong F, Lin X. Finding Lung-Cancer-Related lncRNAs Based on Laplacian Regularized Least Squares With Unbalanced Bi-Random Walk. Front Genet 2022; 13:933009. [PMID: 35938010 PMCID: PMC9355720 DOI: 10.3389/fgene.2022.933009] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Abstract] [Track Full Text] [Download PDF] [Figures] [Journal Information] [Subscribe] [Scholar Register] [Received: 04/30/2022] [Accepted: 06/03/2022] [Indexed: 11/13/2022] Open
Abstract
Lung cancer is one of the leading causes of cancer-related deaths. Thus, it is important to find its biomarkers. Furthermore, there is an increasing number of studies reporting that long noncoding RNAs (lncRNAs) demonstrate dense linkages with multiple human complex diseases. Inferring new lncRNA-disease associations help to identify potential biomarkers for lung cancer and further understand its pathogenesis, design new drugs, and formulate individualized therapeutic options for lung cancer patients. This study developed a computational method (LDA-RLSURW) by integrating Laplacian regularized least squares and unbalanced bi-random walk to discover possible lncRNA biomarkers for lung cancer. First, the lncRNA and disease similarities were computed. Second, unbalanced bi-random walk was, respectively, applied to the lncRNA and disease networks to score associations between diseases and lncRNAs. Third, Laplacian regularized least squares were further used to compute the association probability between each lncRNA-disease pair based on the computed random walk scores. LDA-RLSURW was compared using 10 classical LDA prediction methods, and the best AUC value of 0.9027 on the lncRNADisease database was obtained. We found the top 30 lncRNAs associated with lung cancers and inferred that lncRNAs TUG1, PTENP1, and UCA1 may be biomarkers of lung neoplasms, non-small–cell lung cancer, and LUAD, respectively.
Collapse
|
22
|
Peng L, Yang J, Wang M, Zhou L. Editorial: Machine Learning-Based Methods for RNA Data Analysis. Front Genet 2022; 13:828575. [PMID: 35692815 PMCID: PMC9175173 DOI: 10.3389/fgene.2022.828575] [Citation(s) in RCA: 0] [Impact Index Per Article: 0] [Reference Citation Analysis] [Track Full Text] [Download PDF] [Journal Information] [Subscribe] [Scholar Register] [Received: 12/03/2021] [Accepted: 04/12/2022] [Indexed: 11/13/2022] Open
Affiliation(s)
- Lihong Peng
- College of Life Sciences and Chemistry, Hunan University of Technology, Zhuzhou, China
- School of Computer, Hunan University of Technology, Zhuzhou, China
| | | | - Minxian Wang
- CAS Key Laboratory of Genome Sciences and Information, Beijing Institute of Genomics, Chinese Academy of Sciences, Beijing, China
- University of Chinese Academy of Sciences, Beijing, China
| | - Liqian Zhou
- College of Life Sciences and Chemistry, Hunan University of Technology, Zhuzhou, China
- *Correspondence: Liqian Zhou,
| |
Collapse
|