This page was machine-translated and may differ from the original. View original
KAIST-UCSD Research Team Uses AI to Analyze Genetic Information
Predicting transcription factors that regulate the copying process
Development of 'DeeptiFactor' for rapid protein sequence analysis
A joint research team led by Professor Sang-Yeop Lee of the Department of Chemical and Biomolecular Engineering at the Korea Advanced Institute of Science and Technology (KAIST) and Professor Bernhard Palsson of the Department of Biomolecular Engineering at the University of California, San Diego (UCSD) announced on the 29th that they have developed a 'DeepTFactor' system that uses AI to predict transcription factors that control gene transcription (the process of copying genetic information).
Analyzing gene transcription mediated by transcription factors can help us understand how organisms respond to genetic or environmental changes to control gene expression. In this sense, identifying an organism's transcription factors is the first step toward analyzing its transcriptional regulatory system.
Until now, to find new transcription factors, we have analyzed homology (similar properties) with already known transcription factors or used data-based approaches such as machine learning.
To utilize existing machine learning models, it is necessary to rely on specialized knowledge of the problem to be solved, such as calculating the physicochemical properties of molecules or analyzing the homology of biological sequences, to identify features to use as input values for the model.
Deep learning has recently been utilized in various fields of biology because it can inherently learn potential features for problem solving. However, in the case of prediction systems using deep learning, the inference process cannot be directly confirmed due to the complex calculations within the system.
.jpg)
DeeptiFactor, developed by the joint research team, utilizes three parallel convolutional neural networks (CNNs) to predict transcription factors from protein sequences. Using DeeptiFactor, the joint research team predicted 332 transcription factors in Escherichia coli and identified the genome-wide binding sites of three of them, validating DeeptiFactor's performance.
Furthermore, the joint research team used a saliency map-based deep learning model interpretation methodology to understand DeeptiFactor's inference process. This methodology confirmed that, although information about the DNA binding region of transcription factors is not explicitly provided during DeeptiFactor's learning process, it implicitly learns this information and utilizes it for prediction.
Professor Lee Sang-yeop said, “Using DeeptiFactor, we can quickly analyze newly discovered protein sequences and numerous protein sequences that have not yet been characterized,” and “This will be utilized as a basic technology for analyzing the electronic regulatory networks of organisms.”
Meanwhile, this study was conducted with support from the Ministry of Science and ICT's Climate Change Response Technology Development Project's System Metabolic Engineering for Biorefinery Source Technology Development Project. Additionally, it was published on December 28th in the international academic journal Proceedings of the National Academy of Sciences of the United States of America (PNAS) under the title 'DeepTFactor: A deep learning-based tool for the prediction of transcription factors.'
Predicting transcription factors that regulate the copying process
Development of 'DeeptiFactor' for rapid protein sequence analysis
A joint research team led by Professor Sang-Yeop Lee of the Department of Chemical and Biomolecular Engineering at the Korea Advanced Institute of Science and Technology (KAIST) and Professor Bernhard Palsson of the Department of Biomolecular Engineering at the University of California, San Diego (UCSD) announced on the 29th that they have developed a 'DeepTFactor' system that uses AI to predict transcription factors that control gene transcription (the process of copying genetic information).
Analyzing gene transcription mediated by transcription factors can help us understand how organisms respond to genetic or environmental changes to control gene expression. In this sense, identifying an organism's transcription factors is the first step toward analyzing its transcriptional regulatory system.
Until now, to find new transcription factors, we have analyzed homology (similar properties) with already known transcription factors or used data-based approaches such as machine learning.
To utilize existing machine learning models, it is necessary to rely on specialized knowledge of the problem to be solved, such as calculating the physicochemical properties of molecules or analyzing the homology of biological sequences, to identify features to use as input values for the model.
Deep learning has recently been utilized in various fields of biology because it can inherently learn potential features for problem solving. However, in the case of prediction systems using deep learning, the inference process cannot be directly confirmed due to the complex calculations within the system.
▲ Network structure of a deep learning model for transcription factor prediction
[Image = KAIST]
[Image = KAIST]
DeeptiFactor, developed by the joint research team, utilizes three parallel convolutional neural networks (CNNs) to predict transcription factors from protein sequences. Using DeeptiFactor, the joint research team predicted 332 transcription factors in Escherichia coli and identified the genome-wide binding sites of three of them, validating DeeptiFactor's performance.
Furthermore, the joint research team used a saliency map-based deep learning model interpretation methodology to understand DeeptiFactor's inference process. This methodology confirmed that, although information about the DNA binding region of transcription factors is not explicitly provided during DeeptiFactor's learning process, it implicitly learns this information and utilizes it for prediction.
Professor Lee Sang-yeop said, “Using DeeptiFactor, we can quickly analyze newly discovered protein sequences and numerous protein sequences that have not yet been characterized,” and “This will be utilized as a basic technology for analyzing the electronic regulatory networks of organisms.”
Meanwhile, this study was conducted with support from the Ministry of Science and ICT's Climate Change Response Technology Development Project's System Metabolic Engineering for Biorefinery Source Technology Development Project. Additionally, it was published on December 28th in the international academic journal Proceedings of the National Academy of Sciences of the United States of America (PNAS) under the title 'DeepTFactor: A deep learning-based tool for the prediction of transcription factors.'
본 기사에 대한 정정·반론·추후보도 청구는 보도 청구 안내를, 그간 게재된 보도문은 정정·반론보도 모아보기를 참고해 주세요.














