This page was machine-translated and may differ from the original. View original
Nvidia presented research achievements based on DGX systems by researchers from Nvidia's Artificial Intelligence Lab (NVAIL) at the International Conference on Machine Learning (ICML) 2017, which is being held in Sydney from the 6th to 11th (local time).
NVAIL, operated by Nvidia, brings together leading universities and research institutes from around the world. In particular, research teams from University of California Berkeley, Swiss AI Research Institute IDSIA, and University of Tokyo are leading advances in the field of deep learning based on Nvidia DGX, the world's first artificial intelligence supercomputer.
In current artificial intelligence approaches, robots learn optimal response methods to stimuli through repeated work experience. Professor Levin explained that if robots could learn without such repetitive tasks, not only would robot adaptability improve, but they could also learn much more.
He stated, "Robots must go through thousands of training iterations to learn a single skill. If we could dramatically reduce the number of experiences required for such learning, we could learn thousands of skills with the same number of work iterations that previously took to learn just one skill," and added, "While it is difficult to build a machine that makes no mistakes, it is possible to build a machine that learns more quickly from mistakes, thereby reducing the number of mistakes that must be experienced."
The research team led by Professor Levin is using Nvidia DGX systems to train algorithms that adjust visual recognition and movement. Dr. Chelsea Finn, a doctoral candidate in Professor Levin's research group, presented a research paper on this work at this conference and provided a tutorial on "Deep Reinforcement Learning, Decision Making, and Control" together with Professor Levin.

Swiss AI Research Institute IDSIA: Advanced Deep Learning
The combination of Recurrent Neural Networks (RNN) and Long Short-Term Memory (LSTM) has had a significant impact on researchers in handwriting and speech recognition.
Unlike feedforward networks that automatically pass each computation to the next stage, RNNs can use internal memory to process arbitrary data such as different pronunciations or handwriting variations, and immediately leverage previous decisions and current stimuli in learning.
This means that as RNNs advance further as neural networks, they become increasingly difficult to handle and slow down the deep learning process. Researchers at Swiss AI Research Institute IDSIA provided a solution to this problem through recurrent highway networks.
Rupesh Srivastava, an AI researcher at IDSIA and co-author of a research paper in this field presented at ICML, explained, "Until now, training recurrent networks was very difficult even in situations where layers doubled in sequential propagation, whereas thanks to recurrent highway networks, it is now possible to smoothly train recurrent networks even when layers increase to ten in repeated propagation."
The Srivastava research team enhanced training speed by utilizing Nvidia Tesla K40, K80, TITAN X, and GeForce GTX 1080 GPUs along with CUDA and cuDNN for deep learning. Srivastava noted that with the adoption of the DGX artificial intelligence supercomputer, "experiment cycles have been significantly accelerated, and progress on all lab projects has become very rapid."
University of Tokyo: Deep Learning Deception
By leveraging the capabilities of DGX, "pseudo-labels" were assigned to unlabeled data in target domains. The team reported developing methods to avoid various challenges in autonomous domain adaptation. This allows deep learning models to apply what they learn from a source domain, such as the classification ability of book reviews, to a completely different target domain, such as movie reviews, without needing to train a new model.
The University of Tokyo research team proposed a concept called "asymmetric tri-training." This concept assigns different roles to three classifiers and utilizes three different neural networks. Two networks are used to assign labels to unlabeled target samples, while the remaining network conducts training with target samples that have been assigned pseudo-labels.
The related research paper with Professor Harada as a co-author was presented at ICML. Professor Harada stated, "Similar efforts must be made to realize the potential, and by sharing this research, we hope that related research will proceed more rapidly."
NVAIL, operated by Nvidia, brings together leading universities and research institutes from around the world. In particular, research teams from University of California Berkeley, Swiss AI Research Institute IDSIA, and University of Tokyo are leading advances in the field of deep learning based on Nvidia DGX, the world's first artificial intelligence supercomputer.
In current artificial intelligence approaches, robots learn optimal response methods to stimuli through repeated work experience. Professor Levin explained that if robots could learn without such repetitive tasks, not only would robot adaptability improve, but they could also learn much more.
He stated, "Robots must go through thousands of training iterations to learn a single skill. If we could dramatically reduce the number of experiences required for such learning, we could learn thousands of skills with the same number of work iterations that previously took to learn just one skill," and added, "While it is difficult to build a machine that makes no mistakes, it is possible to build a machine that learns more quickly from mistakes, thereby reducing the number of mistakes that must be experienced."
The research team led by Professor Levin is using Nvidia DGX systems to train algorithms that adjust visual recognition and movement. Dr. Chelsea Finn, a doctoral candidate in Professor Levin's research group, presented a research paper on this work at this conference and provided a tutorial on "Deep Reinforcement Learning, Decision Making, and Control" together with Professor Levin.
Swiss AI Research Institute IDSIA: Advanced Deep Learning
The combination of Recurrent Neural Networks (RNN) and Long Short-Term Memory (LSTM) has had a significant impact on researchers in handwriting and speech recognition.
Unlike feedforward networks that automatically pass each computation to the next stage, RNNs can use internal memory to process arbitrary data such as different pronunciations or handwriting variations, and immediately leverage previous decisions and current stimuli in learning.
This means that as RNNs advance further as neural networks, they become increasingly difficult to handle and slow down the deep learning process. Researchers at Swiss AI Research Institute IDSIA provided a solution to this problem through recurrent highway networks.
Rupesh Srivastava, an AI researcher at IDSIA and co-author of a research paper in this field presented at ICML, explained, "Until now, training recurrent networks was very difficult even in situations where layers doubled in sequential propagation, whereas thanks to recurrent highway networks, it is now possible to smoothly train recurrent networks even when layers increase to ten in repeated propagation."
The Srivastava research team enhanced training speed by utilizing Nvidia Tesla K40, K80, TITAN X, and GeForce GTX 1080 GPUs along with CUDA and cuDNN for deep learning. Srivastava noted that with the adoption of the DGX artificial intelligence supercomputer, "experiment cycles have been significantly accelerated, and progress on all lab projects has become very rapid."
University of Tokyo: Deep Learning Deception
By leveraging the capabilities of DGX, "pseudo-labels" were assigned to unlabeled data in target domains. The team reported developing methods to avoid various challenges in autonomous domain adaptation. This allows deep learning models to apply what they learn from a source domain, such as the classification ability of book reviews, to a completely different target domain, such as movie reviews, without needing to train a new model.
The University of Tokyo research team proposed a concept called "asymmetric tri-training." This concept assigns different roles to three classifiers and utilizes three different neural networks. Two networks are used to assign labels to unlabeled target samples, while the remaining network conducts training with target samples that have been assigned pseudo-labels.
The related research paper with Professor Harada as a co-author was presented at ICML. Professor Harada stated, "Similar efforts must be made to realize the potential, and by sharing this research, we hope that related research will proceed more rapidly."
To request a correction, reply or follow-up report on this article, see how to file a request. Previously published statements are collected in corrections & replies.
김자영 Reporter













