Physical AI HBM Smart Factory SDV AIoT Power Semicon 특수 가스 정정·반론보도 모음 e4ds plus

Smart speaker can achieve clear voice recognition without deep learning

Google 우선 소스 기사입력2019.12.03 17:47

Multi-channel device that increases voice recognition rate and reduces corpus construction rate
ADC, capture and amplification technology makes 1:1 conversations and even long-distance listening a reality

“Alexa, buy me a dollhouse.”

A six-year-old in the US asked her Amazon Echo to buy her the dollhouse and cookies she wanted. The device did as she said, and a few days later, a four-pound bag of cookies and a $170 KidzKraft dollhouse were delivered to her home.

The bigger story happened next. The anchor of the local news station where the child lives was reporting this incident, and as his last comment, he said, “Alexa, buy me a dollhouse too.” Then, all the Amazon Echos in the house listening to the news started ordering dollhouses.

There was an uproar out of nowhere where people who had ordered dollhouses started canceling their orders, and Amazon canceled all dollhouse orders.

For those of us who don't yet use smart speakers as everyday friends, this may seem like a strange sight, but as a conversational platform, smart speakers are the fastest-growing segment of commercial products.

According to the '2018 CIO Survey' published by market research firm Gartner, 4% of the organizations that responded are already investing in conversational platform technology and utilizing conversational interfaces. Seventeen percent of respondents said they are pursuing or actively experimenting with conversational platform technologies in the short term.

Gartner also predicted that the AI speaker market size will reach 2.1 billion dollars in 2020, and in fact, manufacturers, IT companies, and mobile carriers are actively participating in the AI speaker market to secure market share.

AI speakers are noteworthy because they have value as connected devices that can be linked to home IoT, vehicles, etc. The key to the success of this business is the expansion of speaker functions. If AI is the soul, then the speaker can be seen as the body.

The domestic market has been formed around mobile carriers and portal operators, and KT and SK Telecom have approached consumers in a familiar way by linking with IPTV. Recently, there has been a trend of adding visual information to smart speakers to overcome the limitations of voice recognition.

According to a survey by Voicebot.ai, a media outlet specializing in smart speakers, 13.2% of smart speaker owners in the U.S. in 2019 owned smart speakers with added display functions. However, just as voice call usage remains strong despite the advancement of video call technology, it is difficult to rule out the possibility that voice-based speakers will maintain their influence.

It also provides a new perspective on the fact that the reason why AI speakers react differently from human speech is because of deep learning. What if, instead of using deep learning-based system algorithms, we could capture the human voice itself like a photo and remember it?

An alternative could be an ADC (Analog to Digital Converter) device that captures the human voice, which is an analog sound, clearly in any situation through a microphone and then transmits it as a sound signal to a digital device for storage.
Avi Mufi TI Audio Codec & Converter
Product Marketing Manager (Photo = Reporter Lee Su-min)

I met with the engineer behind a device that accurately captures and understands voice commands even in noisy environments, even when the speaker's voice is soft. That engineer is Abhi Muppiri, an audio product marketing engineer at TI.


- I heard there is a product that delivers clearer audio quality from a distance four times farther than existing competing devices. Could you please introduce the product?

The new audio analog-to-digital converter (ADC) is the TLV320ADC5140. It is the smallest 4-channel audio ADC in its class that delivers the same performance in the industry.

The device consists of three TI Burr-Brown audio ADCs, enabling long-distance, high-fidelity recording performance in any environment, as well as low-distortion audio recording even in noisy environments.

It is designed to enhance the far-field audio capture capability and better detect low-volume commands for applications such as high-end speakers, sound bars, wireless speakers, high-definition TVs, IP network cameras, video conferencing systems, and smart home appliances.

The TLV320ADC5140 also has the advantage of reducing the number of microphones in the array, thereby reducing system cost.


- When the transmission strength is weakened due to the distance between the transmitter and receiver, surrounding obstacles, etc., the signal-to-noise ratio (S/N) value also decreases. How can it have a signal-to-noise ratio of over 106 dB while also having low-distortion audio recording and hi-fi recording functions?

As the distance between the user and the device increases, the sound signal naturally weakens, but the TLV320ADC5140 is equipped with a Dynamic Range Enhancer (DRE) to solve this problem.

DRE enhances low-level audio signals at the system level and maintains low-distortion recordings even close to speaker outputs, as well as enabling long-distance hi-fi recording in all environments.

In other words, even if a very weak sound signal enters the device, the user's voice can be accurately captured by repeating the step of amplifying this signal.

The biggest feature of the TLV320ADC5140 product is that it relies on the microphone to capture sound clearly and accurately. The hardware for basic performance is built into the device, and the user can control the level of sound sensitivity through software.

This multi-channel product can install 4-channel analog microphones or 8-channel digital PDM microphones, and can also use a combination of analog and digital microphones.


- One of the major features of this product is that it is the smallest 4-channel audio in the product line that implements the same performance. How were you able to create an ADC of such a small size? If you increased the number and size of the transmitting antennas to increase the S/R, the physical ratio within the ADC would also increase, right?

The TLV320ADC5140 relies primarily on the microphone rather than the transmitting antenna when receiving audio signals. Since it captures and converts sound, it considers the microphone as well as the ADC as important.

We use the latest process technology to control the size because the analog microphone has 4 microphones and the digital microphone has 8 microphones. Analog microphones are usually used in high-end, and digital microphones are used in laptops or PCs.

However, with the advent of smart speakers in 2018, the signal-to-voice ratio of digital microphones has improved. As a result, products that were previously only available in high-end models can now be installed in laptops and PCs, and can handle up to 110 dB.

For example, in the case of digital microphones, the biggest advantage is that they consume less power, and we can choose IP cameras that we install for security purposes. The system is designed to work together with the digital microphones to activate when someone breaks into your home.

Since the analog microphone has good sound detection (audio quality), the TLV320ADC5140 product can be customized by combining analog and digital to suit the customer's needs. It is a device multi-channel input type.


- You said that you can record low-distortion audio from a distance even in a noisy environment. What is the technology for targeting the desired sound?

One of the main roles of the ADC is to deliver the sound signal to the algorithm, so that it can accurately and cleanly capture the desired sound even in a noisy environment.

This is a technology that passes the input sound through recording to an algorithm and the algorithm decides which user's sound to select. The algorithm accurately identifies who is speaking and what the content is within the transmitted signal.


- The TLV320ADC5140 product has the advantage of being small in size, but in the case of electronic products, if the size is smaller, the integration level increases, but doesn't the heat generated actually increase? How did you solve the heat generation problem?

It is true that as performance is integrated, more heat is generated. To solve this problem, TI's TLV320ADC5140 is designed in a way that integrates well with process technology.

We applied a process that maximizes performance while minimizing power consumption, and designed it together with peripherals during system design so that when power consumption or heat generation issues occur, it can be turned on and off appropriately through peripherals.

The current ADC power consumption is 9mW per channel. In particular, this product is designed to consume less power than existing models.

TI has already proven that there is no problem with the heat generation issue through the TLV320ADC3101 product, which is installed in most AI speakers.


- Last December, Karthik Vasanth, TI’s vice president of data converters, announced that TI would invest heavily in both size reduction and speed improvement. Now, a year later, how much has the speed of the newly released TLV320ADC5140 improved?

The TLV320ADC5140 has a sampling rate of 768KHz. This is the fastest sampling rate on the market and is a specification that can be used in current high-performance smart speakers.

Considering that the sampling speed of the previous generation product was 192KHz, it can be said that it is implementing a speed that is 4 times faster than before.


- According to the 'TI Audio Innovation: Trends in Automotive, Smart Home and Pro Audio Applications' data, the microphone array can instantly determine who is speaking, and only the microphone facing the speaker is turned on, while all other microphones are turned off to reduce excessive noise generation. Can you explain what technology is used to instantly determine who is speaking?

It is implemented through a system algorithm. Stereo Acoustic Echo Cancellation (AEC) and Beamforming technology can determine which speaker is speaking and from which direction the sound is coming.

In other words, the core technology of the TI TLV320ADC5140 is to cleanly capture the desired sound regardless of any noise and transmit it to the DSP.


- Could you briefly introduce the features of TI products for Korean engineers?

Korean companies are among the first to adopt the relevant technology. TI's TLV320ADC5140 has a wide range of applications, as its main feature is to capture and transmit audio in a clean state.

Above all, it can be designed with the customer's desired custom combination, so it can be used for all required applications.


Abhi Muppiri is a veteran engineer with over 10 years of experience in systems and applications engineering and design. He is currently responsible for customer engagement, content creation, demand generation and pricing strategies for audio ADC and audio codec investments for personal electronics and industrial applications.
#TI
최인영 기자
기사 전체보기