• bitcoinBitcoin(BTC)$77,332.00-0.08%
  • ethereumEthereum(ETH)$2,530.462.10%
  • tetherTether(USDT)$1.000.02%
  • binancecoinBNB(BNB)$733.792.69%
  • rippleXRP(XRP)$1.370.83%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$101.671.66%
  • tronTRON(TRX)$0.3394020.24%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.02-1.77%
  • zcashZcash(ZEC)$1,147.442.54%
  • HyperliquidHyperliquid(HYPE)$79.24-1.31%
  • dogecoinDogecoin(DOGE)$0.0847040.78%
  • RainRain(RAIN)$0.015138-3.77%
  • moneroMonero(XMR)$544.926.88%
  • USDSUSDS(USDS)$1.000.00%
  • whitebitWhiteBIT Coin(WBT)$80.380.27%
  • chainlinkChainlink(LINK)$11.540.29%
  • leo-tokenLEO Token(LEO)$9.120.37%
  • cardanoCardano(ADA)$0.2088370.36%
  • stellarStellar(XLM)$0.1809942.56%
  • bitcoin-cashBitcoin Cash(BCH)$231.581.83%
  • Ethena USDeEthena USDe(USDE)$1.000.04%
  • daiDai(DAI)$1.00-0.01%
  • USD1USD1(USD1)$1.000.01%
  • litecoinLitecoin(LTC)$54.101.81%
  • uniswapUniswap(UNI)$6.384.76%
  • CantonCanton(CC)$0.0987960.03%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.382.05%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • avalanche-2Avalanche(AVAX)$7.45-0.79%
  • hedera-hashgraphHedera(HBAR)$0.074464-0.22%
  • nearNEAR Protocol(NEAR)$2.37-4.42%
  • shiba-inuShiba Inu(SHIB)$0.0000053.09%
  • suiSui(SUI)$0.73-1.45%
  • crypto-com-chainCronos(CRO)$0.0576041.27%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • MemeCoreMemeCore(M)$1.19-0.31%
  • tether-goldTether Gold(XAUT)$4,349.93-0.12%
  • Circle USYCCircle USYC(USYC)$1.140.03%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • okbOKB(OKB)$114.312.91%
  • BittensorBittensor(TAO)$235.60-0.34%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.11%
  • aaveAave(AAVE)$126.392.81%
  • mantleMantle(MNT)$0.58-2.08%
  • pax-goldPAX Gold(PAXG)$4,355.30-0.09%
  • AsterAster(ASTER)$0.69-2.81%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0567102.00%
  • polkadotPolkadot(DOT)$1.05-6.28%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

This AI Paper from Google Unveils a Groundbreaking Non-Autoregressive, LM-Fused ASR System for Superior Multilingual Speech Recognition

January 29, 2024
in AI & Technology
Reading Time: 5 mins read
A A
This AI Paper from Google Unveils a Groundbreaking Non-Autoregressive, LM-Fused ASR System for Superior Multilingual Speech Recognition
ShareShareShareShareShare

The evolution of technology in speech recognition has been marked by significant strides, but challenges like latency the time delay in processing spoken language, have continually impeded progress. This latency is especially pronounced in autoregressive models, which process speech sequentially, leading to delays. These delays are detrimental in real-time applications like live captioning or virtual assistants, where immediacy is key. Addressing this latency without compromising accuracy remains critical in advancing speech recognition technology.

A pioneering approach in speech recognition is developing a non-autoregressive model, a departure from traditional methods. This model, proposed by a team of researchers from Google Research, is designed to tackle the inherent latency issues found in existing systems. It utilizes large language models and leverages parallel processing, which processes speech segments simultaneously rather than sequentially. This similar processing approach is instrumental in reducing latency, offering a more fluid and responsive user experience.

The core of this innovative model is the fusion of the Universal Speech Model (USM) with the PaLM 2 language model. The USM, a robust model with 2 billion parameters, is designed for accurate speech recognition. It uses a vocabulary of 16,384-word pieces and employs a Connectionist Temporal Classification (CTC) decoder for parallel processing. The USM is trained on an extensive dataset, encompassing over 12 million hours of unlabeled audio and 28 billion sentences of text data, making it incredibly adept at handling multilingual inputs.

The PaLM 2 language model, known for its prowess in natural language processing, complements the USM. It’s trained on diverse data sources, including web documents and books, and employs a large 256,000 wordpiece vocabulary. The model stands out for its ability to score Automatic Speech Recognition (ASR) hypotheses using a prefix language model scoring mode. This method involves prompting the model with a fixed prefix—top hypotheses from previous segments—and scoring several suffix hypotheses for the current segment. 

In practice, the combined system processes long-form audio in 8-second chunks. As soon as the audio is available, the USM encodes it, and these segments are then relayed to the CTC decoder. The decoder forms a confusion network lattice encoding possible word pieces, which the PaLM 2 model scores. The system updates every 8 seconds, providing a near real-time response.

The performance of this model was rigorously evaluated across several languages and datasets, including YouTube captioning and the FLEURS test set. The results were remarkable. An average improvement of 10.8% in relative word error rate (WER) was observed on the multilingual FLEURS test set. For the YouTube captioning dataset, which presents a more challenging scenario, the model achieved an average improvement of 3.6% across all languages. These improvements are a testament to the model’s effectiveness in diverse languages and settings.

The study delved into various factors affecting the model’s performance. It explored the impact of language model size, ranging from 128 million to 340 billion parameters. It found that while larger models reduced sensitivity to fusion weight, the gains in WER might not offset the increasing inference costs. The optimal LLM scoring weight also shifted with model size, suggesting a balance between model complexity and computational efficiency.

In conclusion, this research presents a significant leap in speech recognition technology. Its highlights include:

  • A non-autoregressive model combining the USM and PaLM 2 for reduced latency.
  • Enhanced accuracy and speed, making it suitable for real-time applications.
  • Significant improvements in WER across multiple languages and datasets.

This model’s innovative approach to processing speech in parallel, coupled with its ability to handle multilingual inputs efficiently, makes it a promising solution for various real-world applications. The insights provided into system parameters and their effects on ASR efficacy add valuable knowledge to the field, paving the way for future advancements in speech recognition technology. 


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. Join our 36k+ ML SubReddit, 41k+ Facebook Community, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our Telegram Channel


YOU MAY ALSO LIKE

Kai-Fu Lee Says China Will Win AI Reach Race

Everybody’s Business: Unpacking Apple’s Upcoming Launches

Hello, My name is Adnan Hassan. I am a consulting intern at Marktechpost and soon to be a management trainee at American Express. I am currently pursuing a dual degree at the Indian Institute of Technology, Kharagpur. I am passionate about technology and want to create new products that make a difference.


🧑‍💻 [FREE AI WEBINAR] ‘Build Real-Time Document/Image Analytics with GPT-4 Vision’ (Jan 29, 2024)


Credit: Source link

ShareTweetSendSharePin

Related Posts

Kai-Fu Lee Says China Will Win AI Reach Race
AI & Technology

Kai-Fu Lee Says China Will Win AI Reach Race

September 12, 2026
Everybody’s Business: Unpacking Apple’s Upcoming Launches
AI & Technology

Everybody’s Business: Unpacking Apple’s Upcoming Launches

September 12, 2026
Why Laser Beams Are the Hottest New Tech in Defense
AI & Technology

Why Laser Beams Are the Hottest New Tech in Defense

September 12, 2026
Why Amazon Is Diversifying Its AI Chip Supply
AI & Technology

Why Amazon Is Diversifying Its AI Chip Supply

September 12, 2026
Next Post
Business First Bancshares: TBV Per Share Has Stalled Since 2019 (NASDAQ:BFST)

Business First Bancshares: TBV Per Share Has Stalled Since 2019 (NASDAQ:BFST)

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Teen pleads guilty to deadly 2024 Georgia high school shooting

Teen pleads guilty to deadly 2024 Georgia high school shooting

September 5, 2026
Conservative creators claim TikTok is heavily censoring them

Conservative creators claim TikTok is heavily censoring them

September 10, 2026
Breaking Down Apple’s First Foldable iPhone With Mark Gurman

Breaking Down Apple’s First Foldable iPhone With Mark Gurman

September 12, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!