• bitcoinBitcoin(BTC)$76,827.00-1.91%
  • ethereumEthereum(ETH)$2,446.09-1.12%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$710.96-1.89%
  • rippleXRP(XRP)$1.34-3.58%
  • usd-coinUSDC(USDC)$1.000.02%
  • solanaSolana(SOL)$99.26-2.51%
  • tronTRON(TRX)$0.3404810.31%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.040.98%
  • zcashZcash(ZEC)$1,069.73-14.03%
  • HyperliquidHyperliquid(HYPE)$78.58-6.42%
  • dogecoinDogecoin(DOGE)$0.083436-2.96%
  • RainRain(RAIN)$0.015708-3.53%
  • USDSUSDS(USDS)$1.000.00%
  • moneroMonero(XMR)$507.24-1.23%
  • whitebitWhiteBIT Coin(WBT)$79.50-1.64%
  • chainlinkChainlink(LINK)$11.47-3.14%
  • leo-tokenLEO Token(LEO)$9.10-1.09%
  • cardanoCardano(ADA)$0.206666-2.98%
  • stellarStellar(XLM)$0.175568-2.60%
  • daiDai(DAI)$1.000.01%
  • bitcoin-cashBitcoin Cash(BCH)$226.30-10.14%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • USD1USD1(USD1)$1.00-0.01%
  • litecoinLitecoin(LTC)$52.81-0.36%
  • CantonCanton(CC)$0.097889-6.53%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.35-1.89%
  • uniswapUniswap(UNI)$5.98-0.73%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • hedera-hashgraphHedera(HBAR)$0.075031-2.60%
  • avalanche-2Avalanche(AVAX)$7.46-4.45%
  • nearNEAR Protocol(NEAR)$2.39-3.92%
  • suiSui(SUI)$0.73-4.60%
  • shiba-inuShiba Inu(SHIB)$0.000005-3.30%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.056395-2.93%
  • tether-goldTether Gold(XAUT)$4,322.72-2.06%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • MemeCoreMemeCore(M)$1.14-6.82%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • okbOKB(OKB)$109.07-3.64%
  • BittensorBittensor(TAO)$235.31-7.26%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.07%
  • mantleMantle(MNT)$0.58-4.26%
  • polkadotPolkadot(DOT)$1.11-0.44%
  • AsterAster(ASTER)$0.70-3.75%
  • aaveAave(AAVE)$121.50-3.04%
  • pax-goldPAX Gold(PAXG)$4,328.98-2.01%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.056614-0.12%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Revolutionizing AI’s Listening Skills: Tsinghua University and ByteDance Unveil SALMONN – A Groundbreaking Multimodal Neural Network for Advanced Audio Processing

November 4, 2023
in AI & Technology
Reading Time: 4 mins read
A A
Revolutionizing AI’s Listening Skills: Tsinghua University and ByteDance Unveil SALMONN – A Groundbreaking Multimodal Neural Network for Advanced Audio Processing
ShareShareShareShareShare

In several natural language processing applications, text-based big language models have shown impressive and even human-level performance. In the meanwhile, an LLM training paradigm known as instruction tuning—in which data is arranged as pairs of user instruction and reference response—has evolved that enables LLMs to comply with unrestricted user commands. Increasingly, researchers are interested in equipping LLMs with multimodal sensory skills. Current research focuses on linking LLMs to the encoder of one more input type—such as an image, silent video, audio event, or speech—or to the encoders of many input kinds together. 

To align the encoder output spaces with the LLM input space—which is often taught through cross-modal pre-training and instruction tuning—one can utilize a connection module and LLM adaptors. The speech audio language music open neural network that is proposed in this study is a single audio-text multimodal LLM that can recognize and comprehend speech, audio events, and music—the three main categories of sounds. SALMONN employs a dual encoder framework, comprising a BEATs audio encoder and a speech encoder from the Whisper speech model, to improve performance on both speech and nonspeech audio applications. 

To further enhance Vicuna’s performance, the low-rank adaption strategy is utilized as a cross-modal adaptor to match the augmented input space with the output space. The cross-modal pre-training and instruction tuning phases of the window-level Q-Former and LoRA employ many speech, audio, and music challenges. The resultant multimodal LLMs show little to no cross-modal emergent skills and can be restricted to the specific kinds of tasks utilized in instruction tuning, specifically audio captioning and voice recognition, which they term the task over-fitting problem. The ability to execute cross-modal tasks that are not noticed during training is referred to in this study as cross-modal emergent skills. These abilities are basically the emergent capabilities of LLMs that are lost during instruction tailoring. 

In order to mitigate the significant catastrophic forgetting of the training tasks, they suggest adding an additional few-shot activation tuning stage to SALMONN’s repertoire. SALMONN’s cognitive hearing abilities are assessed using a variety of speech, auditory events, and music standards. There are three levels to the tasks. The first two levels test untrained activities, while the first level benchmarks eight tasks that are taught in instruction tuning, including audio captioning, translation, and voice recognition. Five speech-based natural language processing (NLP) tasks, including slot filling and translation to untrained languages, are included in the second level. These tasks need multilingual and high-quality alignments between voice and text tokens. 

Comprehending non-speech auditory information is necessary for the last set of activities, such as audio-based narrative and speech audio co-reasoning. The results of the experiments demonstrate that SALMONN can complete all of these tasks and perform competitively on industry benchmarks when used as a single model. This suggests that it is possible to create artificial intelligence that is capable of “hearing” and comprehending a wide variety of audio inputs, including speech, audio events, and music. 

This paper’s primary contribution may be summed up as follows. 

• To the best of their knowledge, researchers from Tsinghua University and ByteDance offer SALMONN, the first multimodal LLM that can recognize and comprehend general audio inputs including voice, audio events, and music. 

• By varying the LoRA scaling factor, they investigate the existence of cross-modal emergent skills. They then suggest a low-cost activation tuning technique as an additional training step that can activate these abilities and reduce catastrophic forgetting to tasks encountered during training. 

• They provide two new tasks, audio-based storytelling and spoken audio co-reasoning, and assess SALMONN on a variety of tasks that represent a range of general hearing skills.


Check out the Paper and Github. All Credit For This Research Goes To the Researchers on This Project. Also, don’t forget to join our 32k+ ML SubReddit, 40k+ Facebook Community, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.

If you like our work, you will love our newsletter..


YOU MAY ALSO LIKE

How These XL Phones Compete

CA Governor Signs ‘Landmark’ Laws On Youth Use Of Social Media And AI Chatbots

Aneesh Tickoo is a consulting intern at MarktechPost. He is currently pursuing his undergraduate degree in Data Science and Artificial Intelligence from the Indian Institute of Technology(IIT), Bhilai. He spends most of his time working on projects aimed at harnessing the power of machine learning. His research interest is image processing and is passionate about building solutions around it. He loves to connect with people and collaborate on interesting projects.


🔥 Meet Retouch4me: A Family of Artificial Intelligence-Powered Plug-Ins for Photography Retouching

Credit: Source link

ShareTweetSendSharePin

Related Posts

How These XL Phones Compete
AI & Technology

How These XL Phones Compete

September 10, 2026
CA Governor Signs ‘Landmark’ Laws On Youth Use Of Social Media And AI Chatbots
AI & Technology

CA Governor Signs ‘Landmark’ Laws On Youth Use Of Social Media And AI Chatbots

September 10, 2026
Meet Redis LangCache: A Managed Semantic Cache That Cuts LLM API Costs by Up to 90% and Returns Cache Hits Up to 15x Faster
AI & Technology

Meet Redis LangCache: A Managed Semantic Cache That Cuts LLM API Costs by Up to 90% and Returns Cache Hits Up to 15x Faster

September 10, 2026
Meta Is Testing Community Notes In Latin America. Fact Checkers Are Worried.
AI & Technology

Meta Is Testing Community Notes In Latin America. Fact Checkers Are Worried.

September 10, 2026
Next Post
Israel strikes ambulance at Al-Shifa Hospital in Gaza City

Israel strikes ambulance at Al-Shifa Hospital in Gaza City

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Lightfield Raises M Series A Led by a16z to Accelerate Growth – Unite.AI

Lightfield Raises $47M Series A Led by a16z to Accelerate Growth – Unite.AI

September 9, 2026
Tropical Storm Bertha makes landfall in Louisiana

Tropical Storm Bertha makes landfall in Louisiana

September 6, 2026
Tooley Oil is selling its NorCal convenience stores and car washes

Tooley Oil is selling its NorCal convenience stores and car washes

September 8, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!