• bitcoinBitcoin(BTC)$78,923.000.45%
  • ethereumEthereum(ETH)$2,495.970.95%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$751.210.09%
  • rippleXRP(XRP)$1.432.82%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$104.011.03%
  • tronTRON(TRX)$0.3385810.45%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.040.00%
  • zcashZcash(ZEC)$1,214.207.81%
  • HyperliquidHyperliquid(HYPE)$85.792.04%
  • dogecoinDogecoin(DOGE)$0.0900830.52%
  • RainRain(RAIN)$0.016045-1.19%
  • USDSUSDS(USDS)$1.000.01%
  • whitebitWhiteBIT Coin(WBT)$81.747.49%
  • moneroMonero(XMR)$505.42-2.25%
  • chainlinkChainlink(LINK)$12.48-1.44%
  • leo-tokenLEO Token(LEO)$9.19-0.46%
  • cardanoCardano(ADA)$0.2187990.65%
  • stellarStellar(XLM)$0.189018-0.74%
  • bitcoin-cashBitcoin Cash(BCH)$258.45-0.02%
  • daiDai(DAI)$1.00-0.01%
  • Ethena USDeEthena USDe(USDE)$1.000.02%
  • CantonCanton(CC)$0.1088002.69%
  • USD1USD1(USD1)$1.000.00%
  • uniswapUniswap(UNI)$6.85-2.56%
  • litecoinLitecoin(LTC)$54.02-1.83%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.390.25%
  • hedera-hashgraphHedera(HBAR)$0.079048-2.52%
  • avalanche-2Avalanche(AVAX)$7.98-1.02%
  • suiSui(SUI)$0.81-0.83%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • shiba-inuShiba Inu(SHIB)$0.0000050.11%
  • nearNEAR Protocol(NEAR)$2.321.63%
  • crypto-com-chainCronos(CRO)$0.0597563.17%
  • paypal-usdPayPal USD(PYUSD)$1.000.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,397.68-0.41%
  • MemeCoreMemeCore(M)$1.180.63%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BittensorBittensor(TAO)$256.460.36%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • okbOKB(OKB)$114.24-0.84%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.01%
  • mantleMantle(MNT)$0.631.12%
  • AsterAster(ASTER)$0.76-0.93%
  • polkadotPolkadot(DOT)$1.1911.21%
  • aaveAave(AAVE)$129.00-1.94%
  • pax-goldPAX Gold(PAXG)$4,401.73-0.43%
  • Pump.funPump.fun(PUMP)$0.0044604.20%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Alibaba Researchers Introduce Qwen-Audio Series: A Set of Large-Scale Audio-Language Models with Universal Audio Understanding Abilities

November 22, 2023
in AI & Technology
Reading Time: 4 mins read
A A
Alibaba Researchers Introduce Qwen-Audio Series: A Set of Large-Scale Audio-Language Models with Universal Audio Understanding Abilities
ShareShareShareShareShare

Researchers from Alibaba Group introduced Qwen-Audio, which addresses the challenge of limited pre-trained audio models for diverse tasks. A hierarchical tag-based multi-task framework is designed to avoid interference issues from co-training. Qwen-Audio achieves impressive performance across benchmark tasks without task-specific fine-tuning. Qwen-Audio-Chat, built upon Qwen-Audio, supports multi-turn dialogues and diverse audio-central scenarios, demonstrating its universal audio understanding abilities.

Qwen-Audio overcomes the limitations of previous audio-language models by handling diverse audio types and tasks. Unlike prior works on speech alone, Qwen-Audio incorporates human speech, natural sounds, music, and songs, allowing co-training on datasets with varying granularities. The model excels in speech perception and recognition tasks without task-specific modifications. Qwen-Audio-Chat extends these capabilities to align with human intent, supporting multilingual, multi-turn dialogues from audio and text inputs, showcasing robust and comprehensive audio understanding.

LLMs excel in general artificial intelligence but lack audio comprehension. Qwen-Audio addresses this by scaling pre-training to cover 30 tasks and diverse audio types. A multi-task framework mitigates interference, enabling knowledge sharing. Qwen-Audio performs impressively across benchmarks without task-specific fine-tuning. Qwen-Audio-Chat, an extension, supports multi-turn dialogues and diverse audio-centric scenarios, showcasing comprehensive audio interaction capabilities in LLMs.

Qwen-Audio and Qwen-Audio-Chat are models for universal audio understanding and flexible human interaction. Qwen-Audio adopts a multi-task pre-training approach, optimizing the audio encoder while freezing language model weights. In contrast, Qwen-Audio-Chat employs supervised fine-tuning, optimizing the language model while fixing audio encoder weights. The training process includes multi-task pre-training and supervised fine-tuning. Qwen-Audio-Chat enables versatile human interaction, supporting multilingual, multi-turn dialogues from audio and text inputs, showcasing its adaptability and comprehensive audio understanding.

Qwen-Audio demonstrates remarkable performance across diverse benchmark tasks, surpassing counterparts without task-specific fine-tuning. It consistently outperforms baselines by a substantial margin on jobs like AAC, SWRT ASC, SER, AQA, VSC, and MNA. The model establishes state-of-the-art results on CochlScene, ClothoAQA, and VocalSound, showcasing robust audio understanding capabilities. Qwen-Audio’s superior performance across various analyses highlights its effectiveness and competence in achieving state-of-the-art results in challenging audio tasks.

The Qwen-Audio series introduces large-scale audio-language models with universal understanding across diverse audio types and tasks. Developed through a multi-task training framework, these models facilitate knowledge sharing and overcome interference from varying textual labels in different datasets. Achieving impressive performance across benchmarks without task-specific fine-tuning, Qwen-Audio surpasses prior works. Qwen-Audio-Chat extends these capabilities, enabling multi-turn dialogues and supporting diverse audio scenarios, showcasing robust alignment with human intent and facilitating multilingual interactions.

Qwen-Audio’s future exploration includes expanding capabilities for different audio types, languages, and specific tasks. Refining the multi-task framework or exploring alternative knowledge-sharing approaches could address interference issues in co-training. Investigating task-specific fine-tuning can enhance performance. Continuous updates based on new benchmarks, datasets, and user feedback aim to improve universal audio understanding. Qwen-Audio-Chat is refined to align with human intent, support multilingual interactions, and enable dynamic multi-turn dialogues.


Check out the Paper and Github. All credit for this research goes to the researchers of this project. Also, don’t forget to join our 33k+ ML SubReddit, 41k+ Facebook Community, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.

If you like our work, you will love our newsletter..


YOU MAY ALSO LIKE

Meta Introduces Muse, a Personal AI Agent That Runs on Its Own Dedicated Secure Cloud Computer

NSA, CISA, FBI Warn China-Based AI Firms Distill US Frontier Models – Unite.AI

Sana Hassan, a consulting intern at Marktechpost and dual-degree student at IIT Madras, is passionate about applying technology and AI to address real-world challenges. With a keen interest in solving practical problems, he brings a fresh perspective to the intersection of AI and real-life solutions.


🔥 Join The AI Startup Newsletter To Learn About Latest AI Startups

Credit: Source link

ShareTweetSendSharePin

Related Posts

Meta Introduces Muse, a Personal AI Agent That Runs on Its Own Dedicated Secure Cloud Computer
AI & Technology

Meta Introduces Muse, a Personal AI Agent That Runs on Its Own Dedicated Secure Cloud Computer

September 9, 2026
NSA, CISA, FBI Warn China-Based AI Firms Distill US Frontier Models – Unite.AI
AI & Technology

NSA, CISA, FBI Warn China-Based AI Firms Distill US Frontier Models – Unite.AI

September 9, 2026
How To Change And Customize Your Apple CarPlay Display
AI & Technology

How To Change And Customize Your Apple CarPlay Display

September 8, 2026
Is There Any Benefit To Restarting Your PC Regularly?
AI & Technology

Is There Any Benefit To Restarting Your PC Regularly?

September 8, 2026
Next Post
Sam Bankman-Fried set to testify in his own fraud trial

Sam Bankman-Fried set to testify in his own fraud trial

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Stolen Claude session cookies can reach corporate Gmail through grants no IT admin can revoke

Stolen Claude session cookies can reach corporate Gmail through grants no IT admin can revoke

September 2, 2026
Anthropic’s  Billion Credit Line Sets Stage for IPO

Anthropic’s $15 Billion Credit Line Sets Stage for IPO

September 8, 2026
Night sky September 2026 — The best things to see this month – Space

Night sky September 2026 — The best things to see this month – Space

September 2, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!