• bitcoinBitcoin(BTC)$83,381.00-1.61%
  • ethereumEthereum(ETH)$2,683.48-0.73%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$765.39-1.60%
  • rippleXRP(XRP)$1.52-0.50%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$119.55-2.54%
  • tronTRON(TRX)$0.3346010.19%
  • zcashZcash(ZEC)$1,591.33-3.34%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.00-5.42%
  • HyperliquidHyperliquid(HYPE)$89.72-2.72%
  • dogecoinDogecoin(DOGE)$0.094008-3.40%
  • chainlinkChainlink(LINK)$14.522.65%
  • moneroMonero(XMR)$529.99-3.44%
  • whitebitWhiteBIT Coin(WBT)$83.41-1.37%
  • USDSUSDS(USDS)$1.00-0.04%
  • cardanoCardano(ADA)$0.249801-2.00%
  • RainRain(RAIN)$0.012574-0.84%
  • leo-tokenLEO Token(LEO)$9.060.41%
  • stellarStellar(XLM)$0.2230143.63%
  • nearNEAR Protocol(NEAR)$5.201.26%
  • bitcoin-cashBitcoin Cash(BCH)$312.83-6.72%
  • uniswapUniswap(UNI)$9.01-8.08%
  • litecoinLitecoin(LTC)$70.22-1.06%
  • hedera-hashgraphHedera(HBAR)$0.11923627.14%
  • CantonCanton(CC)$0.131554-3.43%
  • Ethena USDeEthena USDe(USDE)$1.000.01%
  • avalanche-2Avalanche(AVAX)$10.50-3.31%
  • suiSui(SUI)$1.19-5.08%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.642.65%
  • daiDai(DAI)$1.000.01%
  • USD1USD1(USD1)$1.000.00%
  • BitwayBitway(BTW)$1.3315.46%
  • BittensorBittensor(TAO)$306.47-6.52%
  • quant-networkQuant(QNT)$237.5250.67%
  • crypto-com-chainCronos(CRO)$0.0685721.20%
  • shiba-inuShiba Inu(SHIB)$0.000006-3.52%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,152.45-2.97%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.02%
  • EthenaEthena(ENA)$0.268994-1.37%
  • MemeCoreMemeCore(M)$1.17-2.20%
  • OndoOndo(ONDO)$0.53-2.58%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • okbOKB(OKB)$118.05-2.87%
  • Pump.funPump.fun(PUMP)$0.00519215.08%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • aaveAave(AAVE)$149.94-2.80%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.10%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

UT Austin and ServiceNow Research Team Releases AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs

September 14, 2025
in AI & Technology
Reading Time: 11 mins read
A A
UT Austin and ServiceNow Research Team Releases AU-Harness: An Open-Source Toolkit for Holistic Evaluation of Audio LLMs
ShareShareShareShareShare

Voice AI is becoming one of the most important frontiers in multimodal AI. From intelligent assistants to interactive agents, the ability to understand and reason over audio is reshaping how machines engage with humans. Yet while models have grown rapidly in capability, the tools for evaluating them have not kept pace. Existing benchmarks remain fragmented, slow, and narrowly focused, often making it difficult to compare models or test them in realistic, multi-turn settings.

To address this gap, UT Austin and ServiceNow Research Team has released AU-Harness, a new open-source toolkit built to evaluate Large Audio Language Models (LALMs) at scale. AU-Harness is designed to be fast, standardized, and extensible, enabling researchers to test models across a wide range of tasks—from speech recognition to complex audio reasoning—within a single unified framework.

YOU MAY ALSO LIKE

A Modular, Repairable GPS Watch Is A Good First Step

Fireworks AI Releases Ember-1: A Post-Trained Kimi K3 That Uses About 40% Fewer Tokens

Why do we need a new audio evaluation framework?

Current audio benchmarks have focused on applications like speech-to-text or emotion recognition. Frameworks such as AudioBench, VoiceBench, and DynamicSUPERB-2.0 broadened coverage, but they left some really critical gaps.

Three issues stand out. First is throughput bottlenecks: many toolkits don’t take advantage of batching or parallelism, making large-scale evaluations painfully slow. Second is prompting inconsistency, which makes results across models hard to compare. Third is restricted task scope: key areas like diarization (who spoke when) and spoken reasoning (following instructions delivered in audio) are missing in many cases.

These gaps limit the progress of LALMs, especially as they evolve into multimodal agents that must handle long, context-heavy, and multi-turn interactions.

https://arxiv.org/pdf/2509.08031

How does AU-Harness improve efficiency?

The research team designed AU-Harness with focus on speed. By integrating with the vLLM inference engine, it introduces a token-based request scheduler that manages concurrent evaluations across multiple nodes. It also shards datasets so that workloads are distributed proportionally across compute resources.

This design allows near-linear scaling of evaluations and keeps hardware fully utilized. In practice, AU-Harness delivers 127% higher throughput and reduces the real-time factor (RTF) by nearly 60% compared to existing kits. For researchers, this translates into evaluations that once took days now completing in hours.

Can evaluations be customized?

Flexibility is another core feature of AU-Harness. Each model in an evaluation run can have its own hyperparameters, such as temperature or max token settings, without breaking standardization. Configurations allow for dataset filtering (e.g., by accent, audio length, or noise profile), enabling targeted diagnostics.

Perhaps most importantly, AU-Harness supports multi-turn dialogue evaluation. Earlier toolkits were limited to single-turn tasks, but modern voice agents operate in extended conversations. With AU-Harness, researchers can benchmark dialogue continuity, contextual reasoning, and adaptability across multi-step exchanges.

What tasks does AU-Harness cover?

AU-Harness dramatically expands task coverage, supporting 50+ datasets, 380+ subsets, and 21 tasks across six categories:

  • Speech Recognition: from simple ASR to long-form and code-switching speech.
  • Paralinguistics: emotion, accent, gender, and speaker recognition.
  • Audio Understanding: scene and music comprehension.
  • Spoken Language Understanding: question answering, translation, and dialogue summarization.
  • Spoken Language Reasoning: speech-to-coding, function calling, and multi-step instruction following.
  • Safety & Security: robustness evaluation and spoofing detection.

Two innovations stand out:

  • LLM-Adaptive Diarization, which evaluates diarization through prompting rather than specialized neural models.
  • Spoken Language Reasoning, which tests models’ ability to process and reason about spoken instructions, rather than just transcribe them.
https://arxiv.org/pdf/2509.08031

What do the benchmarks reveal about today’s models?

When applied to leading systems like GPT-4o, Qwen2.5-Omni, and Voxtral-Mini-3B, AU-Harness highlights both strengths and weaknesses.

Models excel at ASR and question answering, showing strong accuracy in speech recognition and spoken QA tasks. But they lag in temporal reasoning tasks, such as diarization, and in complex instruction-following, particularly when instructions are given in audio form.

A key finding is the instruction modality gap: when identical tasks are presented as spoken instructions instead of text, performance drops by as much as 9.5 points. This suggests that while models are adept at processing text-based reasoning, adapting those skills to the audio modality remains an open challenge.

https://arxiv.org/pdf/2509.08031

Summary

AU-Harness marks an important step toward standardized and scalable evaluation of audio language models. By combining efficiency, reproducibility, and broad task coverage—including diarization and spoken reasoning—it addresses the long-standing gaps in benchmarking voice-enabled AI. Its open-source release and public leaderboard invite the community to collaborate, compare, and push the boundaries of what voice-first AI systems can achieve.


Check out the Paper, Project and GitHub Page. Feel free to check out our GitHub Page for Tutorials, Codes and Notebooks. Also, feel free to follow us on Twitter and don’t forget to join our 100k+ ML SubReddit and Subscribe to our Newsletter.


Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.

Credit: Source link

ShareTweetSendSharePin

Related Posts

A Modular, Repairable GPS Watch Is A Good First Step
AI & Technology

A Modular, Repairable GPS Watch Is A Good First Step

September 28, 2026
Fireworks AI Releases Ember-1: A Post-Trained Kimi K3 That Uses About 40% Fewer Tokens
AI & Technology

Fireworks AI Releases Ember-1: A Post-Trained Kimi K3 That Uses About 40% Fewer Tokens

September 28, 2026
You Can Now Preorder The Tiny Boox Picco Ereader
AI & Technology

You Can Now Preorder The Tiny Boox Picco Ereader

September 28, 2026
20 Agentic Use Cases of TypeSafe AI’s Jev
AI & Technology

20 Agentic Use Cases of TypeSafe AI’s Jev

September 28, 2026
Next Post
No. 16 Texas A&M survives slugfest vs. No. 8 Notre Dame with game-winning drive: Takeaways – The New York Times

No. 16 Texas A&M survives slugfest vs. No. 8 Notre Dame with game-winning drive: Takeaways - The New York Times

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Social Security recipients may get bigger benefit boost in 2027

Social Security recipients may get bigger benefit boost in 2027

September 22, 2026
Ukrainian President Volodymyr Zelenskyy gifted puppy

Ukrainian President Volodymyr Zelenskyy gifted puppy

September 22, 2026
Nebraska community speaking out over use of shock gloves in school

Nebraska community speaking out over use of shock gloves in school

September 27, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!