• bitcoinBitcoin(BTC)$62,905.00-3.00%
  • ethereumEthereum(ETH)$1,860.88-3.30%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$586.95-1.30%
  • usd-coinUSDC(USDC)$1.000.00%
  • rippleXRP(XRP)$1.06-2.70%
  • solanaSolana(SOL)$72.96-2.50%
  • tronTRON(TRX)$0.325904-0.90%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.032.30%
  • whitebitWhiteBIT Coin(WBT)$54.92-3.00%
  • HyperliquidHyperliquid(HYPE)$52.71-4.90%
  • dogecoinDogecoin(DOGE)$0.069603-2.00%
  • USDSUSDS(USDS)$1.000.00%
  • leo-tokenLEO Token(LEO)$9.75-0.20%
  • RainRain(RAIN)$0.012745-4.40%
  • zcashZcash(ZEC)$456.09-3.90%
  • moneroMonero(XMR)$355.12-2.10%
  • cardanoCardano(ADA)$0.169048-2.50%
  • chainlinkChainlink(LINK)$8.14-4.20%
  • stellarStellar(XLM)$0.171975-0.90%
  • CantonCanton(CC)$0.117540-2.80%
  • daiDai(DAI)$1.000.00%
  • bitcoin-cashBitcoin Cash(BCH)$208.67-4.70%
  • USD1USD1(USD1)$1.000.00%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.39-2.40%
  • litecoinLitecoin(LTC)$44.60-1.90%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • Circle USYCCircle USYC(USYC)$1.130.10%
  • hedera-hashgraphHedera(HBAR)$0.068176-0.70%
  • shiba-inuShiba Inu(SHIB)$0.0000051.90%
  • avalanche-2Avalanche(AVAX)$6.40-1.30%
  • suiSui(SUI)$0.68-3.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • uniswapUniswap(UNI)$4.30-2.80%
  • crypto-com-chainCronos(CRO)$0.054216-1.40%
  • tether-goldTether Gold(XAUT)$4,035.50-1.70%
  • nearNEAR Protocol(NEAR)$1.68-1.30%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.10%
  • OndoOndo(ONDO)$0.392962-7.00%
  • BittensorBittensor(TAO)$192.85-0.60%
  • okbOKB(OKB)$86.180.40%
  • pax-goldPAX Gold(PAXG)$4,040.45-1.60%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.054894-0.40%
  • AsterAster(ASTER)$0.60-1.40%
  • HTX DAOHTX DAO(HTX)$0.000002-1.30%
  • usddUSDD(USDD)$1.000.00%
  • MemeCoreMemeCore(M)$1.1516.50%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

PolyAI Releases Dialog-RSN-1: An Audio-Native Dialog Model That Fuses Turn-Taking, Speech Recognition, Function Calling, And Response

July 31, 2026
in AI & Technology
Reading Time: 4 mins read
A A
PolyAI Releases Dialog-RSN-1: An Audio-Native Dialog Model That Fuses Turn-Taking, Speech Recognition, Function Calling, And Response
ShareShareShareShareShare

PolyAI has introduced Dialog-RSN-1, a dialog model that perceives the caller’s audio directly instead of reading a transcript. It fuses turn-taking, speech recognition, function calling and response generation into one audio-native model, and is already handling live production calls.

Key Takeaways

  • Dialog-RSN-1 is audio-aware on the input side only; TTS stays separate, so the output voice remains controllable.
  • It runs as a request-based LLM probed on demand, not an always-on stream that pins a GPU.
  • Turn-taking is the model’s first output token: EMPTY, ONGOING or COMPLETE.
  • PolyAI reports sub-300ms responses, +11% relative containment at a restaurant group, and −37% latency at an insurer.
  • English only at launch, delivered through PolyAI’s platform rather than open weights or a public API.

Is it deployable, and by whom?

Yes, but only through PolyAI: no open weights, no public API yet. Existing customers can enable it today; new customers can request early access.

YOU MAY ALSO LIKE

Claude Turned a Cyber Benchmark Into Three Real Intrusions – Unite.AI

Australia’s Social Media Ban For Under-16s Has Had Limited Impact So Far

  • Company level: large, high-call-volume enterprises. PolyAI reports 100+ enterprise customers and 2,000+ live deployments at its $86M Series D in December 2025. Self-serve developers and SMBs are not the target.
  • Industries: restaurants, insurance, financial services, healthcare, hotels, retail, telecom, travel and utilities.
  • Applications: booking and reservations, billing and payments, authentication, call routing, order management and troubleshooting.

The architecture

Two architectures dominate, and each concedes something. A cascaded stack sends only the ASR’s best guess to the LLM, so tone, hesitation and recognition uncertainty are gone before the LLM sees anything. Tuning means hand-adjusting end-pointing parameters and ASR biasing that rarely generalize across use cases. Speech-to-speech models such as GPT Realtime and Gemini Live keep the audio but bake the voice into the model, limiting pronunciation control, and always-on full-duplex variants pin a GPU for the entire call.

Dialog-RSN-1 is audio-aware on input only: one model reasons over raw audio and hands generation to a separate, promptable TTS system. It is probed on demand rather than streamed: a high-recall VAD plus a few timers decide when to run it, and the first token of the reply settles whether the agent should speak. Cheap acoustic cues only choose when to ask; the model, with full context, makes the actual turn-taking call.

How it was built

PolyAI post-trained open-weight multimodal models with supervised and reinforcement finetuning on in-house data. The pipeline is broadly base-model agnostic; PolyAI evaluated Gemma, GPT-OSS, Qwen and Mistral. Targeting sub-300ms on A100 GPUs puts candidates in the 8B dense to 30B sparse range. Latency work includes prefilling the attention cache while the user speaks, an append-only prompt template to minimize cache invalidation, routing each caller to the same GPU, a finetuned speculative drafter with mean acceptance of 3.9 tokens, and auto-reasoning learned during RFT. Transcription runs last, after the response or tool call, in parallel with speech generation.

Results

PolyAI evaluated on Dialog-Eval, an internal benchmark it plans to open-source. Each example is a call truncated at one decision point, scoring a single atomic next step rather than a full rollout. PolyAI reports Dialog-RSN-1 as the highest-scoring real-time capable model, puts the cascaded Audio-score ceiling near 77, and notes GPT Realtime 2.1 scoring on par with cascades on audio-aware examples. On transcription, gpt-4o-transcribe’s WER improved from 7.8% to 6.9% once given the same context, with Dialog-RSN-1 lower still.

For this release PolyAI focused on English; Raven 3.5 remains its recommendation for non-English and rich web chat. A technical report and a Dialog-Eval paper are planned.


Check out the Technical details here. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.

Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us


Michal Sutter is a data science professional with a Master of Science in Data Science from the University of Padova. With a solid foundation in statistical analysis, machine learning, and data engineering, Michal excels at transforming complex datasets into actionable insights.

Credit: Source link

ShareTweetSendSharePin

Related Posts

Claude Turned a Cyber Benchmark Into Three Real Intrusions – Unite.AI
AI & Technology

Claude Turned a Cyber Benchmark Into Three Real Intrusions – Unite.AI

July 31, 2026
Australia’s Social Media Ban For Under-16s Has Had Limited Impact So Far
AI & Technology

Australia’s Social Media Ban For Under-16s Has Had Limited Impact So Far

July 31, 2026
Why Your Biomedical RAG Is Hiding Contradictions From You – Unite.AI
AI & Technology

Why Your Biomedical RAG Is Hiding Contradictions From You – Unite.AI

July 31, 2026
JetBrains Open-Sources KotlinLLM: Smart Macros That Generate Kotlin Source Code at Runtime and Hot-Reload It Through JDI
AI & Technology

JetBrains Open-Sources KotlinLLM: Smart Macros That Generate Kotlin Source Code at Runtime and Hot-Reload It Through JDI

July 31, 2026
Next Post
AI Bubble Burst. Is it Finally Over?

AI Bubble Burst. Is it Finally Over?

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Most Traders Make This Mistake Every Single Day

Most Traders Make This Mistake Every Single Day

July 26, 2026
Katie Couric reveals recent diagnosis of transient global amnesia

Katie Couric reveals recent diagnosis of transient global amnesia

July 31, 2026
Folarin Balogun talks World Cup loss, red card and the future of soccer in USA

Folarin Balogun talks World Cup loss, red card and the future of soccer in USA

July 26, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!