• bitcoinBitcoin(BTC)$80,377.002.99%
  • ethereumEthereum(ETH)$2,522.282.97%
  • tetherTether(USDT)$1.000.01%
  • binancecoinBNB(BNB)$711.872.12%
  • rippleXRP(XRP)$1.477.01%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$107.0811.44%
  • tronTRON(TRX)$0.3375350.82%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.00-0.69%
  • HyperliquidHyperliquid(HYPE)$84.644.80%
  • dogecoinDogecoin(DOGE)$0.0889125.18%
  • zcashZcash(ZEC)$803.784.15%
  • RainRain(RAIN)$0.0174900.32%
  • USDSUSDS(USDS)$1.000.00%
  • chainlinkChainlink(LINK)$11.936.27%
  • moneroMonero(XMR)$465.327.63%
  • whitebitWhiteBIT Coin(WBT)$74.112.74%
  • leo-tokenLEO Token(LEO)$9.350.88%
  • cardanoCardano(ADA)$0.2144594.55%
  • stellarStellar(XLM)$0.1884094.93%
  • bitcoin-cashBitcoin Cash(BCH)$271.213.27%
  • CantonCanton(CC)$0.1167361.17%
  • daiDai(DAI)$1.00-0.01%
  • USD1USD1(USD1)$1.000.01%
  • Ethena USDeEthena USDe(USDE)$1.000.01%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.433.64%
  • litecoinLitecoin(LTC)$50.151.17%
  • hedera-hashgraphHedera(HBAR)$0.0791542.97%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • avalanche-2Avalanche(AVAX)$7.534.01%
  • suiSui(SUI)$0.796.95%
  • shiba-inuShiba Inu(SHIB)$0.0000054.73%
  • crypto-com-chainCronos(CRO)$0.0610465.95%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • tether-goldTether Gold(XAUT)$4,594.490.12%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • uniswapUniswap(UNI)$4.506.64%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • MemeCoreMemeCore(M)$1.12-0.61%
  • nearNEAR Protocol(NEAR)$1.924.72%
  • BittensorBittensor(TAO)$253.3811.20%
  • okbOKB(OKB)$113.882.64%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.27%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • pax-goldPAX Gold(PAXG)$4,597.120.02%
  • aaveAave(AAVE)$127.993.83%
  • Pump.funPump.fun(PUMP)$0.0048933.36%
  • AsterAster(ASTER)$0.712.74%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0593541.66%
  • OndoOndo(ONDO)$0.3784445.46%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

PolyAI Releases Dialog-RSN-1: An Audio-Native Dialog Model That Fuses Turn-Taking, Speech Recognition, Function Calling, And Response

July 31, 2026
in AI & Technology
Reading Time: 4 mins read
A A
PolyAI Releases Dialog-RSN-1: An Audio-Native Dialog Model That Fuses Turn-Taking, Speech Recognition, Function Calling, And Response
ShareShareShareShareShare

PolyAI has introduced Dialog-RSN-1, a dialog model that perceives the caller’s audio directly instead of reading a transcript. It fuses turn-taking, speech recognition, function calling and response generation into one audio-native model, and is already handling live production calls.

Key Takeaways

  • Dialog-RSN-1 is audio-aware on the input side only; TTS stays separate, so the output voice remains controllable.
  • It runs as a request-based LLM probed on demand, not an always-on stream that pins a GPU.
  • Turn-taking is the model’s first output token: EMPTY, ONGOING or COMPLETE.
  • PolyAI reports sub-300ms responses, +11% relative containment at a restaurant group, and −37% latency at an insurer.
  • English only at launch, delivered through PolyAI’s platform rather than open weights or a public API.

Is it deployable, and by whom?

Yes, but only through PolyAI: no open weights, no public API yet. Existing customers can enable it today; new customers can request early access.

YOU MAY ALSO LIKE

From In-Silico to Wet-Lab: Evaluating AI Protein Design Performance

Which Saves You The Most Money On An iPhone?

  • Company level: large, high-call-volume enterprises. PolyAI reports 100+ enterprise customers and 2,000+ live deployments at its $86M Series D in December 2025. Self-serve developers and SMBs are not the target.
  • Industries: restaurants, insurance, financial services, healthcare, hotels, retail, telecom, travel and utilities.
  • Applications: booking and reservations, billing and payments, authentication, call routing, order management and troubleshooting.

The architecture

Two architectures dominate, and each concedes something. A cascaded stack sends only the ASR’s best guess to the LLM, so tone, hesitation and recognition uncertainty are gone before the LLM sees anything. Tuning means hand-adjusting end-pointing parameters and ASR biasing that rarely generalize across use cases. Speech-to-speech models such as GPT Realtime and Gemini Live keep the audio but bake the voice into the model, limiting pronunciation control, and always-on full-duplex variants pin a GPU for the entire call.

Dialog-RSN-1 is audio-aware on input only: one model reasons over raw audio and hands generation to a separate, promptable TTS system. It is probed on demand rather than streamed: a high-recall VAD plus a few timers decide when to run it, and the first token of the reply settles whether the agent should speak. Cheap acoustic cues only choose when to ask; the model, with full context, makes the actual turn-taking call.

How it was built

PolyAI post-trained open-weight multimodal models with supervised and reinforcement finetuning on in-house data. The pipeline is broadly base-model agnostic; PolyAI evaluated Gemma, GPT-OSS, Qwen and Mistral. Targeting sub-300ms on A100 GPUs puts candidates in the 8B dense to 30B sparse range. Latency work includes prefilling the attention cache while the user speaks, an append-only prompt template to minimize cache invalidation, routing each caller to the same GPU, a finetuned speculative drafter with mean acceptance of 3.9 tokens, and auto-reasoning learned during RFT. Transcription runs last, after the response or tool call, in parallel with speech generation.

Results

PolyAI evaluated on Dialog-Eval, an internal benchmark it plans to open-source. Each example is a call truncated at one decision point, scoring a single atomic next step rather than a full rollout. PolyAI reports Dialog-RSN-1 as the highest-scoring real-time capable model, puts the cascaded Audio-score ceiling near 77, and notes GPT Realtime 2.1 scoring on par with cascades on audio-aware examples. On transcription, gpt-4o-transcribe’s WER improved from 7.8% to 6.9% once given the same context, with Dialog-RSN-1 lower still.

For this release PolyAI focused on English; Raven 3.5 remains its recommendation for non-English and rich web chat. A technical report and a Dialog-Eval paper are planned.


Check out the Technical details here. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.

Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us


Michal Sutter is a data science professional with a Master of Science in Data Science from the University of Padova. With a solid foundation in statistical analysis, machine learning, and data engineering, Michal excels at transforming complex datasets into actionable insights.

Credit: Source link

ShareTweetSendSharePin

Related Posts

From In-Silico to Wet-Lab: Evaluating AI Protein Design Performance
AI & Technology

From In-Silico to Wet-Lab: Evaluating AI Protein Design Performance

August 27, 2026
Which Saves You The Most Money On An iPhone?
AI & Technology

Which Saves You The Most Money On An iPhone?

August 27, 2026
Ransomware Operator Ran Cursor Agent Inside Ten Victim Networks – Unite.AI
AI & Technology

Ransomware Operator Ran Cursor Agent Inside Ten Victim Networks – Unite.AI

August 27, 2026
Visa ships a security AI that patches production code before any human reviews it
AI & Technology

Visa ships a security AI that patches production code before any human reviews it

August 27, 2026
Next Post
AI Bubble Burst. Is it Finally Over?

AI Bubble Burst. Is it Finally Over?

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Luigi Mangione’s state trial postponed after double jeopardy claim

Luigi Mangione’s state trial postponed after double jeopardy claim

August 21, 2026
San Diego man pleads guilty to kicking protected sea lion

San Diego man pleads guilty to kicking protected sea lion

August 22, 2026
GOP Sen. Jim Banks says GOP Rep. Max Miller should resign if ‘allegations are true’: Full interview

GOP Sen. Jim Banks says GOP Rep. Max Miller should resign if ‘allegations are true’: Full interview

August 27, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!