• bitcoinBitcoin(BTC)$78,909.002.08%
  • ethereumEthereum(ETH)$2,531.741.06%
  • tetherTether(USDT)$1.000.01%
  • binancecoinBNB(BNB)$725.980.70%
  • rippleXRP(XRP)$1.435.79%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$103.001.98%
  • tronTRON(TRX)$0.340692-0.11%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.010.00%
  • zcashZcash(ZEC)$1,148.263.59%
  • HyperliquidHyperliquid(HYPE)$81.183.48%
  • dogecoinDogecoin(DOGE)$0.0847770.42%
  • RainRain(RAIN)$0.014371-6.41%
  • USDSUSDS(USDS)$1.000.00%
  • moneroMonero(XMR)$512.42-4.00%
  • whitebitWhiteBIT Coin(WBT)$81.601.81%
  • chainlinkChainlink(LINK)$11.601.73%
  • leo-tokenLEO Token(LEO)$9.00-0.59%
  • cardanoCardano(ADA)$0.2118191.58%
  • stellarStellar(XLM)$0.1965019.49%
  • Ethena USDeEthena USDe(USDE)$1.000.02%
  • daiDai(DAI)$1.00-0.02%
  • bitcoin-cashBitcoin Cash(BCH)$226.190.87%
  • USD1USD1(USD1)$1.000.03%
  • litecoinLitecoin(LTC)$54.14-0.90%
  • uniswapUniswap(UNI)$6.431.62%
  • CantonCanton(CC)$0.0968561.25%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.36-0.10%
  • hedera-hashgraphHedera(HBAR)$0.0779442.06%
  • nearNEAR Protocol(NEAR)$2.589.76%
  • avalanche-2Avalanche(AVAX)$7.612.48%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • shiba-inuShiba Inu(SHIB)$0.0000051.74%
  • suiSui(SUI)$0.732.02%
  • crypto-com-chainCronos(CRO)$0.0593871.78%
  • paypal-usdPayPal USD(PYUSD)$1.000.02%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,312.32-0.83%
  • BittensorBittensor(TAO)$236.30-0.06%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • MemeCoreMemeCore(M)$1.10-3.46%
  • okbOKB(OKB)$114.300.94%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.27%
  • aaveAave(AAVE)$127.530.42%
  • BitwayBitway(BTW)$0.700.79%
  • AsterAster(ASTER)$0.700.32%
  • mantleMantle(MNT)$0.571.19%
  • pax-goldPAX Gold(PAXG)$4,317.58-0.79%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0575270.92%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Tencent AI Open Sources Covo-Audio: A 7B Speech Language Model and Inference Pipeline for Real-Time Audio Conversations and Reasoning

March 26, 2026
in AI & Technology
Reading Time: 8 mins read
A A
Tencent AI Open Sources Covo-Audio: A 7B Speech Language Model and Inference Pipeline for Real-Time Audio Conversations and Reasoning
ShareShareShareShareShare

Tencent AI Lab has released Covo-Audio, a 7B-parameter end-to-end Large Audio Language Model (LALM). The model is designed to unify speech processing and language intelligence by directly processing continuous audio inputs and generating audio outputs within a single architecture.

System Architecture

The Covo-Audio framework consists of four primary components designed for seamless cross-modal interaction:

YOU MAY ALSO LIKE

Anthropic Launches Claude for Financial Advisors With Partner Connectors – Unite.AI

How To Fix Outlook’s “Your Message Can’t Be Displayed Right Now” Error

  • Audio Encoder: The model utilizes Whisper-large-v3 as its primary encoder due to its robustness against background noise and varied accents. This component operates at a frame rate of 50 Hz.
  • Audio Adapter: To bridge the encoder and the LLM, a specialized adapter employs three downsampling modules, integrating linear and convolution layers to reduce the frame rate from 50 Hz to 6.25 Hz.
  • LLM Backbone: The system is built upon Qwen2.5-7B-Base, which has been adapted to process interleaved sequences of continuous acoustic features and textual tokens.
  • Speech Tokenizer and Decoder: The tokenizer, based on WavLM-large, uses a codebook size of 16,384 to produce discrete audio tokens at 25 Hz. The decoder employs a Flow-Matching (FM) based framework and a BigVGAN vocoder to reconstruct high-fidelity 24K waveforms.
https://arxiv.org/pdf/2602.09823

Hierarchical Tri-modal Interleaving

A core contribution of this work is the Hierarchical Tri-modal Speech-Text Interleaving strategy. Unlike traditional methods that operate solely at the word or character level, this framework aligns continuous acoustic features (ac)(a_c), discrete speech tokens (ad)(a_d), and natural language text (t)(t).

The model utilizes two primary patterns:

  1. Sequential Interleaving (ac→t→ad)(a_c \rightarrow t \rightarrow a_d): Continuous features, text, and discrete tokens are arranged in a progressive chain.
  2. Parallel Integration (ac→t|ad)(a_c \rightarrow t | a_d): Continuous features are aligned with a coupled text-discrete unit.

The hierarchical aspect ensures structural coherence by using phrase-level interleaving for fine-grained alignment and sentence-level interleaving to preserve global semantic integrity in long-form utterances. The training process involved a two-stage pre-training pipeline processing a total of 2T tokens.

Intelligence-Speaker Decoupling

To mitigate the high cost of constructing large-scale dialogue data for specific speakers, the research team proposed an Intelligence Speaker Decoupling strategy. This technique separates dialogue intelligence from voice rendering, allowing for flexible voice customization using minimal text-to-speech (TTS) data.

The method reformats high-quality TTS recordings into pseudo-conversations with masked text loss. By excluding the text response portion from the loss calculation, the model preserves its reasoning abilities while inheriting the naturalness of the TTS speaker. This enables personalized interaction without the need for extensive, speaker-specific dialogue datasets.

Full-Duplex Voice Interaction

Covo-Audio evolved into Covo-Audio-Chat-FD, a variant capable of simultaneous dual-stream communication. The audio encoder is reformatted into a chunk-streaming manner, and the user and model streams are chunk-interleaved in a 1:4 ratio. Each chunk represents 0.16s of audio.

The system manages conversational states through specific architectural tokens:

  • THINK Token: Indicates a listening-only state while the model waits to respond.
  • SHIFT Token: Signifies the transition to the model’s speaking turn.
  • BREAK Token: Detects interruption signals (barge-ins), triggering the model to terminate speaking immediately and switch back to listening.

For multi-turn scenarios, the model implements a recursive context-filling strategy, where continuous audio features from user input and generated tokens from previous turns are prefixed as historical context.

Audio Reasoning and Reinforcement Learning

To enhance complex reasoning, the model incorporates Chain-of-Thought (CoT) reasoning and Group Relative Policy Optimization (GRPO). The model is optimized using a verifiable composite reward function:

$$R_{total} = R_{accuracy} + R_{format} + R_{consistency} + R_{thinking}$$

/* <![CDATA[ */
wp.i18n.setLocaleData( { 'text direction\u0004ltr': [ 'ltr' ] } );
//# sourceURL=wp-i18n-js-after
/* ]]> */

This structure allows the model to optimize for correctness (Raccuracy)(R_{accuracy}), structured output adherence (Rformat)(R_{format}), logical coherence (Rconsistency)(R_{consistency}), and reasoning depth (Rthinking)(R_{thinking}).

Evaluation and Performance

Covo-Audio (7B) shows competitive or superior results on several evaluated benchmarks, with strongest claims made for models of comparable scale and selected speech/audio tasks. On the MMAU benchmark, it achieved an average score of 75.30%, the highest among evaluated 7B-scale models. It notably excelled in music understanding with a score of 76.05%. On the MMSU benchmark, Covo-Audio achieved a leading 66.64% average accuracy.

Regarding its conversational variants, Covo-Audio-Chat demonstrated strong performance on URO-Bench, particularly in speech reasoning and spoken dialogue tasks, outperforming models like Qwen3-Omni on the Chinese track. For empathetic interaction on the VStyle benchmark, it achieved state-of-the-art results in Mandarin for anger (4.89), sadness (4.93), and anxiety (5.00).

The research team notes an ‘early-response’ issue on the GaokaoEval full-duplex setting, where unusually long silent pauses between vocal fragments can cause premature responses. This ‘early-response’ behavior correlates with the model’s pause-handling success metric and is identified as a critical direction for future optimization.

Key Takeaways

  • Unified End-to-End Architecture: Covo-Audio is a 7B-parameter model that natively processes continuous audio inputs and generates high-fidelity audio outputs within a single, unified architecture. It eliminates the need for cascaded ASR-LLM-TTS pipelines, reducing error propagation and information loss.
  • Hierarchical Tri-modal Interleaving: The model employs a specialized strategy to align continuous acoustic features, discrete speech tokens, and natural language text. By interleaving these modalities at both phrase and sentence levels, it preserves global semantic integrity while capturing fine-grained prosodic nuances.
  • Intelligence-Speaker Decoupling: Tencent research team introduces a technique to decouple dialogue intelligence from specific voice rendering. This allows for flexible voice customization using lightweight Text-to-Speech (TTS) data, significantly lowering the cost of developing personalized conversational agents.
  • Native Full-Duplex Interaction: The Covo-Audio-Chat-FD variant supports simultaneous listening and speaking. It utilizes specific architectural tokens—THINK, SHIFT, and BREAK—to manage complex real-time dynamics such as smooth turn-taking, backchanneling, and user barge-ins.
  • Superior Parameter Efficiency: Despite its compact 7B scale, Covo-Audio achieves state-of-the-art or highly competitive performance across core benchmarks, including MMAU, MMSU, and URO-Bench. It frequently matches or exceeds the performance of much larger systems, such as 32B-parameter models, in audio and speech understanding tasks.

Check out the Paper, Model on HF and Repo. Also, feel free to follow us on Twitter and don’t forget to join our 120k+ ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.

The post Tencent AI Open Sources Covo-Audio: A 7B Speech Language Model and Inference Pipeline for Real-Time Audio Conversations and Reasoning appeared first on MarkTechPost.

Credit: Source link

ShareTweetSendSharePin

Related Posts

Anthropic Launches Claude for Financial Advisors With Partner Connectors – Unite.AI
AI & Technology

Anthropic Launches Claude for Financial Advisors With Partner Connectors – Unite.AI

September 14, 2026
How To Fix Outlook’s “Your Message Can’t Be Displayed Right Now” Error
AI & Technology

How To Fix Outlook’s “Your Message Can’t Be Displayed Right Now” Error

September 14, 2026
Temporal Raises 0M Series E at .55B Valuation to Expand Operations – Unite.AI
AI & Technology

Temporal Raises $550M Series E at $12.55B Valuation to Expand Operations – Unite.AI

September 14, 2026
What Is MSI Mode On Windows PCs And Does It Speed Up Your GPU?
AI & Technology

What Is MSI Mode On Windows PCs And Does It Speed Up Your GPU?

September 14, 2026
Next Post
Donald Trump to visit Xi Jinping in May after Iran war postponement – BBC

Donald Trump to visit Xi Jinping in May after Iran war postponement - BBC

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
U.S. flag unfurled at the Pentagon to commemorate 9/11

U.S. flag unfurled at the Pentagon to commemorate 9/11

September 13, 2026
Chip Suppliers Bullish on AI Buildout

Chip Suppliers Bullish on AI Buildout

September 8, 2026
Capcom Is Reviving More Dormant Franchises After The Success Of Onimusha: Way Of The Sword

Capcom Is Reviving More Dormant Franchises After The Success Of Onimusha: Way Of The Sword

September 7, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!