• bitcoinBitcoin(BTC)$83,403.00-1.17%
  • ethereumEthereum(ETH)$2,652.76-1.67%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$771.38-0.05%
  • rippleXRP(XRP)$1.50-1.28%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$120.10-0.50%
  • tronTRON(TRX)$0.3336750.26%
  • zcashZcash(ZEC)$1,566.68-4.71%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.06-0.38%
  • HyperliquidHyperliquid(HYPE)$90.09-3.37%
  • dogecoinDogecoin(DOGE)$0.094811-1.43%
  • chainlinkChainlink(LINK)$14.01-0.69%
  • moneroMonero(XMR)$543.12-3.04%
  • whitebitWhiteBIT Coin(WBT)$83.20-1.07%
  • USDSUSDS(USDS)$1.00-0.01%
  • cardanoCardano(ADA)$0.2525190.23%
  • RainRain(RAIN)$0.012540-1.81%
  • leo-tokenLEO Token(LEO)$9.040.07%
  • stellarStellar(XLM)$0.213579-0.64%
  • nearNEAR Protocol(NEAR)$5.252.74%
  • bitcoin-cashBitcoin Cash(BCH)$314.25-5.56%
  • uniswapUniswap(UNI)$9.39-3.63%
  • CantonCanton(CC)$0.1425834.28%
  • litecoinLitecoin(LTC)$70.89-0.93%
  • suiSui(SUI)$1.256.13%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • avalanche-2Avalanche(AVAX)$10.790.34%
  • daiDai(DAI)$1.00-0.01%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.623.80%
  • USD1USD1(USD1)$1.000.00%
  • hedera-hashgraphHedera(HBAR)$0.0951501.90%
  • quant-networkQuant(QNT)$262.5051.86%
  • BittensorBittensor(TAO)$309.84-3.42%
  • BitwayBitway(BTW)$1.2621.93%
  • shiba-inuShiba Inu(SHIB)$0.000006-1.66%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.064722-2.24%
  • tether-goldTether Gold(XAUT)$4,201.62-1.83%
  • EthenaEthena(ENA)$0.2740832.67%
  • OndoOndo(ONDO)$0.575.95%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • MemeCoreMemeCore(M)$1.19-3.48%
  • okbOKB(OKB)$118.68-1.82%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • Pump.funPump.fun(PUMP)$0.00510416.18%
  • aaveAave(AAVE)$151.17-2.61%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.10%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

This AI Paper Proposes a Novel Dual-Branch Encoder-Decoder Architecture for Unsupervised Speech Enhancement (SE)

October 5, 2025
in AI & Technology
Reading Time: 5 mins read
A A
This AI Paper Proposes a Novel Dual-Branch Encoder-Decoder Architecture for Unsupervised Speech Enhancement (SE)
ShareShareShareShareShare

Can a speech enhancer trained only on real noisy recordings cleanly separate speech and noise—without ever seeing paired data? A team of researchers from Brno University of Technology and Johns Hopkins University proposes Unsupervised Speech Enhancement using Data-defined Priors (USE-DDP), a dual-stream encoder–decoder that separates any noisy input into two waveforms—estimated clean speech and residual noise—and learns both solely from unpaired datasets (clean-speech corpus and optional noise corpus). Training enforces that the sum of the two outputs reconstructs the input waveform, avoiding degenerate solutions and aligning the design with neural audio codec objectives.

https://arxiv.org/pdf/2509.22942

Why this is important?

Most learning-based speech enhancement pipelines depend on paired clean–noisy recordings, which are expensive or impossible to collect at scale in real-world conditions. Unsupervised routes like MetricGAN-U remove the need for clean data but couple model performance to external, non-intrusive metrics used during training. USE-DDP keeps the training data-only, imposing priors with discriminators over independent clean-speech and noise datasets and using reconstruction consistency to tie estimates back to the observed mixture.

YOU MAY ALSO LIKE

20 Agentic Use Cases of TypeSafe AI’s Jev

Which Is Better To Use?

How it works?

  • Generator: A codec-style encoder compresses the input audio into a latent sequence; this is split into two parallel transformer branches (RoFormer) that target clean speech and noise respectively, decoded by a shared decoder back to waveforms. The input is reconstructed as the least-squares combination of the two outputs (scalars α, β compensate for amplitude errors). Reconstruction uses multi-scale mel/STFT and SI-SDR losses, as in neural audio codecs.
  • Priors via adversaries: Three discriminator ensembles—clean, noise, and noisy—impose distributional constraints: the clean branch must resemble the clean-speech corpus; the noise branch must resemble a noise corpus; the reconstructed mixture must sound natural. LS-GAN and feature-matching losses are used.
  • Initialization: Initializing encoder/decoder from a pretrained Descript Audio Codec improves convergence and final quality vs. training from scratch.

How it compares?

On the standard VCTK+DEMAND simulated setup, USE-DDP reports parity with the strongest unsupervised baselines (e.g., unSE/unSE+ based on optimal transport) and competitive DNSMOS vs. MetricGAN-U (which directly optimizes DNSMOS). Example numbers from the paper’s Table 1 (input vs. systems): DNSMOS improves from 2.54 (noisy) to ~3.03 (USE-DDP), PESQ from 1.97 to ~2.47; CBAK trails some baselines due to more aggressive noise attenuation in non-speech segments—consistent with the explicit noise prior.

https://arxiv.org/pdf/2509.22942

Data choice is not a detail—it’s the result

A central finding: which clean-speech corpus defines the prior can swing outcomes and even create over-optimistic results on simulated tests.

  • In-domain prior (VCTK clean) on VCTK+DEMAND → best scores (DNSMOS ≈3.03), but this configuration unrealistically “peeks” at the target distribution used to synthesize the mixtures.
  • Out-of-domain prior → notably lower metrics (e.g., PESQ ~2.04), reflecting distribution mismatch and some noise leakage into the clean branch.
  • Real-world CHiME-3: using a “close-talk” channel as in-domain clean prior actually hurts—because the “clean” reference itself contains environment bleed; an out-of-domain truly clean corpus yields higher DNSMOS/UTMOS on both dev and test, albeit with some intelligibility trade-off under stronger suppression.

This clarifies discrepancies across prior unsupervised results and argues for careful, transparent prior selection when claiming SOTA on simulated benchmarks.

Our Comments

The proposed dual-branch encoder-decoder architecture treats enhancement as explicit two-source estimation with data-defined priors, not metric-chasing. The reconstruction constraint (clean + noise = input) plus adversarial priors over independent clean/noise corpora gives a clear inductive bias, and initializing from a neural audio codec is a pragmatic way to stabilize training. The results look competitive with unsupervised baselines while avoiding DNSMOS-guided objectives; the caveat is that “clean prior” choice materially affects reported gains, so claims should specify corpus selection.


Check out the PAPER. Feel free to check out our GitHub Page for Tutorials, Codes and Notebooks. Also, feel free to follow us on Twitter and don’t forget to join our 100k+ ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.

The post This AI Paper Proposes a Novel Dual-Branch Encoder-Decoder Architecture for Unsupervised Speech Enhancement (SE) appeared first on MarkTechPost.

Credit: Source link

ShareTweetSendSharePin

Related Posts

20 Agentic Use Cases of TypeSafe AI’s Jev
AI & Technology

20 Agentic Use Cases of TypeSafe AI’s Jev

September 28, 2026
Which Is Better To Use?
AI & Technology

Which Is Better To Use?

September 28, 2026
Are 3D Printers Worth Buying In 2026?
AI & Technology

Are 3D Printers Worth Buying In 2026?

September 28, 2026
Bill Gates Says It’s ‘Completely Irresponsible’ For AI To Not Have Safeguards
AI & Technology

Bill Gates Says It’s ‘Completely Irresponsible’ For AI To Not Have Safeguards

September 27, 2026
Next Post
Former FBI Director James Comey charged with making a false statement and obstruction

Former FBI Director James Comey charged with making a false statement and obstruction

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Start your engines: NBC News takes you inside IndyCar’s D.C. Grand Prix

Start your engines: NBC News takes you inside IndyCar’s D.C. Grand Prix

September 26, 2026
Panda Express CEOs share relatable morning routines

Panda Express CEOs share relatable morning routines

September 23, 2026
Judge in Lindsay Clancy case denies defense’s motion for mistrial after witness brings up religion

Judge in Lindsay Clancy case denies defense’s motion for mistrial after witness brings up religion

September 24, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!