• bitcoinBitcoin(BTC)$84,132.000.22%
  • ethereumEthereum(ETH)$2,691.240.02%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$774.20-0.18%
  • rippleXRP(XRP)$1.55-1.68%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$121.710.88%
  • tronTRON(TRX)$0.3366140.07%
  • zcashZcash(ZEC)$1,553.340.28%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.02-0.38%
  • HyperliquidHyperliquid(HYPE)$92.431.51%
  • dogecoinDogecoin(DOGE)$0.0984190.59%
  • chainlinkChainlink(LINK)$14.242.82%
  • moneroMonero(XMR)$554.93-0.11%
  • whitebitWhiteBIT Coin(WBT)$83.970.17%
  • USDSUSDS(USDS)$1.000.01%
  • cardanoCardano(ADA)$0.2590511.62%
  • RainRain(RAIN)$0.0127917.95%
  • leo-tokenLEO Token(LEO)$8.981.76%
  • stellarStellar(XLM)$0.2197480.65%
  • bitcoin-cashBitcoin Cash(BCH)$335.800.40%
  • nearNEAR Protocol(NEAR)$4.85-6.31%
  • uniswapUniswap(UNI)$9.681.18%
  • litecoinLitecoin(LTC)$72.994.61%
  • CantonCanton(CC)$0.13922912.26%
  • Ethena USDeEthena USDe(USDE)$1.00-0.02%
  • avalanche-2Avalanche(AVAX)$10.975.38%
  • suiSui(SUI)$1.186.31%
  • daiDai(DAI)$1.000.00%
  • USD1USD1(USD1)$1.000.00%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.484.41%
  • hedera-hashgraphHedera(HBAR)$0.0950061.40%
  • BittensorBittensor(TAO)$335.3910.21%
  • shiba-inuShiba Inu(SHIB)$0.0000062.98%
  • crypto-com-chainCronos(CRO)$0.0660170.07%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • BitwayBitway(BTW)$1.06-14.70%
  • MemeCoreMemeCore(M)$1.223.89%
  • EthenaEthena(ENA)$0.2746917.15%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • tether-goldTether Gold(XAUT)$4,279.61-0.24%
  • OndoOndo(ONDO)$0.552.36%
  • okbOKB(OKB)$122.001.54%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • aaveAave(AAVE)$154.563.39%
  • mantleMantle(MNT)$0.705.91%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.15-0.11%
  • polkadotPolkadot(DOT)$1.288.20%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Noise-Augmented CAM (Continuous Autoregressive Models): Advancing Real-Time Audio Generation

December 9, 2024
in AI & Technology
Reading Time: 5 mins read
A A
Noise-Augmented CAM (Continuous Autoregressive Models): Advancing Real-Time Audio Generation
ShareShareShareShareShare

Autoregressive models are used to generate sequences of discrete tokens. The next token is conditioned by the preceding tokens in a given sequence in the approach. Recent research showed that generating sequences of continuous embeddings autoregressively is also feasible. However, such Continuous Autoregressive Models (CAMs) generate these embeddings similarly sequentially, but they face challenges such as a decline in generation quality over extended sequences. This decline occurs because of error accumulation during the inference process, where small prediction errors compound as the sequence length increases, resulting in degraded output.

Traditional models for autoregressive image and audio generation relied on discretizing data into tokens using VQ-VAEs to enable models to work within a discrete probability space. Such an approach introduces significant drawbacks, including additional losses when training VAEs and added complexity. Although continuous embeddings are more efficient, they tend to accumulate errors during inference, causing distribution shifts and lowering the generated output’s quality. Recent attempts to bypass quantization by training on continuous embeddings have failed to produce convincing results due to cumbersome non-sequential masking and fine-tuning techniques impair efficiency and restrict further usage within the research community.

YOU MAY ALSO LIKE

This App Lets You Use An Apple Watch With An Android Phone

These Xbox Players Got GTA 6 For Free The Hard Way

To solve this, a group of researchers from Queen Mary University and Sony Computer Science Laboratories conducted detailed research and proposed a method to counteract error accumulation and train purely autoregressive models on ordered sequences of continuous embeddings without adding complexity. To overcome the drawbacks of standard AMs, CAM introduced a noise augmentation strategy during training to simulate the errors that occur during inference. This method combined the strengths of Rectified Flow (RF) and AMs for continuous embeddings.

The main concept behind the CAM proposed was injecting noise in the sequence during training to simulate error-prone inference conditions. It then applied iterative reverse diffusion to generate sequences autoregressively, progressively improving predictions while correcting mistakes. CAM was pre-trained to be robust for error accumulation during the generation of longer sequences through training with noisy sequences. This process improved the general quality of the generated sequences, especially for tasks such as music generation, for which the quality of each predicted element proved crucial to the overall output.

The method was tested on a music dataset and compared with the experiment’s autoregressive and non-autoregressive baselines. The researchers used a dataset of about 20,000 single-instrument recordings with 48 kHz stereo audio for training and evaluation. They processed the data with Music2Latent to create continuous latent embeddings with a 12 Hz sampling rate. Based on a transformer with 16 layers and 150 million parameters, CAM was trained using AdamW for 400k iterations. CAM performed better than the other models, with FAD of 0.405 and FADacc of 0.394, compared to baselines like GIVT or MAR. CAM provided better quality basics for reconstructing the sound spectrum and avoiding the error buildup in long sequences; the noise augmentation approach also helped to enhance the GIVT scores.

In summary, the proposed method trains purely autoregressive models on continuous embeddings that directly address the error accumulation problem. A noise injection technique calibrated carefully at inference time further reduces error accumulation. This method opens the path for real-time and interactive audio applications that benefit from the efficiency and sequential nature of autoregressive models and can be used as a baseline for further research in the domain!


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter and join our Telegram Channel and LinkedIn Group. If you like our work, you will love our newsletter.. Don’t Forget to join our 60k+ ML SubReddit.

🚨 [Must Attend Webinar]: ‘Transform proofs-of-concept into production-ready AI applications and agents’ (Promoted)


Divyesh is a consulting intern at Marktechpost. He is pursuing a BTech in Agricultural and Food Engineering from the Indian Institute of Technology, Kharagpur. He is a Data Science and Machine learning enthusiast who wants to integrate these leading technologies into the agricultural domain and solve challenges.

🚨🚨FREE AI WEBINAR: ‘Fast-Track Your LLM Apps with deepset & Haystack'(Promoted)


Credit: Source link

ShareTweetSendSharePin

Related Posts

This App Lets You Use An Apple Watch With An Android Phone
AI & Technology

This App Lets You Use An Apple Watch With An Android Phone

September 26, 2026
These Xbox Players Got GTA 6 For Free The Hard Way
AI & Technology

These Xbox Players Got GTA 6 For Free The Hard Way

September 26, 2026
Exa Launches Agent Ultra: A Subagent Swarm Deep Research API Built for Exhaustive List Building
AI & Technology

Exa Launches Agent Ultra: A Subagent Swarm Deep Research API Built for Exhaustive List Building

September 26, 2026
End-to-End Multimodal Data Augmentation and Adversarial Robustness Benchmark with AugLy for Images, Text, Audio, and PyTorch
AI & Technology

End-to-End Multimodal Data Augmentation and Adversarial Robustness Benchmark with AugLy for Images, Text, Audio, and PyTorch

September 26, 2026
Next Post
TOP 5 STOCKS TO WATCH AS BITCOIN HITS 0,000

TOP 5 STOCKS TO WATCH AS BITCOIN HITS $100,000

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Hideo Kojima explains the surprise split with Sony – The Washington Post

Hideo Kojima explains the surprise split with Sony – The Washington Post

September 20, 2026
Sen. Ted Cruz comments on possible 2028 presidential run

Sen. Ted Cruz comments on possible 2028 presidential run

September 21, 2026
How To Hide Or Replace The Audio Button In iMessages

How To Hide Or Replace The Audio Button In iMessages

September 22, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!