• bitcoinBitcoin(BTC)$76,614.001.16%
  • ethereumEthereum(ETH)$2,455.672.28%
  • tetherTether(USDT)$1.00-0.02%
  • binancecoinBNB(BNB)$727.462.25%
  • rippleXRP(XRP)$1.302.03%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$100.473.33%
  • tronTRON(TRX)$0.3347810.05%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.032.65%
  • zcashZcash(ZEC)$1,357.9511.40%
  • HyperliquidHyperliquid(HYPE)$80.493.03%
  • dogecoinDogecoin(DOGE)$0.0811342.57%
  • USDSUSDS(USDS)$1.000.04%
  • moneroMonero(XMR)$502.430.48%
  • whitebitWhiteBIT Coin(WBT)$78.981.48%
  • RainRain(RAIN)$0.013085-2.15%
  • chainlinkChainlink(LINK)$11.265.05%
  • leo-tokenLEO Token(LEO)$8.920.44%
  • cardanoCardano(ADA)$0.2007254.35%
  • stellarStellar(XLM)$0.1847555.63%
  • Ethena USDeEthena USDe(USDE)$1.000.04%
  • daiDai(DAI)$1.00-0.03%
  • bitcoin-cashBitcoin Cash(BCH)$225.203.49%
  • uniswapUniswap(UNI)$7.0012.17%
  • USD1USD1(USD1)$1.00-0.02%
  • litecoinLitecoin(LTC)$52.814.44%
  • CantonCanton(CC)$0.10067810.28%
  • nearNEAR Protocol(NEAR)$2.8716.41%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.342.10%
  • avalanche-2Avalanche(AVAX)$7.594.51%
  • hedera-hashgraphHedera(HBAR)$0.0747541.26%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • shiba-inuShiba Inu(SHIB)$0.0000055.10%
  • suiSui(SUI)$0.735.58%
  • crypto-com-chainCronos(CRO)$0.0579914.91%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • tether-goldTether Gold(XAUT)$4,360.510.27%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • BittensorBittensor(TAO)$226.954.55%
  • MemeCoreMemeCore(M)$1.132.72%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • okbOKB(OKB)$112.092.11%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.22%
  • AsterAster(ASTER)$0.749.05%
  • aaveAave(AAVE)$124.195.73%
  • pax-goldPAX Gold(PAXG)$4,362.350.20%
  • mantleMantle(MNT)$0.573.57%
  • Pump.funPump.fun(PUMP)$0.00396710.31%
  • BitwayBitway(BTW)$0.68-10.66%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

TRAMBA: A Novel Hybrid Transformer and Mamba-based Architecture for Speech Super Resolution and Enhancement for Mobile and Wearable Platforms

May 9, 2024
in AI & Technology
Reading Time: 5 mins read
A A
TRAMBA: A Novel Hybrid Transformer and Mamba-based Architecture for Speech Super Resolution and Enhancement for Mobile and Wearable Platforms
ShareShareShareShareShare

Wearables have transformed human-technology interaction, facilitating continuous health monitoring. The wearables market is projected to surge from 70 billion USD in 2023 to 230 billion USD by 2032, with head-worn devices, including earphones and glasses, experiencing rapid growth (71 billion USD in 2023 to 172 billion USD by 2030). This growth is propelled by the rising significance of wearables, augmented reality (AR), and virtual reality (VR). Head-worn wearables uniquely capture speech signals, traditionally collected by over-the-air (OTA) microphones near or on the head, converting air pressure fluctuations into electrical signals for various applications. However, OTA microphones, typically located near the mouth, easily capture background noise, potentially compromising speech quality, particularly in noisy environments.

Various studies have tackled the challenge of separating speech from background noise through denoising, sound source separation, and speech enhancement techniques. However, this approach is hindered by the model’s inability to anticipate the diverse types of background noises and the prevalence of noisy environments, such as bustling cafeterias or construction sites. Unlike OTA microphones, bone conduction microphones (BCM) placed directly on the head are resilient to ambient noise, detecting vibrations from the skin and skull during speech. Although BCMs offer noise robustness, vibration-based methods suffer from frequency attenuation, affecting speech intelligibility. Some research endeavors explore vibration and bone-conduction super-resolution methods to reconstruct higher frequencies for improved speech quality, yet practical implementation for real-time wearable systems faces challenges. These include the heavy processing demands of state-of-the-art speech super-resolution models like generative adversarial networks (GANs), which require substantial memory and computational resources, resulting in performance gaps compared to smaller footprint methods. Optimization considerations, such as sampling rate and deployment strategies, remain crucial for enhancing real-time system efficiency.

Researchers from Northwestern University and Columbia University introduced TRAMBA, a hybrid transformer, and Mamba architecture for enhancing acoustic and bone conduction speech in mobile and wearable platforms. Previously, adopting bone conduction speech enhancement in such platforms faced challenges due to labor-intensive data collection and performance gaps between models. TRAMBA addresses this by pre-training with widely available audio speech datasets and fine-tuning with a small amount of bone conduction data. It achieves reconstructing intelligible speech using a single wearable accelerometer, demonstrating generalizability across multiple acoustic modalities. Integrated into wearable and mobile platforms, TRAMBA enables real-time speech super-resolution and significant power consumption reduction. This is also the first study to sense intelligible speech using only a single head-worn accelerometer.

At a macro level, TRAMBA architecture integrates a modified U-Net structure with self-attention in the downsampling and upsampling layers, along with Mamba in the narrow bottleneck layer. TRAMBA operates on 512ms windows of single-channel audio and preprocesses acceleration data from an accelerometer. Each downsampling block consists of a 1D convolutional layer with LeakyReLU activations, followed by a robust conditioning layer called Scale-only Attention-based Feature-wise Linear Modulation (SAFiLM). SAFiLM utilizes a multi-head attention mechanism to learn scaling factors for enhancing feature representations. The bottleneck layer employs Mamba, known for its efficient memory usage and attention mechanisms akin to transformers. However, due to gradient vanishing issues, transformers are retained only in the downsampling and upsampling blocks. Residual connections are employed to facilitate gradient flow and optimize deeper networks, enhancing training efficiency.

TRAMBA exhibits superior performance across various metrics and sampling rates compared to other models, including U-Net architectures. Although the Aero GAN method slightly outperforms TRAMBA in the LSD metric, TRAMBA excels in perceptual and noise metrics such as SNR, PESQ, and STOI. This highlights the effectiveness of integrating transformers and Mamba in enhancing local speech formants compared to traditional architectures. Also, transformer and Mamba-based models demonstrate superior performance over state-of-the-art GANs with significantly reduced memory and inference time requirements. Notably, TRAMBA’s efficient processing allows for real-time operation, unlike Aero GAN, which exceeds the window size, making it impractical for real-time applications. Comparisons with the top-performing U-Net architecture (TUNet) are also made.

In conclusion, this study presents TRAMBA, a hybrid architecture combining transformer and Mamba elements for speech super-resolution and enhancement on mobile and wearable platforms. It surpasses existing methods across various acoustic modalities while maintaining a compact memory footprint of only 19.7 MBs, contrasting with GANs requiring at least hundreds of MBs. Integrated into real mobile and head-worn wearable systems, TRAMBA exhibits superior speech quality in noisy environments compared to traditional denoising approaches. Also, it extends battery life by up to 160% by reducing the resolution of audio that needs to be sampled and transmitted. TRAMBA represents a crucial advancement for integrating speech enhancement into practical mobile and wearable platforms.


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. Join our Telegram Channel, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 41k+ ML SubReddit


YOU MAY ALSO LIKE

NVIDIA And Google’s New Coalition Wants To Speed Up AI Data Center Power Grid Connections

Z.ai Details GLM-5.3-Flash Inference Build on 100,000 Chinese Chips – Unite.AI

Asjad is an intern consultant at Marktechpost. He is persuing B.Tech in mechanical engineering at the Indian Institute of Technology, Kharagpur. Asjad is a Machine learning and deep learning enthusiast who is always researching the applications of machine learning in healthcare.


✅ [FREE AI WEBINAR Alert] Live RAG Comparison Test: Pinecone vs Mongo vs Postgres vs SingleStore: May 9, 2024 10:00am – 11:00am PDT


Credit: Source link

ShareTweetSendSharePin

Related Posts

NVIDIA And Google’s New Coalition Wants To Speed Up AI Data Center Power Grid Connections
AI & Technology

NVIDIA And Google’s New Coalition Wants To Speed Up AI Data Center Power Grid Connections

September 17, 2026
Z.ai Details GLM-5.3-Flash Inference Build on 100,000 Chinese Chips – Unite.AI
AI & Technology

Z.ai Details GLM-5.3-Flash Inference Build on 100,000 Chinese Chips – Unite.AI

September 17, 2026
OpenAI Releases a Model Misalignment Disclosure Framework With 3 Review Tracks and 6 Incident Reports From RL Training
AI & Technology

OpenAI Releases a Model Misalignment Disclosure Framework With 3 Review Tracks and 6 Incident Reports From RL Training

September 17, 2026
Google Research Introduces Retrieve-for-Train (R4T): An RL-Compiled Diffusion Retriever for 12× to 20× Faster Query Fan-Out
AI & Technology

Google Research Introduces Retrieve-for-Train (R4T): An RL-Compiled Diffusion Retriever for 12× to 20× Faster Query Fan-Out

September 17, 2026
Next Post
Florida man wrongly convicted of murder awarded  million

Florida man wrongly convicted of murder awarded $14 million

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Supreme Court rejects use of GOP Missouri congressional map

Supreme Court rejects use of GOP Missouri congressional map

September 14, 2026
Q3 Estimated Tax Is Due September 15

Q3 Estimated Tax Is Due September 15

September 14, 2026
Reddington: Lindsay Clancy is ‘not good’ following mistrial

Reddington: Lindsay Clancy is ‘not good’ following mistrial

September 17, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!