• bitcoinBitcoin(BTC)$80,492.00-0.71%
  • ethereumEthereum(ETH)$2,579.24-1.76%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$749.25-1.63%
  • rippleXRP(XRP)$1.38-2.70%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$108.98-3.18%
  • tronTRON(TRX)$0.3399500.46%
  • zcashZcash(ZEC)$1,455.50-4.87%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.02-1.50%
  • HyperliquidHyperliquid(HYPE)$90.98-1.85%
  • dogecoinDogecoin(DOGE)$0.085452-2.53%
  • moneroMonero(XMR)$524.59-7.45%
  • whitebitWhiteBIT Coin(WBT)$81.87-1.62%
  • USDSUSDS(USDS)$1.00-0.01%
  • RainRain(RAIN)$0.0134990.75%
  • chainlinkChainlink(LINK)$12.03-2.97%
  • cardanoCardano(ADA)$0.221059-2.37%
  • leo-tokenLEO Token(LEO)$8.89-0.05%
  • stellarStellar(XLM)$0.190668-1.78%
  • uniswapUniswap(UNI)$8.77-2.19%
  • bitcoin-cashBitcoin Cash(BCH)$247.19-0.21%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • daiDai(DAI)$1.00-0.02%
  • nearNEAR Protocol(NEAR)$3.48-6.47%
  • litecoinLitecoin(LTC)$56.96-2.45%
  • USD1USD1(USD1)$1.00-0.01%
  • avalanche-2Avalanche(AVAX)$9.6113.53%
  • CantonCanton(CC)$0.105380-5.30%
  • MemeCoreMemeCore(M)$1.7132.11%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.370.16%
  • hedera-hashgraphHedera(HBAR)$0.0806061.67%
  • suiSui(SUI)$0.82-0.53%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • shiba-inuShiba Inu(SHIB)$0.000005-0.78%
  • crypto-com-chainCronos(CRO)$0.058871-0.23%
  • BittensorBittensor(TAO)$252.30-1.12%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • tether-goldTether Gold(XAUT)$4,368.43-0.14%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • okbOKB(OKB)$115.88-0.72%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.15-0.01%
  • aaveAave(AAVE)$137.54-3.79%
  • AsterAster(ASTER)$0.74-4.55%
  • EthenaEthena(ENA)$0.19674811.82%
  • OndoOndo(ONDO)$0.4069550.69%
  • mantleMantle(MNT)$0.59-3.31%
  • pax-goldPAX Gold(PAXG)$4,360.77-0.15%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Kyutai Open Sources Moshi: A Real-Time Native Multimodal Foundation AI Model that can Listen and Speak

July 3, 2024
in AI & Technology
Reading Time: 4 mins read
A A
Kyutai Open Sources Moshi: A Real-Time Native Multimodal Foundation AI Model that can Listen and Speak
ShareShareShareShareShare

In a stunning announcement reverberating through the tech world, Kyutai introduced Moshi, a revolutionary real-time native multimodal foundation model. This innovative model mirrors and surpasses some of the functionalities showcased by OpenAI’s GPT-4o in May.

Moshi is designed to understand and express emotions, offering capabilities like speaking with different accents, including French. It can listen and generate audio and speech while maintaining a seamless flow of textual thoughts, as it says. One of Moshi’s standout features is its ability to handle two audio streams simultaneously, allowing it to listen and talk simultaneously. This real-time interaction is underpinned by joint pre-training on a mix of text and audio, leveraging synthetic text data from Helium, a 7 billion parameter language model developed by Kyutai.

YOU MAY ALSO LIKE

How Long Can You Expect Your Old Cassette Tapes To Last?

How To Record Audio On Your iPhone

The fine-tuning process of Moshi involved 100,000 “oral-style” synthetic conversations, converted using Text-to-Speech (TTS) technology. The model’s voice was trained on synthetic data generated by a separate TTS model, achieving an impressive end-to-end latency of 200 milliseconds. Remarkably, Kyutai has also developed a smaller variant of Moshi that can run on a MacBook or a consumer-sized GPU, making it accessible to a broader range of users.

Kyutai has emphasized the importance of responsible AI use by incorporating watermarking to detect AI-generated audio, a feature that is currently a work in progress. The decision to release Moshi as an open-source project highlights Kyutai’s commitment to transparency and collaborative development within the AI community.

At its core, Moshi is powered by a 7-billion-parameter multimodal language model that processes speech input and output. The model operates with a two-channel I/O system, generating text tokens and audio codecs concurrently. The base text language model, Helium 7B, was trained from scratch and then jointly trained with text and audio codecs. Based on Kyutai’s in-house Mimi model, the speech codec boasts a 300x compression factor, capturing semantic and acoustic information.

Training Moshi involved rigorous processes, fine-tuning 100,000 highly detailed transcripts annotated with emotion and style. The Text-to-Speech Engine, which supports 70 different emotions and styles, was fine-tuned on 20 hours of audio recorded by a licensed voice talent named Alice. The model is designed for adaptability and can be fine-tuned with less than 30 minutes of audio.

Moshi’s deployment showcases its efficiency. The demo model, hosted on Scaleway and Hugging Face platforms, can handle two batch sizes at 24 GB VRAM. It supports various backends, including CUDA, Metal, and CPU, and benefits from optimizations in inference code through Rust. Enhanced KV caching and prompt caching are anticipated to improve performance further.

Looking ahead, Kyutai has ambitious plans for Moshi. The team intends to release a technical report and open model versions, including the inference codebase, the 7B model, the audio codec, and the full optimized stack. Future iterations, such as Moshi 1.1, 1.2, and 2.0, will refine the model based on user feedback. Moshi’s licensing aims to be as permissive as possible, fostering widespread adoption and innovation.

In conclusion, Moshi exemplifies the potential of small, focused teams to achieve extraordinary advancements in AI technology. This model opens up new avenues for research assistance, brainstorming, language learning, and more, demonstrating the transformative power of AI when deployed on-device with unparalleled flexibility. As an open-source model, it invites collaboration and innovation, ensuring that the benefits of this groundbreaking technology are accessible to all.


Check out the Announcement, Keynote, and Demo Chat. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. 

Paper, Code, and Model are coming…

Join our Telegram Channel and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 46k+ ML SubReddit


Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.

🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

How Long Can You Expect Your Old Cassette Tapes To Last?
AI & Technology

How Long Can You Expect Your Old Cassette Tapes To Last?

September 20, 2026
How To Record Audio On Your iPhone
AI & Technology

How To Record Audio On Your iPhone

September 20, 2026
What Is The Difference Between Apple CarPlay And CarPlay Ultra?
AI & Technology

What Is The Difference Between Apple CarPlay And CarPlay Ultra?

September 19, 2026
OpenClaw Releases 2026.9.5 With Atomic Updates, Plugin Hot Reload, Conversation Sharing, and Expanded GPT Live
AI & Technology

OpenClaw Releases 2026.9.5 With Atomic Updates, Plugin Hot Reload, Conversation Sharing, and Expanded GPT Live

September 19, 2026
Next Post
White House defends Biden’s mental and physical fitness after debate performance

White House defends Biden’s mental and physical fitness after debate performance

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
GOP Sen. John Kennedy says the U.S. must ‘choke [Putin] to death’ as Trump admin meets with him

GOP Sen. John Kennedy says the U.S. must ‘choke [Putin] to death’ as Trump admin meets with him

September 16, 2026
Gemini AI guessed passwords, accessed protected systems of 3 companies

Gemini AI guessed passwords, accessed protected systems of 3 companies

September 19, 2026
Ryanair CEO apologizes for comparing European airline rivals to ‘rapists’ in wild rant

Ryanair CEO apologizes for comparing European airline rivals to ‘rapists’ in wild rant

September 18, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!