• bitcoinBitcoin(BTC)$77,198.00-0.04%
  • ethereumEthereum(ETH)$2,519.340.29%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$726.50-0.99%
  • rippleXRP(XRP)$1.360.00%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$101.680.01%
  • tronTRON(TRX)$0.339385-0.13%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.00-3.07%
  • zcashZcash(ZEC)$1,140.250.52%
  • HyperliquidHyperliquid(HYPE)$79.080.76%
  • dogecoinDogecoin(DOGE)$0.0846560.39%
  • RainRain(RAIN)$0.0157513.07%
  • moneroMonero(XMR)$531.960.33%
  • USDSUSDS(USDS)$1.00-0.01%
  • whitebitWhiteBIT Coin(WBT)$80.220.09%
  • chainlinkChainlink(LINK)$11.49-0.05%
  • leo-tokenLEO Token(LEO)$9.06-0.75%
  • cardanoCardano(ADA)$0.207101-0.44%
  • stellarStellar(XLM)$0.179827-0.36%
  • Ethena USDeEthena USDe(USDE)$1.00-0.02%
  • daiDai(DAI)$1.000.00%
  • bitcoin-cashBitcoin Cash(BCH)$225.63-1.74%
  • USD1USD1(USD1)$1.000.00%
  • litecoinLitecoin(LTC)$53.840.32%
  • uniswapUniswap(UNI)$6.373.48%
  • CantonCanton(CC)$0.097828-0.76%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.370.41%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • hedera-hashgraphHedera(HBAR)$0.0751020.78%
  • avalanche-2Avalanche(AVAX)$7.40-0.75%
  • shiba-inuShiba Inu(SHIB)$0.0000050.59%
  • nearNEAR Protocol(NEAR)$2.350.05%
  • suiSui(SUI)$0.72-0.11%
  • crypto-com-chainCronos(CRO)$0.0596113.99%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • MemeCoreMemeCore(M)$1.18-1.46%
  • tether-goldTether Gold(XAUT)$4,348.19-0.02%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • okbOKB(OKB)$114.15-0.20%
  • BittensorBittensor(TAO)$236.020.88%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.09%
  • aaveAave(AAVE)$127.011.61%
  • pax-goldPAX Gold(PAXG)$4,354.130.00%
  • AsterAster(ASTER)$0.691.51%
  • mantleMantle(MNT)$0.56-4.62%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0571923.77%
  • polkadotPolkadot(DOT)$1.01-4.11%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Meta and UC Berkeley Researchers Present Audio2Photoreal: An Artificial Intelligence Framework for Generating Full-Bodied Photorealistic Avatars that Gesture According to the Conversational Dynamics

January 12, 2024
in AI & Technology
Reading Time: 6 mins read
A A
Meta and UC Berkeley Researchers Present Audio2Photoreal: An Artificial Intelligence Framework for Generating Full-Bodied Photorealistic Avatars that Gesture According to the Conversational Dynamics
ShareShareShareShareShare

Avatar technology has become ubiquitous in platforms like Snapchat, Instagram, and video games, enhancing user engagement by replicating human actions and emotions. However, the quest for a more immersive experience led researchers from Meta and BAIR to introduce “Audio2Photoreal,” a groundbreaking method for synthesizing photorealistic avatars capable of natural conversations.

Imagine engaging in a telepresent conversation with a friend represented by a photorealistic 3D model, dynamically expressing emotions aligned with their speech. The challenge lies in overcoming the limitations of non-textured meshes, which fail to capture subtle nuances like eye gaze or smirking, resulting in a robotic and uncanny interaction (see Figure 1, middle). The research aims to bridge this gap, presenting a method for generating photorealistic avatars based on the speech audio of a dyadic conversation.

YOU MAY ALSO LIKE

Implementation of Machine Learning Workflows with NVIDIA cuML, RAPIDS, GPU Benchmarking, Explainability, Clustering, and Model Inference

Hyundai Motor Group Puts Data Flywheel Into Full Operation – Unite.AI

Reference: https://arxiv.org/pdf/2401.01885.pdf 

The approach involves synthesizing diverse high-frequency gestures and expressive facial movements synchronized with speech. Leveraging both an autoregressive VQ-based method and a diffusion model for body and hands, the researchers achieve a balance between frame rate and motion details. The result is a system that renders photorealistic avatars capable of conveying intricate facial, body, and hand motions in real time.

To support this research, the team introduces a unique multi-view conversational dataset, providing a photorealistic reconstruction of non-scripted, long-form conversations. Unlike previous datasets focused on upper body or facial motion, this dataset captures the dynamics of interpersonal conversations, offering a more comprehensive understanding of conversational gestures.

Reference: https://arxiv.org/pdf/2401.01885.pdf 

The system employs a two-model (shown in Figure 3) approach for face and body motion synthesis, each addressing the unique dynamics of these components. The face motion model (Figure 4(a)), a diffusion model conditioned on input audio and lip vertices, focuses on generating speech-consistent facial details. In contrast, the body motion model uses an autoregressive audio-conditioned transformer to predict coarse guide poses (Figure 4(b)) at 1fps, later refined by the diffusion model (Figure 4(c)) for diverse yet plausible body motions.

Reference: https://arxiv.org/pdf/2401.01885.pdf 

The evaluation demonstrates the model’s effectiveness (shown in Figure 6) in generating realistic and diverse conversational motions, outperforming various baselines. Photorealism proves crucial in capturing subtle nuances, as highlighted in perceptual evaluations. The quantitative results showcase the method’s ability to balance realism and diversity, surpassing prior works in terms of motion quality.

While the model excels in generating compelling and plausible gestures, it operates on short-range audio, limiting its capability for long-range language understanding. Additionally, the ethical considerations of consent are addressed by rendering only consenting participants in the dataset.

Reference: https://arxiv.org/pdf/2401.01885.pdf 

In conclusion, “Audio2Photoreal” represents a significant leap in synthesizing conversational avatars, offering a more immersive and realistic experience. The research not only introduces a novel dataset and methodology but also opens avenues for exploring ethical considerations in photorealistic motion synthesis.


Check out the Paper and Project. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. Join our 36k+ ML SubReddit, 41k+ Facebook Community, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..


Vineet Kumar is a consulting intern at MarktechPost. He is currently pursuing his BS from the Indian Institute of Technology(IIT), Kanpur. He is a Machine Learning enthusiast. He is passionate about research and the latest advancements in Deep Learning, Computer Vision, and related fields.


[Partnership and Promotion on Marktechpost] 🐝 Now you can partner with Marktechpost to promote your Research Paper, Github Repo and even add your pro commentary in any trending research article on marktechpost.com. Elevate your and your company’s AI research visibility in the tech community…Learn more


Credit: Source link

ShareTweetSendSharePin

Related Posts

Implementation of Machine Learning Workflows with NVIDIA cuML, RAPIDS, GPU Benchmarking, Explainability, Clustering, and Model Inference
AI & Technology

Implementation of Machine Learning Workflows with NVIDIA cuML, RAPIDS, GPU Benchmarking, Explainability, Clustering, and Model Inference

September 13, 2026
Hyundai Motor Group Puts Data Flywheel Into Full Operation – Unite.AI
AI & Technology

Hyundai Motor Group Puts Data Flywheel Into Full Operation – Unite.AI

September 13, 2026
What Is The Difference Between A Dead Pixel And A Stuck Pixel?
AI & Technology

What Is The Difference Between A Dead Pixel And A Stuck Pixel?

September 13, 2026
Cognition Releases SWE-2: A Kimi K3 Post-Trained Coding Model That Matches Fable 5.1 on FrontierCode at 64% Lower Cost
AI & Technology

Cognition Releases SWE-2: A Kimi K3 Post-Trained Coding Model That Matches Fable 5.1 on FrontierCode at 64% Lower Cost

September 12, 2026
Next Post
Nigerians plead not guilty in Michigan sextortion suicide case

Nigerians plead not guilty in Michigan sextortion suicide case

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Reported tornadoes hit New York City area

Reported tornadoes hit New York City area

September 7, 2026
What Is The Anker ‘Smart Display Charger’ And What Does That Screen Even Do?

What Is The Anker ‘Smart Display Charger’ And What Does That Screen Even Do?

September 8, 2026
Bertha pushes west after battering Louisiana

Bertha pushes west after battering Louisiana

September 6, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!