• bitcoinBitcoin(BTC)$78,442.00-1.47%
  • ethereumEthereum(ETH)$2,472.09-1.45%
  • tetherTether(USDT)$1.00-0.03%
  • binancecoinBNB(BNB)$751.790.62%
  • rippleXRP(XRP)$1.40-0.78%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$102.97-2.58%
  • tronTRON(TRX)$0.3388350.99%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.060.00%
  • zcashZcash(ZEC)$1,157.18-3.04%
  • HyperliquidHyperliquid(HYPE)$83.24-5.36%
  • dogecoinDogecoin(DOGE)$0.089317-2.10%
  • RainRain(RAIN)$0.0167591.36%
  • USDSUSDS(USDS)$1.00-0.01%
  • moneroMonero(XMR)$505.10-5.92%
  • whitebitWhiteBIT Coin(WBT)$79.758.71%
  • chainlinkChainlink(LINK)$12.53-5.75%
  • leo-tokenLEO Token(LEO)$9.180.01%
  • cardanoCardano(ADA)$0.218767-2.56%
  • stellarStellar(XLM)$0.188546-3.25%
  • bitcoin-cashBitcoin Cash(BCH)$256.11-1.94%
  • daiDai(DAI)$1.000.02%
  • Ethena USDeEthena USDe(USDE)$1.00-0.02%
  • uniswapUniswap(UNI)$6.99-1.93%
  • litecoinLitecoin(LTC)$55.48-5.17%
  • USD1USD1(USD1)$1.00-0.03%
  • CantonCanton(CC)$0.104070-3.99%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.39-1.68%
  • hedera-hashgraphHedera(HBAR)$0.080196-3.37%
  • avalanche-2Avalanche(AVAX)$8.03-0.28%
  • suiSui(SUI)$0.81-2.84%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • shiba-inuShiba Inu(SHIB)$0.000005-1.77%
  • nearNEAR Protocol(NEAR)$2.30-3.82%
  • crypto-com-chainCronos(CRO)$0.0607565.26%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,399.940.04%
  • MemeCoreMemeCore(M)$1.184.52%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • BittensorBittensor(TAO)$252.90-4.86%
  • okbOKB(OKB)$115.11-1.36%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.15-0.41%
  • mantleMantle(MNT)$0.63-2.12%
  • AsterAster(ASTER)$0.76-5.12%
  • aaveAave(AAVE)$129.15-4.30%
  • pax-goldPAX Gold(PAXG)$4,404.450.05%
  • polkadotPolkadot(DOT)$1.085.27%
  • OndoOndo(ONDO)$0.376707-3.95%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Uni3D: Exploring Unified 3D Representation at Scale

October 27, 2023
in AI & Technology
Reading Time: 6 mins read
A A
Uni3D: Exploring Unified 3D Representation at Scale
ShareShareShareShareShare

Scaling up representations of text and visuals has been a major focus of research in recent years. Developments and research conducted in the recent past have led to numerous revolutions in language learning and vision. However, despite the popularity of scaling text and visual representations, the scaling of representations for 3D scenes and objects has not been sufficiently discussed.

Today, we will discuss Uni3D, a 3D foundation model that aims to explore unified 3D representations. The Uni3D framework employs a 2D-initialized ViT framework, pretrained end-to-end, to align image-text features with their corresponding 3D point cloud features.

The Uni3D framework uses pretext tasks and a simple architecture to leverage the abundance of pretrained 2D models and image-text-aligned models as initializations and targets, respectively. This approach unleashes the full potential of 2D models and strategies to scale them to the 3D world.

In this article, we will delve deeper into 3D computer vision and the Uni3D framework, exploring the essential concepts and the architecture of the model. So, let’s begin.

In the past few years, computer vision has emerged as one of the most heavily invested domains in the AI industry. Following significant advancements in 2D computer vision frameworks, developers have shifted their focus to 3D computer vision. This field, particularly 3D representation learning, merges aspects of computer graphics, machine learning, computer vision, and mathematics to automate the processing and understanding of 3D geometry. The rapid development of 3D sensors like LiDAR, along with their widespread applications in the AR/VR industry, has resulted in 3D representation learning gaining increased attention. Its potential applications continue to grow daily.

Although existing frameworks have shown remarkable progress in 3D model architecture, task-oriented modeling, and learning objectives, most explore 3D architecture on a relatively small scale with limited data, parameters, and task scenarios. The challenge of learning scalable 3D representations, which can then be applied to real-time applications in diverse environments, remains largely unexplored.

Moving along, in the past few years, scaling large language models that are pre-trained has helped in revolutionizing the natural language processing domain, and recent works have indicated a translation in the progress to 2D from language using data and model scaling which makes way for developers to try & reattempt this success to learn a 3D representation that can be scaled & be transferred to applications in real world. 

Uni3D is a scalable and unified pretraining 3D framework developed with the aim to learn large-scale 3D representations that tests its limits at the scale of over a billion parameters, over 10 million images paired with over 70 million texts, and over a million 3D shapes. The figure below compares the zero-shot accuracy against parameters in the Uni3D framework. The Uni3D framework successfully scales 3D representations from 6 million to over a billion. 

The Uni3D framework consists of a 2D ViT or Vision Transformer as the 3D encoder that is then pre-trained end-to-end to align the image-text aligned features with the 3D point cloud features. The Uni3D framework makes use of pretext tasks and  simple architecture to leverage the abundance of pretrained 2D models and image text aligned models as initialization and targets respectively, thus unleashing the full potential of 2D models, and strategies to scale them to the 3D world. The flexibility & scalability of the Uni3D framework is measured in terms of

  1. Scaling the model from 6M to over a billion parameters. 
  2. 2D initialization to text supervised from visual self-supervised learning. 
  3. Text-image target model scaling from 150 million to over a billion parameters. 

Under the flexible and unified framework offered by Uni3D, developers observe a coherent boost in the performance when it comes to scaling each component. The large-scale 3D representation learning also benefits immensely from the sharable 2D and scale-up strategies. 

As it can be seen in the figure below, the Uni3D framework displays a boost in the performance when compared to prior art in few-shot and zero-shot settings. It is worth noting that the Uni3D framework returns a zero-shot classification accuracy score of over 88% on ModelNet which is at par with the performance of several state of the art supervision methods. 

Furthermore, the Uni3D framework also delivers top notch accuracy & performance when performing other representative 3D tasks like part segmentation, and open world understanding. The Uni3D framework aims to bridge the gap between 2D vision and 3D vision by scaling 3D foundational models with a unified yet simple pre-training approach to learn more robust 3D representations across a wide array of tasks, that might ultimately help in the convergence of 2D and 3D vision across a wide array of modalities.

Uni3D : Related Work

The Uni3D framework draws inspiration, and learns from the developments made by previous 3D representation learning, and Foundational models especially under different modalities. 

3D Representation Learning

The 3D representation learning method uses cloud points for 3D understanding of the object, and this field has been explored by developers a lot in the recent past, and it has been observed that these cloud points can be pre-trained under self-supervision using specific 3D pretext tasks including mask point modeling, self-reconstruction, and contrastive learning. 

It is worth noting that these methods work with limited data, and they often do not investigate multimodal representations to 3D from 2D or NLP. However, the recent success of the CLIP framework that returns high efficiency in learning visual concepts from raw text using the contrastive learning method, and further seeks to learn 3D representations by aligning image, text, and cloud point features using the same contrastive learning method. 

Foundation Models

Developers have exhaustively been working on designing foundation models to scale up and unify multimodal representations. For example, in the NLP domain, developers have been working on frameworks that can scale up pre-trained language models, and it is slowly revolutionizing the NLP industry. Furthermore, advancements can be observed in the 2D vision domain as well because developers are working on frameworks that use data & model scaling techniques to help in the progress of language to 2D models, although such frameworks are difficult to replicate for 3D models because of the limited availability of 3D data, and the challenges encountered when unifying & scaling up the 3D frameworks. 

By learning from the above two work domains, developers have created the Uni3D framework, the first 3D foundation model with over a billion parameters that makes use of a unified ViT or Vision Transformer architecture that allows developers to scale the Uni3D model using unified 3D or NLP strategies for scaling up the models. Developers hope that this method will allow the Uni3D framework to bridge the gap that currently separates 2D and 3D vision along with facilitating multimodal convergence. 

Uni3D : Method and Architecture

The above image demonstrates the generic overview of the Uni3D framework, a scalable and unified pre-training 3D framework for large-scale 3D representation learning. Developers make use of over 70 million texts, and 10 million images paired with over a million 3D shapes to scale the Uni3D framework to over a billion parameters. The Uni3D framework uses a 2D ViT or Vision Transformer as a 3D encoder that is then trained end-to-end to align the text-image data with the 3D cloud point features, allowing the Uni3D framework to deliver the desired efficiency & accuracy across a wide array of benchmarks. Let us now have a detailed look at the working of the Uni3D framework. 

Scaling the Uni3D Framework

Prior studies on cloud point representation learning have traditionally focused heavily on designing particular model architectures that deliver better performance across a wide range of applications, and work on a limited amount of data thanks to small-scale datasets. However, recent studies have tried exploring the possibility of using scalable pre-training in 3D but there were no major outcomes thanks to the availability of limited 3D data. To solve the scalability problem of 3D frameworks, the Uni3D framework leverages the power of a vanilla transformer structure that almost mirrors a Vision Transformer, and can solve the scaling problems by using unified 2D or NLP scaling-up strategies to scale the model size. 

YOU MAY ALSO LIKE

What Is The Anker ‘Smart Display Charger’ And What Does That Screen Even Do?

What Is Retrieval-Augmented Generation (RAG)? How AI Answers with External Knowledge – Unite.AI

Credit: Source link

ShareTweetSendSharePin

Related Posts

What Is The Anker ‘Smart Display Charger’ And What Does That Screen Even Do?
AI & Technology

What Is The Anker ‘Smart Display Charger’ And What Does That Screen Even Do?

September 8, 2026
What Is Retrieval-Augmented Generation (RAG)? How AI Answers with External Knowledge – Unite.AI
AI & Technology

What Is Retrieval-Augmented Generation (RAG)? How AI Answers with External Knowledge – Unite.AI

September 8, 2026
Motional Releases nuReasoning Dataset and Launches ECCV Challenge – Unite.AI
AI & Technology

Motional Releases nuReasoning Dataset and Launches ECCV Challenge – Unite.AI

September 8, 2026
Renault Is Building Its €17,900 Dacia Spring EV In Europe To Qualify For Local Subsidies
AI & Technology

Renault Is Building Its €17,900 Dacia Spring EV In Europe To Qualify For Local Subsidies

September 8, 2026
Next Post
Giant #Sandstorm Sweeps Across Northwestern China

Giant #Sandstorm Sweeps Across Northwestern China

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Axis Robotics Releases AXIS: A Browser-Based Data Engine With 207 Robot Manipulation Tasks and 50,129 Trajectories

Axis Robotics Releases AXIS: A Browser-Based Data Engine With 207 Robot Manipulation Tasks and 50,129 Trajectories

September 7, 2026
Retail Is Beating the S&P — What They’re Buying Right Now

Retail Is Beating the S&P — What They’re Buying Right Now

September 3, 2026
OpenBMB Releases MiniCPM5-2B: A 2.52B Dense Model Averaging 53.9 Across 34 Benchmarks and Built to Run On Device

OpenBMB Releases MiniCPM5-2B: A 2.52B Dense Model Averaging 53.9 Across 34 Benchmarks and Built to Run On Device

September 7, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!