• bitcoinBitcoin(BTC)$84,557.00-1.84%
  • ethereumEthereum(ETH)$2,690.04-2.33%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$767.65-2.49%
  • rippleXRP(XRP)$1.50-4.50%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$115.07-2.87%
  • tronTRON(TRX)$0.341372-0.06%
  • zcashZcash(ZEC)$1,495.03-3.68%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.03-0.89%
  • HyperliquidHyperliquid(HYPE)$93.98-3.27%
  • dogecoinDogecoin(DOGE)$0.092597-7.80%
  • moneroMonero(XMR)$550.27-2.80%
  • whitebitWhiteBIT Coin(WBT)$84.93-1.98%
  • USDSUSDS(USDS)$1.00-0.01%
  • chainlinkChainlink(LINK)$12.36-4.96%
  • cardanoCardano(ADA)$0.238620-5.74%
  • RainRain(RAIN)$0.012250-6.43%
  • leo-tokenLEO Token(LEO)$9.010.38%
  • stellarStellar(XLM)$0.202137-6.61%
  • bitcoin-cashBitcoin Cash(BCH)$338.380.11%
  • uniswapUniswap(UNI)$9.27-9.15%
  • nearNEAR Protocol(NEAR)$4.33-1.49%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • litecoinLitecoin(LTC)$61.66-2.28%
  • daiDai(DAI)$1.000.01%
  • avalanche-2Avalanche(AVAX)$10.32-8.46%
  • USD1USD1(USD1)$1.000.00%
  • CantonCanton(CC)$0.109633-4.35%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.42-2.93%
  • hedera-hashgraphHedera(HBAR)$0.090473-8.68%
  • suiSui(SUI)$0.96-5.80%
  • shiba-inuShiba Inu(SHIB)$0.000006-7.58%
  • BittensorBittensor(TAO)$287.04-8.20%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • crypto-com-chainCronos(CRO)$0.061317-8.04%
  • BitwayBitway(BTW)$1.0313.59%
  • MemeCoreMemeCore(M)$1.22-7.04%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • tether-goldTether Gold(XAUT)$4,290.05-1.64%
  • okbOKB(OKB)$118.67-3.30%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.29%
  • mantleMantle(MNT)$0.65-2.12%
  • aaveAave(AAVE)$138.95-5.23%
  • EthenaEthena(ENA)$0.205647-4.64%
  • OndoOndo(ONDO)$0.414017-5.44%
  • AsterAster(ASTER)$0.69-5.25%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Researchers from Meta AI and UT Austin Explored Scaling in Auto-Encoders and Introduced ViTok: A ViT-Style Auto-Encoder to Perform Exploration

January 18, 2025
in AI & Technology
Reading Time: 6 mins read
A A
Researchers from Meta AI and UT Austin Explored Scaling in Auto-Encoders and Introduced ViTok: A ViT-Style Auto-Encoder to Perform Exploration
ShareShareShareShareShare

Modern image and video generation methods rely heavily on tokenization to encode high-dimensional data into compact latent representations. While advancements in scaling generator models have been substantial, tokenizers—primarily based on convolutional neural networks (CNNs)—have received comparatively less attention. This raises questions about how scaling tokenizers might improve reconstruction accuracy and generative tasks. Challenges include architectural limitations and constrained datasets, which affect scalability and broader applicability. There is also a need to understand how design choices in auto-encoders influence performance metrics such as fidelity, compression, and generation.

Researchers from Meta and UT Austin have addressed these issues by introducing ViTok, a Vision Transformer (ViT)-based auto-encoder. Unlike traditional CNN-based tokenizers, ViTok employs a Transformer-based architecture enhanced by the Llama framework. This design supports large-scale tokenization for images and videos, overcoming dataset constraints by training on extensive and diverse data.

YOU MAY ALSO LIKE

Microsoft’s New Surface Pro 12 And Surface Laptop 13 Feature Snapdragon X2 Plus Chips

Google Releases Gemini 3.8 Flash TTS and Flash-Lite TTS With Prompt-Based Voice Design

ViTok focuses on three aspects of scaling:

  1. Bottleneck scaling: Examining the relationship between latent code size and performance.
  2. Encoder scaling: Evaluating the impact of increasing encoder complexity.
  3. Decoder scaling: Assessing how larger decoders influence reconstruction and generation.

These efforts aim to optimize visual tokenization for both images and videos by addressing inefficiencies in existing architectures.

Technical Details and Advantages of ViTok

ViTok uses an asymmetric auto-encoder framework with several distinctive features:

  1. Patch and Tubelet Embedding: Inputs are divided into patches (for images) or tubelets (for videos) to capture spatial and spatiotemporal details.
  2. Latent Bottleneck: The size of the latent space, defined by the number of floating points (E), determines the balance between compression and reconstruction quality.
  3. Encoder and Decoder Design: ViTok employs a lightweight encoder for efficiency and a more computationally intensive decoder for robust reconstruction.

By leveraging Vision Transformers, ViTok improves scalability. Its enhanced decoder incorporates perceptual and adversarial losses to produce high-quality outputs. Together, these components enable ViTok to:

  • Achieve effective reconstruction with fewer computational FLOPs.
  • Handle image and video data efficiently, taking advantage of the redundancy in video sequences.
  • Balance trade-offs between fidelity (e.g., PSNR, SSIM) and perceptual quality (e.g., FID, IS).

Results and Insights

ViTok’s performance was evaluated using benchmarks such as ImageNet-1K, COCO for images, and UCF-101 for videos. Key findings include:

  • Bottleneck Scaling: Increasing bottleneck size improves reconstruction but can complicate generative tasks if the latent space is too large.
  • Encoder Scaling: Larger encoders show minimal benefits for reconstruction and may hinder generative performance due to increased decoding complexity.
  • Decoder Scaling: Larger decoders enhance reconstruction quality, but their benefits for generative tasks vary. A balanced design is often required.

Results highlight ViTok’s strengths in efficiency and accuracy:

  • State-of-the-art metrics for image reconstruction at 256p and 512p resolutions.
  • Improved video reconstruction scores, demonstrating adaptability to spatiotemporal data.
  • Competitive generative performance in class-conditional tasks with reduced computational demands.

Conclusion

ViTok offers a scalable, Transformer-based alternative to traditional CNN tokenizers, addressing key challenges in bottleneck design, encoder scaling, and decoder optimization. Its robust performance across reconstruction and generation tasks highlights its potential for a wide range of applications. By effectively handling both image and video data, ViTok underscores the importance of thoughtful architectural design in advancing visual tokenization.


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter and join our Telegram Channel and LinkedIn Group. Don’t Forget to join our 65k+ ML SubReddit.

🚨 Recommend Open-Source Platform: Parlant is a framework that transforms how AI agents make decisions in customer-facing scenarios. (Promoted)


Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.

📄 Meet ‘Height’:The only autonomous project management tool (Sponsored)

Credit: Source link

ShareTweetSendSharePin

Related Posts

Microsoft’s New Surface Pro 12 And Surface Laptop 13 Feature Snapdragon X2 Plus Chips
AI & Technology

Microsoft’s New Surface Pro 12 And Surface Laptop 13 Feature Snapdragon X2 Plus Chips

September 23, 2026
Google Releases Gemini 3.8 Flash TTS and Flash-Lite TTS With Prompt-Based Voice Design
AI & Technology

Google Releases Gemini 3.8 Flash TTS and Flash-Lite TTS With Prompt-Based Voice Design

September 23, 2026
NVIDIA Releases Nemotron 3 Diarization: A 100M-Parameter Open-Weight Model That Tracks 8 Speakers in Real Time
AI & Technology

NVIDIA Releases Nemotron 3 Diarization: A 100M-Parameter Open-Weight Model That Tracks 8 Speakers in Real Time

September 23, 2026
Disney+ And Hulu Are Getting Even More Expensive (Again)
AI & Technology

Disney+ And Hulu Are Getting Even More Expensive (Again)

September 23, 2026
Next Post
Range Resources Stock: 2025 Free Cash Flow May Exceed 0 Million (NYSE:RRC)

Range Resources Stock: 2025 Free Cash Flow May Exceed $900 Million (NYSE:RRC)

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
XTN: Transportation Likely To Lag Into 2027 Amid Macro Pressures And Factor Weaknesses

XTN: Transportation Likely To Lag Into 2027 Amid Macro Pressures And Factor Weaknesses

September 22, 2026
Berkshire Hathaway names Warren Buffett as chairman emeritus

Berkshire Hathaway names Warren Buffett as chairman emeritus

September 18, 2026
GSA Extends Anthropic’s Claude OneGov Offer for Federal Agencies – Unite.AI

GSA Extends Anthropic’s Claude OneGov Offer for Federal Agencies – Unite.AI

September 17, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!