• bitcoinBitcoin(BTC)$81,261.000.48%
  • ethereumEthereum(ETH)$2,634.630.94%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$762.210.07%
  • rippleXRP(XRP)$1.411.13%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$111.07-1.49%
  • tronTRON(TRX)$0.3397840.44%
  • zcashZcash(ZEC)$1,471.76-6.29%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.02-1.43%
  • HyperliquidHyperliquid(HYPE)$92.00-0.63%
  • dogecoinDogecoin(DOGE)$0.0879030.29%
  • moneroMonero(XMR)$546.08-3.12%
  • whitebitWhiteBIT Coin(WBT)$82.91-0.13%
  • RainRain(RAIN)$0.0137542.90%
  • USDSUSDS(USDS)$1.00-0.02%
  • chainlinkChainlink(LINK)$12.431.66%
  • cardanoCardano(ADA)$0.2278891.58%
  • leo-tokenLEO Token(LEO)$8.90-0.09%
  • stellarStellar(XLM)$0.1966732.12%
  • uniswapUniswap(UNI)$8.66-2.12%
  • bitcoin-cashBitcoin Cash(BCH)$256.440.61%
  • Ethena USDeEthena USDe(USDE)$1.00-0.02%
  • nearNEAR Protocol(NEAR)$3.58-4.14%
  • daiDai(DAI)$1.000.02%
  • litecoinLitecoin(LTC)$58.090.38%
  • avalanche-2Avalanche(AVAX)$10.1423.39%
  • USD1USD1(USD1)$1.00-0.02%
  • CantonCanton(CC)$0.109411-1.78%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.380.61%
  • hedera-hashgraphHedera(HBAR)$0.0815913.05%
  • suiSui(SUI)$0.865.84%
  • MemeCoreMemeCore(M)$1.5115.23%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • shiba-inuShiba Inu(SHIB)$0.0000061.05%
  • BittensorBittensor(TAO)$264.066.68%
  • crypto-com-chainCronos(CRO)$0.059523-0.36%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.03%
  • tether-goldTether Gold(XAUT)$4,373.76-0.06%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • okbOKB(OKB)$117.751.17%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.15-0.29%
  • aaveAave(AAVE)$141.441.40%
  • AsterAster(ASTER)$0.76-0.68%
  • EthenaEthena(ENA)$0.20393020.90%
  • mantleMantle(MNT)$0.62-0.37%
  • OndoOndo(ONDO)$0.4205555.61%
  • Pump.funPump.fun(PUMP)$0.004248-0.54%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

NYU’s new AI architecture makes high-quality image generation faster and cheaper

November 7, 2025
in AI & Technology
Reading Time: 4 mins read
A A
NYU’s new AI architecture makes high-quality image generation faster and cheaper
ShareShareShareShareShare

Researchers at New York University have developed a new architecture for diffusion models that improves the semantic representation of the images they generate. “Diffusion Transformer with Representation Autoencoders” (RAE) challenges some of the accepted norms of building diffusion models. The NYU researcher's model is more efficient and accurate than standard diffusion models, takes advantage of the latest research in representation learning and could pave the way for new applications that were previously too difficult or expensive.

YOU MAY ALSO LIKE

Trump Proposes Renaming Artificial Intelligence, Announces AI Force – Unite.AI

SpaceX Targets September 28 For Starship’s First Orbital Flight

This breakthrough could unlock more reliable and powerful features for enterprise applications. "To edit images well, a model has to really understand what’s in them," paper co-author Saining Xie told VentureBeat. "RAE helps connect that understanding part with the generation part." He also pointed to future applications in "RAG-based generation, where you use RAE encoder features for search and then generate new images based on the search results," as well as in "video generation and action-conditioned world models."

The state of generative modeling

Diffusion models, the technology behind most of today’s powerful image generators, frame generation as a process of learning to compress and decompress images. A variational autoencoder (VAE) learns a compact representation of an image’s key features in a so-called “latent space.” The model is then trained to generate new images by reversing this process from random noise.

While the diffusion part of these models has advanced, the autoencoder used in most of them has remained largely unchanged in recent years. According to the NYU researchers, this standard autoencoder (SD-VAE) is suitable for capturing low-level features and local appearance, but lacks the “global semantic structure crucial for generalization and generative performance.”

At the same time, the field has seen impressive advances in image representation learning with models such as DINO, MAE and CLIP. These models learn semantically-structured visual features that generalize across tasks and can serve as a natural basis for visual understanding. However, a widely-held belief has kept devs from using these architectures in image generation: Models focused on semantics are not suitable for generating images because they don’t capture granular, pixel-level features. Practitioners also believe that diffusion models do not work well with the kind of high-dimensional representations that semantic models produce.

Diffusion with representation encoders

The NYU researchers propose replacing the standard VAE with “representation autoencoders” (RAE). This new type of autoencoder pairs a pretrained representation encoder, like Meta’s DINO, with a trained vision transformer decoder. This approach simplifies the training process by using existing, powerful encoders that have already been trained on massive datasets.

To make this work, the team developed a variant of the diffusion transformer (DiT), the backbone of most image generation models. This modified DiT can be trained efficiently in the high-dimensional space of RAEs without incurring huge compute costs. The researchers show that frozen representation encoders, even those optimized for semantics, can be adapted for image generation tasks. Their method yields reconstructions that are superior to the standard SD-VAE without adding architectural complexity.

However, adopting this approach requires a shift in thinking. "RAE isn’t a simple plug-and-play autoencoder; the diffusion modeling part also needs to evolve," Xie explained. "One key point we want to highlight is that latent space modeling and generative modeling should be co-designed rather than treated separately."

With the right architectural adjustments, the researchers found that higher-dimensional representations are an advantage, offering richer structure, faster convergence and better generation quality. In their paper, the researchers note that these "higher-dimensional latents introduce effectively no extra compute or memory costs." Furthermore, the standard SD-VAE is more computationally expensive, requiring about six times more compute for the encoder and three times more for the decoder, compared to RAE.

Stronger performance and efficiency

The new model architecture delivers significant gains in both training efficiency and generation quality. The team's improved diffusion recipe achieves strong results after only 80 training epochs. Compared to prior diffusion models trained on VAEs, the RAE-based model achieves a 47x training speedup. It also outperforms recent methods based on representation alignment with a 16x training speedup. This level of efficiency translates directly into lower training costs and faster model development cycles.

For enterprise use, this translates into more reliable and consistent outputs. Xie noted that RAE-based models are less prone to semantic errors seen in classic diffusion, adding that RAE gives the model "a much smarter lens on the data." He observed that leading models like ChatGPT-4o and Google's Nano Banana are moving toward "subject-driven, highly consistent and knowledge-augmented generation," and that RAE's semantically rich foundation is key to achieving this reliability at scale and in open source models.

The researchers demonstrated this performance on the ImageNet benchmark. Using the Fréchet Inception Distance (FID) metric, where a lower score indicates higher-quality images, the RAE-based model achieved a state-of-the-art score of 1.51 without guidance. With AutoGuidance, a technique that uses a smaller model to steer the generation process, the FID score dropped to an even more impressive 1.13 for both 256×256 and 512×512 images.

By successfully integrating modern representation learning into the diffusion framework, this work opens a new path for building more capable and cost-effective generative models. This unification points toward a future of more integrated AI systems.

"We believe that in the future, there will be a single, unified representation model that captures the rich, underlying structure of reality… capable of decoding into many different output modalities," Xie said. He added that RAE offers a unique path toward this goal: "The high-dimensional latent space should be learned separately to provide a strong prior that can then be decoded into various modalities — rather than relying on a brute-force approach of mixing all data and training with multiple objectives at once."

Credit: Source link

ShareTweetSendSharePin

Related Posts

Trump Proposes Renaming Artificial Intelligence, Announces AI Force – Unite.AI
AI & Technology

Trump Proposes Renaming Artificial Intelligence, Announces AI Force – Unite.AI

September 19, 2026
SpaceX Targets September 28 For Starship’s First Orbital Flight
AI & Technology

SpaceX Targets September 28 For Starship’s First Orbital Flight

September 19, 2026
Now Trump Says He’s Creating An AI Force
AI & Technology

Now Trump Says He’s Creating An AI Force

September 19, 2026
TypeSafe AI Releases Jev: A System One Model That Returns Typed, Calibrated Decisions Instead of Text
AI & Technology

TypeSafe AI Releases Jev: A System One Model That Returns Typed, Calibrated Decisions Instead of Text

September 19, 2026
Next Post
Rise of Hydra has been delayed with no new release window

Rise of Hydra has been delayed with no new release window

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Starbucks resolves Florida DEI lawsuit with blockbuster companywide agreement

Starbucks resolves Florida DEI lawsuit with blockbuster companywide agreement

September 17, 2026
Car Manufacturers Are Ditching CarPlay In 2026: Here’s Why

Car Manufacturers Are Ditching CarPlay In 2026: Here’s Why

September 13, 2026
Man rescued after getting stuck in garbage truck

Man rescued after getting stuck in garbage truck

September 19, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!