• bitcoinBitcoin(BTC)$80,975.004.17%
  • ethereumEthereum(ETH)$2,512.704.53%
  • tetherTether(USDT)$1.000.03%
  • binancecoinBNB(BNB)$723.604.06%
  • rippleXRP(XRP)$1.456.17%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$103.833.08%
  • tronTRON(TRX)$0.3285530.65%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.031.92%
  • HyperliquidHyperliquid(HYPE)$85.844.48%
  • zcashZcash(ZEC)$955.6215.03%
  • dogecoinDogecoin(DOGE)$0.0869944.94%
  • RainRain(RAIN)$0.0170652.14%
  • USDSUSDS(USDS)$1.000.00%
  • moneroMonero(XMR)$509.66-0.73%
  • chainlinkChainlink(LINK)$11.886.11%
  • whitebitWhiteBIT Coin(WBT)$73.903.80%
  • leo-tokenLEO Token(LEO)$9.24-0.11%
  • cardanoCardano(ADA)$0.2214467.71%
  • stellarStellar(XLM)$0.1832333.45%
  • bitcoin-cashBitcoin Cash(BCH)$255.052.29%
  • daiDai(DAI)$1.000.00%
  • CantonCanton(CC)$0.1100930.65%
  • Ethena USDeEthena USDe(USDE)$1.000.02%
  • USD1USD1(USD1)$1.000.03%
  • litecoinLitecoin(LTC)$50.891.51%
  • uniswapUniswap(UNI)$6.3010.00%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.361.76%
  • hedera-hashgraphHedera(HBAR)$0.0779392.64%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • avalanche-2Avalanche(AVAX)$7.472.55%
  • suiSui(SUI)$0.770.07%
  • shiba-inuShiba Inu(SHIB)$0.0000051.84%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.0571804.16%
  • tether-goldTether Gold(XAUT)$4,458.420.93%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • nearNEAR Protocol(NEAR)$1.921.30%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • MemeCoreMemeCore(M)$1.04-1.83%
  • okbOKB(OKB)$109.583.64%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.12%
  • BittensorBittensor(TAO)$228.183.27%
  • aaveAave(AAVE)$133.684.69%
  • AsterAster(ASTER)$0.72-0.65%
  • pax-goldPAX Gold(PAXG)$4,468.000.86%
  • mantleMantle(MNT)$0.571.56%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0583003.91%
  • OndoOndo(ONDO)$0.3602892.58%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Meet MAGE, MIT’s unified system for image generation and recognition

June 21, 2023
in AI & Technology
Reading Time: 4 mins read
A A
Meet MAGE, MIT’s unified system for image generation and recognition
ShareShareShareShareShare

Join top executives in San Francisco on July 11-12, to hear how leaders are integrating and optimizing AI investments for success. Learn More


In a major development, researchers from MIT’s Computer Science and Artificial Intelligence Laboratory (CSAIL) have announced a framework that can handle both image recognition and image generation tasks with high accuracy. Officially dubbed Masked Generative Encoder, or MAGE, the unified computer vision system promises wide-ranging applications and can cut down on the overhead of training two separate systems for identifying images and generating fresh ones.

YOU MAY ALSO LIKE

The Ternus Era At Apple Begins, But Cook Isn’t Leaving

Mobile Games Designed to Be Addictive Get More Kid-Friendly

>>Follow VentureBeat’s ongoing generative AI coverage<<

The news comes at a time when enterprises are going all-in on AI, particularly generative technologies, for improving workflows. However, as the researchers explain, the MIT system still has some flaws and will need to be perfected in the coming months if it is to see adoption.

The team told VentureBeat that they also plan to expand the model’s capabilities.

Event

Transform 2023

Join us in San Francisco on July 11-12, where top executives will share how they have integrated and optimized AI investments for success and avoided common pitfalls.

 

Register Now

So, how does MAGE work?

Today, building image generation and recognition systems largely revolves around two processes: state-of-the-art generative modeling and self-supervised representation learning. In the former, the system learns to produce high-dimensional data from low-dimensional inputs such as class labels, text embeddings or random noise. In the latter, a high-dimensional image is used as an input to create a low-dimensional embedding for feature detection or classification. 

>>Don’t miss our special issue: Building the foundation for customer data quality.<<

These two techniques, currently used independently of each other, both require a visual and semantic understanding of data. So the team at MIT decided to bring them together in a unified architecture. MAGE is the result. 

To develop the system, the group used a pre-training approach called masked token modeling. They converted sections of image data into abstracted versions represented by semantic tokens. Each of these tokens represented a 16×16-token patch of the original image, acting like mini jigsaw puzzle pieces. 

Once the tokens were ready, some of them were randomly masked and a neural network was trained to predict the hidden ones by gathering the context from the surrounding tokens. That way, the system learned to understand the patterns in an image (image recognition) as well as generate new ones (image generation).

“Our key insight in this work is that generation is viewed as ‘reconstructing’ images that are 100% masked, while representation learning is viewed as ‘encoding’ images that are 0% masked,” the researchers wrote in a paper detailing the system. “The model is trained to reconstruct over a wide range of masking ratios covering high masking ratios that enable generation capabilities, and lower masking ratios that enable representation learning. This simple but very effective approach allows a smooth combination of generative training and representation learning in the same framework: same architecture, training scheme, and loss function.”

In addition to producing images from scratch, the system supports conditional image generation, where users can specify criteria for the images and the tool will cook up the appropriate image.

“The user can input a whole image and the system can understand and recognize the image, outputting the class of the image,” Tianhong Li, one of the researchers behind the system, told VentureBeat. “In other scenarios, the user can input an image with partial crops, and the system can recover the cropped image. They can also ask the system to generate a random image or generate an image given a certain class, such as a fish or dog.”

Potential for many applications

When pre-trained on data from the ImageNet image database, which consists of 1.3 million images, the model obtained a fréchet inception distance score (used to assess the quality of images) of 9.1, outperforming previous models. For recognition, it achieved an 80.9% accuracy rating in linear probing and a 71.9% 10-shot accuracy rating when it had only 10 labeled examples from each class.

“Our method can naturally scale up to any unlabeled image dataset,” Li said, noting that the model’s image understanding capabilities can be beneficial in scenarios where limited labeled data is available, such as in niche industries or emerging technologies.

Similarly, he said, the generation side of the model can help in industries like photo editing, visual effects and post-production with the its ability to remove elements from an image while maintaining a realistic appearance, or, given a specific class, replace an element with another generated element.

“It has [long] been a dream to achieve image generation and image recognition in one single system. MAGE is a [result of] groundbreaking research which successfully harnesses the synergy of these two tasks and achieves the state of the art of them in one single system,” said Huisheng Wang, senior software engineer for research and machine intelligence at Google, who participated in the MAGE project.

“This innovative system has wide-ranging applications, and has the potential to inspire many future works in the field of computer vision,” he added.

More work needed

Moving ahead, the team plans to streamline the MAGE system, especially the token conversion part of the process. Currently, when the image data is converted into tokens, some of the information is lost. Li and team plan to change that through other ways of compression.

Beyond this, Li said they also plan to scale up MAGE on real-world, large-scale unlabeled image datasets, and to apply it to multi-modality tasks, such as image-to-text and text-to-image generation.

VentureBeat’s mission is to be a digital town square for technical decision-makers to gain knowledge about transformative enterprise technology and transact. Discover our Briefings.

Credit: Source link

ShareTweetSendSharePin

Related Posts

The Ternus Era At Apple Begins, But Cook Isn’t Leaving
AI & Technology

The Ternus Era At Apple Begins, But Cook Isn’t Leaving

September 4, 2026
Mobile Games Designed to Be Addictive Get More Kid-Friendly
AI & Technology

Mobile Games Designed to Be Addictive Get More Kid-Friendly

September 4, 2026
Google DeepMind’s WeatherNext 3 Trains on Weather Station Observations to Deliver 5 km Global Forecasts, Refreshed Every Hour
AI & Technology

Google DeepMind’s WeatherNext 3 Trains on Weather Station Observations to Deliver 5 km Global Forecasts, Refreshed Every Hour

September 4, 2026
Nvidia Makes .5 Billion Bet on MediaTek
AI & Technology

Nvidia Makes $3.5 Billion Bet on MediaTek

September 4, 2026
Next Post
Is It Worth Moving For A Pay Increase for Work?

Is It Worth Moving For A Pay Increase for Work?

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
The Pros And Cons Of Flash Memory

The Pros And Cons Of Flash Memory

August 28, 2026
UFC Shanghai Results: Nurmagomedov vs. Song – MMA Fighting

UFC Shanghai Results: Nurmagomedov vs. Song – MMA Fighting

August 29, 2026
Income-Covered Closed-End Fund Report, August 2026

Income-Covered Closed-End Fund Report, August 2026

August 30, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!