• bitcoinBitcoin(BTC)$77,268.00-1.66%
  • ethereumEthereum(ETH)$2,412.64-2.28%
  • tetherTether(USDT)$1.00-0.02%
  • binancecoinBNB(BNB)$680.62-1.46%
  • rippleXRP(XRP)$1.35-2.52%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$99.64-3.46%
  • tronTRON(TRX)$0.322632-3.01%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.01-2.67%
  • HyperliquidHyperliquid(HYPE)$82.39-2.44%
  • zcashZcash(ZEC)$827.07-2.54%
  • dogecoinDogecoin(DOGE)$0.081503-1.80%
  • RainRain(RAIN)$0.016683-0.36%
  • USDSUSDS(USDS)$1.00-0.01%
  • moneroMonero(XMR)$495.81-5.10%
  • leo-tokenLEO Token(LEO)$9.37-2.20%
  • whitebitWhiteBIT Coin(WBT)$71.07-1.91%
  • chainlinkChainlink(LINK)$11.18-1.32%
  • cardanoCardano(ADA)$0.195475-1.63%
  • stellarStellar(XLM)$0.175001-1.31%
  • bitcoin-cashBitcoin Cash(BCH)$244.79-0.81%
  • daiDai(DAI)$1.00-0.02%
  • CantonCanton(CC)$0.113676-6.72%
  • Ethena USDeEthena USDe(USDE)$1.00-0.03%
  • USD1USD1(USD1)$1.00-0.02%
  • litecoinLitecoin(LTC)$49.552.02%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.31-6.42%
  • uniswapUniswap(UNI)$5.8311.42%
  • Global DollarGlobal Dollar(USDG)$1.00-0.04%
  • hedera-hashgraphHedera(HBAR)$0.0739090.43%
  • avalanche-2Avalanche(AVAX)$7.19-0.50%
  • shiba-inuShiba Inu(SHIB)$0.0000051.13%
  • suiSui(SUI)$0.72-1.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.04%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • crypto-com-chainCronos(CRO)$0.054723-3.72%
  • tether-goldTether Gold(XAUT)$4,332.18-2.38%
  • nearNEAR Protocol(NEAR)$1.89-0.87%
  • MemeCoreMemeCore(M)$1.06-3.24%
  • Ripple USDRipple USD(RLUSD)$1.00-0.02%
  • okbOKB(OKB)$110.26-1.90%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.22%
  • BittensorBittensor(TAO)$219.85-4.37%
  • aaveAave(AAVE)$126.782.12%
  • pax-goldPAX Gold(PAXG)$4,339.34-2.36%
  • AsterAster(ASTER)$0.69-1.47%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.056909-1.23%
  • MorphoMorpho(MORPHO)$2.582.70%
  • mantleMantle(MNT)$0.53-3.64%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Meet LLMScore: A New LLM-based Instruction-Following Matching Pipeline to Evaluate the Alignment Between Text Prompts and Synthesized Images in Text-to-Image Synthesis

May 31, 2023
in AI & Technology
Reading Time: 6 mins read
A A
Meet LLMScore: A New LLM-based Instruction-Following Matching Pipeline to Evaluate the Alignment Between Text Prompts and Synthesized Images in Text-to-Image Synthesis
ShareShareShareShareShare

Text-to-image synthesis research has advanced significantly in recent years. However, assessment measures have lagged due to difficulties adapting assessments with different purposes, effectively capturing composite text-image alignment (for example, color, counting, and position) and producing the score understandably. Despite being extensively used and successful, established assessment metrics for text-to-image synthesis like CLIPScore and BLIP have needed help capturing object-level alignment between text and picture. 

The text prompt “A red book and a yellow vase” is shown in Figure 1 as an example from the Concept Conjunction dataset. The left vision aligns with the text query. At the same time, the right image fails to provide a red book, the correct color for the vase, and an additional yellow flower. While the existing metrics (CLIP, NegCLIP, BLIP) predict similar scores for both images, failing to distinguish the correct image (on the left) from the incorrect one (on the right), human judges can make the correct and clear assessment (1.00 v.s. 0.45/0.55) of these two images on both overall and error counting objectives. 

Figure 1: Based on a text prompt taken from the Concept Conjunction dataset, two visuals were created using the Stable-Diffusion-2 algorithm. Baseline part displays the results of the currently used model-based assessment metrics. Human section displays the results of the human evaluation’s rating scores. The LLMScore-generated justification is likewise displayed in the right column.

Additionally, these measures offer a single, opaque score that hides the underlying logic behind how the synthesized pictures were aligned with the provided text prompts. Additionally, these model-based measures are rigid and cannot adhere to diverse standards prioritizing distinct text-to-image assessment objectives. For instance, the evaluation might access semantics at the level of an image (Overall) or more minute information at the level of an item (Error Counting). These problems prevent the current measurements from being in line with subjective assessments. In this study researchers from the University of California, the University of Washington and the University of California uncover the potent reasoning capabilities of large language models (LLMs), introducing LLMScore, a unique framework to evaluate text-image alignment in text-to-image synthesis. 

🚀 JOIN the fastest ML Subreddit Community

The human method of assessing text-image alignment, which entails verifying the accuracy of the items and characteristics mentioned in the text prompt, served as their model. LLMScore may mimic the human review by accessing compositionality at many granularities and producing alignment scores with justifications. This gives users a deeper understanding of the model’s performance and the motivations behind the results. Their LLMScore collects grounded Visio-linguistic information from vision and language models and LLMs, so capturing multi-granularity compositionality in the text and image to improve the evaluation of composite text-to-image synthesis. 

Their method uses language and vision models to convert a picture into multi-granularity (image- and object-level) visual descriptions, enabling us to express the compositional characteristics of numerous objects in language. When reasoning the alignment between text prompts and visuals, they combine these descriptions with text prompts and input them into large language models (LLMs), like GPT-4. Existing metrics struggle to capture compositionality, but their LLMScore does so by detecting the object-level alignment of text and picture (Figure 1). This results in scores that are well associated with human evaluation and have logical justifications (Figure 1). 

Additionally, by tailoring the evaluation instruction for LLMs, their LLMScore can adaptively follow different standards (overall or mistake counting). For instance, they may ask the LLMs to rate the overall alignment of the text prompt and the picture to assess the overall objective. Alternatively, they could ask them to confirm the error counting objective by asking, “How many compositional errors are in the image?” To maintain the determinism of the LLM’s conclusion, they also explicitly provide information on the different forms of text-to-image model errors in the assessment instruction. Because of its adaptability, their system may be used for various text-to-image jobs and assessment criteria. 

Modern text-to-image models like Stable Diffusion and DALLE are tested in their experimental setup using a variety of datasets, including prompt datasets for general use (MSCOCO, DrawBench, PaintSkills), as well as for compositional purposes (Abstract Concept Conjunction, Attribute Binding Contrast). They conducted numerous trials to confirm using LLMScore and show that it is aligned with human judgments without needing extra training. Across all datasets, their LLMS score had the strongest human correlation. On compositional datasets, they outperform the commonly used metrics CLIP and BLIP, respectively, by 58.8% and 31.27% Kendall’s. 

In conclusion, they provide LLMScore, the first effort to demonstrate the effectiveness of large language models for text-to-image assessment. Specifically, their article contributes the following: 

• They suggest the LLMScore. This brand-new framework provides scores that precisely express multi-granularity compositionality (image-level and object-level) for evaluating the alignment between text prompts and synthesized pictures in text-to-image synthesis. 

• Their LLMScore generates precise alignment scores with justifications following several evaluation directives (overall and mistake counting). 

• They use a variety of datasets (both compositional and general purpose) to verify the LLMScore. Among the widely utilized measures (CLIP, BLIP), their suggested LLMScore gets the strongest human correlation.


Check out the Paper and Github Link. Don’t forget to join our 22k+ ML SubReddit, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more. If you have any questions regarding the above article or if we missed anything, feel free to email us at [email protected]

🚀 Check Out 100’s AI Tools in AI Tools Club


YOU MAY ALSO LIKE

Anthropic Releases Claude Fable 5.1 and Claude Mythos 5.1: 52.6% on Terminal-Bench-Science and 75% Cheaper Cache Reads

Frontier models can recover up to 65% of facts they can’t directly recall — just by thinking longer

Aneesh Tickoo is a consulting intern at MarktechPost. He is currently pursuing his undergraduate degree in Data Science and Artificial Intelligence from the Indian Institute of Technology(IIT), Bhilai. He spends most of his time working on projects aimed at harnessing the power of machine learning. His research interest is image processing and is passionate about building solutions around it. He loves to connect with people and collaborate on interesting projects.


➡️ Ultimate Guide to Data Labeling in Machine Learning

Credit: Source link

ShareTweetSendSharePin

Related Posts

Anthropic Releases Claude Fable 5.1 and Claude Mythos 5.1: 52.6% on Terminal-Bench-Science and 75% Cheaper Cache Reads
AI & Technology

Anthropic Releases Claude Fable 5.1 and Claude Mythos 5.1: 52.6% on Terminal-Bench-Science and 75% Cheaper Cache Reads

September 1, 2026
Frontier models can recover up to 65% of facts they can’t directly recall — just by thinking longer
AI & Technology

Frontier models can recover up to 65% of facts they can’t directly recall — just by thinking longer

September 1, 2026
The New Street Fighter Movie Trailer Looks Fun In All The Right Ways
AI & Technology

The New Street Fighter Movie Trailer Looks Fun In All The Right Ways

September 1, 2026
Anthropic Announces Enterprise Frontier Safeguards, Customer-Held Data – Unite.AI
AI & Technology

Anthropic Announces Enterprise Frontier Safeguards, Customer-Held Data – Unite.AI

September 1, 2026
Next Post
EHang Holdings Limited (EH) Q1 2023 Earnings Call Transcript

EHang Holdings Limited (EH) Q1 2023 Earnings Call Transcript

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
How To Make Your Android Alarm Ring At Full Volume Even When Your Calls Are Muted

How To Make Your Android Alarm Ring At Full Volume Even When Your Calls Are Muted

August 28, 2026
Boat captain charged after vessel capsizes near Statue of Liberty, killing woman and her child

Boat captain charged after vessel capsizes near Statue of Liberty, killing woman and her child

August 26, 2026
Prudential plc (PUK) Q2 2026 Earnings Call Prepared Remarks Transcript

Prudential plc (PUK) Q2 2026 Earnings Call Prepared Remarks Transcript

August 27, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!