• bitcoinBitcoin(BTC)$77,099.000.23%
  • ethereumEthereum(ETH)$2,516.532.70%
  • tetherTether(USDT)$1.000.01%
  • binancecoinBNB(BNB)$724.121.63%
  • rippleXRP(XRP)$1.350.24%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$101.932.41%
  • tronTRON(TRX)$0.338180-0.67%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.03-0.29%
  • zcashZcash(ZEC)$1,160.513.63%
  • HyperliquidHyperliquid(HYPE)$79.10-1.00%
  • dogecoinDogecoin(DOGE)$0.0838760.07%
  • RainRain(RAIN)$0.015464-2.16%
  • USDSUSDS(USDS)$1.000.01%
  • moneroMonero(XMR)$515.960.86%
  • whitebitWhiteBIT Coin(WBT)$80.090.63%
  • chainlinkChainlink(LINK)$11.51-0.43%
  • leo-tokenLEO Token(LEO)$9.15-0.40%
  • cardanoCardano(ADA)$0.204775-1.35%
  • stellarStellar(XLM)$0.1776650.71%
  • Ethena USDeEthena USDe(USDE)$1.000.03%
  • daiDai(DAI)$1.000.01%
  • bitcoin-cashBitcoin Cash(BCH)$226.570.58%
  • USD1USD1(USD1)$1.000.04%
  • litecoinLitecoin(LTC)$52.981.03%
  • CantonCanton(CC)$0.096810-1.77%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.360.32%
  • uniswapUniswap(UNI)$5.99-0.50%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • avalanche-2Avalanche(AVAX)$7.42-1.67%
  • hedera-hashgraphHedera(HBAR)$0.074091-1.73%
  • nearNEAR Protocol(NEAR)$2.41-3.14%
  • shiba-inuShiba Inu(SHIB)$0.0000050.81%
  • suiSui(SUI)$0.72-1.72%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.0560850.00%
  • MemeCoreMemeCore(M)$1.181.78%
  • tether-goldTether Gold(XAUT)$4,346.800.63%
  • Circle USYCCircle USYC(USYC)$1.140.03%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • okbOKB(OKB)$112.831.82%
  • BittensorBittensor(TAO)$234.32-2.09%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.02%
  • aaveAave(AAVE)$124.031.50%
  • mantleMantle(MNT)$0.581.25%
  • pax-goldPAX Gold(PAXG)$4,352.510.73%
  • AsterAster(ASTER)$0.68-3.67%
  • polkadotPolkadot(DOT)$1.03-7.39%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.054492-3.41%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Meet Q-Align: The All-in-One Visual Scorer Based on Large Multi-Modality Models

January 7, 2024
in AI & Technology
Reading Time: 5 mins read
A A
Meet Q-Align: The All-in-One Visual Scorer Based on Large Multi-Modality Models
ShareShareShareShareShare

With the vast amount of visual content available online, it is essential to assess images and videos accurately. The challenge is to develop robust machine assessment tools that can determine various types of visual content and align closely with human opinions. This need spans different domains, such as image and video quality assessment (IQA and VQA) and image aesthetic assessment (IAA), each requiring unique approaches to effectively rate and understand visual content.

Traditional methods, ranging from handcrafted algorithms to advanced deep-learning models, have focused on assessing visual content by regressing from mean opinion scores (MOS). However, these methods must be revised, particularly when dealing with new content types and diverse scoring scenarios. Their inadequacy largely stems from poor out-of-distribution generalization abilities, an issue that becomes increasingly prominent with the complexity and variety of modern visual content.

A breakthrough in this field is the introduction of Q-ALIGN, a novel methodology developed by researchers from Nanyang Technological University, Shanghai Jiao Tong University, and SenseTime Research. Q-ALIGN represents a departure from conventional approaches and educates Large Multi-Modality Models (LMMs) to rate visual content using text-defined rating levels, not direct numerical scores. This approach is more akin to how human raters evaluate and judge in subjective studies, marking a significant shift in machine-based visual assessment.

The methodology of Q-ALIGN is intricate and carefully designed. It converts existing score labels into discrete text-defined rating levels during the training phase. This process is analogous to how human raters learn and judge in subjective studies. They typically work with predefined levels like ‘excellent,’ ‘good,’ ‘fair,’ etc., rather than specific numerical scores. The innovation here is teaching LMMs to understand and use these text-defined levels for visual rating, which aligns more with human cognitive processes.

https://arxiv.org/abs/2312.17090

In the inference phase, Q-ALIGN emulates the strategy of collecting MOS from human ratings. It extracts the log probabilities on different rating levels and employs softmax pooling to obtain the close-set probabilities of each level. The final score is then derived from a weighted average of these probabilities. This process mirrors how human ratings are converted into MOS in subjective visual assessments.

The performance and results of Q-ALIGN are noteworthy. It has achieved state-of-the-art performance in IQA, IAA, and VQA tasks. Compared to existing methods that struggle with novel content types and diverse scoring scenarios, Q-ALIGN’s discrete-level-based syllabus has shown superior performance, especially in out-of-distribution settings. These results indicate its effectiveness in accurately assessing a wide range of visual content.

Q-ALIGN’s ability to generalize effectively to new types of content underlines its potential for broad application across various fields. It represents a paradigm shift in the domain of visual content assessment. Adopting a methodology that aligns more closely with human judgment offers a robust, accurate, and more intuitive tool for scoring diverse types of visual content. The work addresses the limitations of existing methods and opens up new possibilities for future advancements in the field.


Check out the Paper and Github. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. Join our 35k+ ML SubReddit, 41k+ Facebook Community, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..


YOU MAY ALSO LIKE

Anthropic Adds Plugin Evals to Claude Code: 6 Grader Types, a No-Plugin Baseline, and a CI Gate for Skills

Lenovo’s Googlebook 15 Seems Decidedly Premium Based On A New Leak

Muhammad Athar Ganaie, a consulting intern at MarktechPost, is a proponet of Efficient Deep Learning, with a focus on Sparse Training. Pursuing an M.Sc. in Electrical Engineering, specializing in Software Engineering, he blends advanced technical knowledge with practical applications. His current endeavor is his thesis on “Improving Efficiency in Deep Reinforcement Learning,” showcasing his commitment to enhancing AI’s capabilities. Athar’s work stands at the intersection “Sparse Training in DNN’s” and “Deep Reinforcemnt Learning”.


⬆️ Join Our 35k+ ML SubReddit


Credit: Source link

ShareTweetSendSharePin

Related Posts

Anthropic Adds Plugin Evals to Claude Code: 6 Grader Types, a No-Plugin Baseline, and a CI Gate for Skills
AI & Technology

Anthropic Adds Plugin Evals to Claude Code: 6 Grader Types, a No-Plugin Baseline, and a CI Gate for Skills

September 11, 2026
Lenovo’s Googlebook 15 Seems Decidedly Premium Based On A New Leak
AI & Technology

Lenovo’s Googlebook 15 Seems Decidedly Premium Based On A New Leak

September 11, 2026
New Images Show A Detailed View Of Meta’s Upcoming Mixed Reality Headset
AI & Technology

New Images Show A Detailed View Of Meta’s Upcoming Mixed Reality Headset

September 11, 2026
Dzmitry Lazerka, Co-Founder of VictoriaMetrics – Interview Series – Unite.AI
AI & Technology

Dzmitry Lazerka, Co-Founder of VictoriaMetrics – Interview Series – Unite.AI

September 11, 2026
Next Post
Salesforce Research Proposes MoonShot: A New Video Generation AI Model that Conditions Simultaneously on Multimodal Inputs of Image and Text

Salesforce Research Proposes MoonShot: A New Video Generation AI Model that Conditions Simultaneously on Multimodal Inputs of Image and Text

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
How These XL Phones Compete

How These XL Phones Compete

September 10, 2026
Paramount Skydance agrees to halt Warner Bros. merger

Paramount Skydance agrees to halt Warner Bros. merger

September 5, 2026
Falcons name Tagovailoa as starter vs. Steelers – NBC Sports

Falcons name Tagovailoa as starter vs. Steelers – NBC Sports

September 7, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!