• bitcoinBitcoin(BTC)$76,815.00-0.68%
  • ethereumEthereum(ETH)$2,493.84-1.57%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$717.66-2.25%
  • rippleXRP(XRP)$1.35-1.45%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$99.92-1.83%
  • tronTRON(TRX)$0.3395260.02%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.00-1.59%
  • zcashZcash(ZEC)$1,108.66-3.61%
  • HyperliquidHyperliquid(HYPE)$77.84-1.93%
  • dogecoinDogecoin(DOGE)$0.083411-1.58%
  • RainRain(RAIN)$0.0154161.77%
  • moneroMonero(XMR)$538.85-0.67%
  • USDSUSDS(USDS)$1.00-0.01%
  • whitebitWhiteBIT Coin(WBT)$79.77-0.78%
  • chainlinkChainlink(LINK)$11.31-2.04%
  • leo-tokenLEO Token(LEO)$9.06-0.65%
  • cardanoCardano(ADA)$0.203862-2.39%
  • stellarStellar(XLM)$0.177785-1.79%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • daiDai(DAI)$1.000.02%
  • bitcoin-cashBitcoin Cash(BCH)$222.35-4.00%
  • USD1USD1(USD1)$1.00-0.02%
  • litecoinLitecoin(LTC)$53.69-0.77%
  • uniswapUniswap(UNI)$6.19-2.60%
  • CantonCanton(CC)$0.095246-3.64%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.35-1.99%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • hedera-hashgraphHedera(HBAR)$0.0746970.29%
  • avalanche-2Avalanche(AVAX)$7.30-2.00%
  • shiba-inuShiba Inu(SHIB)$0.000005-1.20%
  • nearNEAR Protocol(NEAR)$2.30-3.14%
  • suiSui(SUI)$0.71-2.61%
  • crypto-com-chainCronos(CRO)$0.0581750.94%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,347.37-0.06%
  • MemeCoreMemeCore(M)$1.15-2.65%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • okbOKB(OKB)$113.82-0.26%
  • BittensorBittensor(TAO)$232.60-1.33%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.18%
  • aaveAave(AAVE)$124.40-1.67%
  • pax-goldPAX Gold(PAXG)$4,354.57-0.02%
  • AsterAster(ASTER)$0.69-0.02%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0574331.09%
  • mantleMantle(MNT)$0.55-4.09%
  • polkadotPolkadot(DOT)$1.00-4.32%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Google AI Introduces ScreenAI: A Vision-Language Model for User interfaces (UI) and Infographics Understanding

February 21, 2024
in AI & Technology
Reading Time: 4 mins read
A A
Google AI Introduces ScreenAI: A Vision-Language Model for User interfaces (UI) and Infographics Understanding
ShareShareShareShareShare

The capacity of infographics to strategically arrange and use visual signals to clarify complicated concepts has made them essential for efficient communication. Infographics include various visual elements such as charts, diagrams, illustrations, maps, tables, and document layouts. This has been a long-standing technique that makes the material easier to understand. User interfaces (UIs) on desktop and mobile platforms share design concepts and visual languages with infographics in the modern digital world. 

Though there is a lot of overlap between UIs and infographics, creating a cohesive model is made more difficult by the complexity of each. It is difficult to develop a single model that can efficiently analyze and interpret the visual information encoded in pixels because of the intricacy required in understanding, reasoning, and engaging with the various aspects of infographics and user interfaces.

To address this, in a recent Google Research, a team of researchers proposed ScreenAI as a solution. ScreenAI is a Vision-Language Model (VLM) that has the ability to comprehend both UIs and infographics fully. Tasks like graphical question-answering (QA), which may contain charts, pictures, maps, and more, have been included in its scope.

The team has shared that ScreenAI can manage jobs like element annotation, summarization, navigation, and additional UI-specific QA. To accomplish this, the model combines the flexible patching method taken from Pix2struct with the PaLI architecture, which allows it to tackle vision-related tasks by converting them into text or image-to-text problems.

Several tests have been carried out to demonstrate how these design decisions affect the model’s functionality. Upon evaluation, ScreenAI produced new state-of-the-art results on tasks like Multipage DocVQA, WebSRC, MoTIF, and Widget Captioning with under 5 billion parameters. It achieved remarkable performance on tasks including DocVQA, InfographicVQA, and Chart QA, outperforming models of comparable size. 

The team has made available three additional datasets: Screen Annotation, ScreenQA Short, and Complex ScreenQA. One of these datasets specifically focuses on the screen annotation task for future research, while the other two datasets are focused on question-answering, thus further expanding the resources available to advance the field. 

The team has summarized their primary contributions as follows:

  1. The Vision-Language Model (VLM) ScreenAI concept is a step towards a holistic solution that focuses on infographic and user interface comprehension. By utilizing the common visual language and sophisticated design of these components, ScreenAI offers a comprehensive method for understanding digital material.
  1. One significant advancement is the development of a textual representation for UIs. During the pretraining stage, this representation has been used to teach the model how to comprehend user interfaces, improving its capacity to comprehend and process visual data.
  1. To automatically create training data at scale, ScreenAI has used LLMs and the new UI representation, making training more effective and comprehensive.
  1. Three new datasets, Screen Annotation, ScreenQA Short, and Complex ScreenQA, have been released. These datasets allow for thorough model benchmarking for screen-based question answering and the suggested textual representation.
  1. ScreenAI has outperformed larger models by a factor of ten or more on four public infographics QA benchmarks, even with its low number of 4.6 billion parameters. 

Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter and Google News. Join our 37k+ ML SubReddit, 41k+ Facebook Community, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our Telegram Channel


YOU MAY ALSO LIKE

Implementation of Machine Learning Workflows with NVIDIA cuML, RAPIDS, GPU Benchmarking, Explainability, Clustering, and Model Inference

Hyundai Motor Group Puts Data Flywheel Into Full Operation – Unite.AI

Tanya Malhotra is a final year undergrad from the University of Petroleum & Energy Studies, Dehradun, pursuing BTech in Computer Science Engineering with a specialization in Artificial Intelligence and Machine Learning.
She is a Data Science enthusiast with good analytical and critical thinking, along with an ardent interest in acquiring new skills, leading groups, and managing work in an organized manner.


🚀 LLMWare Launches SLIMs: Small Specialized Function-Calling Models for Multi-Step Automation [Check out all the models]


Credit: Source link

ShareTweetSendSharePin

Related Posts

Implementation of Machine Learning Workflows with NVIDIA cuML, RAPIDS, GPU Benchmarking, Explainability, Clustering, and Model Inference
AI & Technology

Implementation of Machine Learning Workflows with NVIDIA cuML, RAPIDS, GPU Benchmarking, Explainability, Clustering, and Model Inference

September 13, 2026
Hyundai Motor Group Puts Data Flywheel Into Full Operation – Unite.AI
AI & Technology

Hyundai Motor Group Puts Data Flywheel Into Full Operation – Unite.AI

September 13, 2026
What Is The Difference Between A Dead Pixel And A Stuck Pixel?
AI & Technology

What Is The Difference Between A Dead Pixel And A Stuck Pixel?

September 13, 2026
Cognition Releases SWE-2: A Kimi K3 Post-Trained Coding Model That Matches Fable 5.1 on FrontierCode at 64% Lower Cost
AI & Technology

Cognition Releases SWE-2: A Kimi K3 Post-Trained Coding Model That Matches Fable 5.1 on FrontierCode at 64% Lower Cost

September 12, 2026
Next Post
Minnesota man charged in crash that killed 5 young Muslim women

Minnesota man charged in crash that killed 5 young Muslim women

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Should You Buy Your Leased Car – Run These Two Numbers First

Should You Buy Your Leased Car – Run These Two Numbers First

September 11, 2026
NVIDIA Details BioNeMo Inference Runtime (BioIR): 2.90x Higher Boltz-2 Folding Throughput and 58.5K Residues per GPU-Hour on 8xH100

NVIDIA Details BioNeMo Inference Runtime (BioIR): 2.90x Higher Boltz-2 Folding Throughput and 58.5K Residues per GPU-Hour on 8xH100

September 10, 2026
Roblox Will Soon Add Offline And Browser-Based Play Modes

Roblox Will Soon Add Offline And Browser-Based Play Modes

September 11, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!