• bitcoinBitcoin(BTC)$78,291.001.56%
  • ethereumEthereum(ETH)$2,499.820.35%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$720.400.18%
  • rippleXRP(XRP)$1.403.59%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$101.350.91%
  • tronTRON(TRX)$0.340782-0.06%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.040.00%
  • zcashZcash(ZEC)$1,144.294.17%
  • HyperliquidHyperliquid(HYPE)$80.072.19%
  • dogecoinDogecoin(DOGE)$0.0837810.18%
  • RainRain(RAIN)$0.014956-2.38%
  • USDSUSDS(USDS)$1.000.00%
  • moneroMonero(XMR)$510.99-3.97%
  • whitebitWhiteBIT Coin(WBT)$80.851.18%
  • chainlinkChainlink(LINK)$11.380.41%
  • leo-tokenLEO Token(LEO)$8.98-0.78%
  • cardanoCardano(ADA)$0.2090190.71%
  • stellarStellar(XLM)$0.1892505.35%
  • Ethena USDeEthena USDe(USDE)$1.000.01%
  • daiDai(DAI)$1.000.00%
  • bitcoin-cashBitcoin Cash(BCH)$222.79-1.35%
  • USD1USD1(USD1)$1.00-0.01%
  • litecoinLitecoin(LTC)$53.70-1.14%
  • uniswapUniswap(UNI)$6.33-0.03%
  • CantonCanton(CC)$0.0955280.03%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.350.09%
  • hedera-hashgraphHedera(HBAR)$0.0765790.06%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • avalanche-2Avalanche(AVAX)$7.450.71%
  • nearNEAR Protocol(NEAR)$2.403.61%
  • shiba-inuShiba Inu(SHIB)$0.0000050.34%
  • suiSui(SUI)$0.720.57%
  • crypto-com-chainCronos(CRO)$0.0592051.17%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,268.37-1.86%
  • BittensorBittensor(TAO)$232.14-1.97%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • MemeCoreMemeCore(M)$1.10-4.65%
  • okbOKB(OKB)$113.670.49%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.13%
  • aaveAave(AAVE)$125.78-1.24%
  • AsterAster(ASTER)$0.69-1.23%
  • mantleMantle(MNT)$0.57-0.27%
  • BitwayBitway(BTW)$0.69-1.47%
  • pax-goldPAX Gold(PAXG)$4,274.31-1.79%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.057152-0.64%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Researchers from Alibaba and the Renmin University of China Present mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

March 23, 2024
in AI & Technology
Reading Time: 6 mins read
A A
Researchers from Alibaba and the Renmin University of China Present mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding
ShareShareShareShareShare

Harnessing the strong language understanding and generation potential of Large Language Models (LLMs), Multimodal Large Language Models (MLLMs) have been developed in recent years for vision-and-language understanding tasks. MLLMs have shown promising results in understanding general images by aligning a pre-trained visual encoder (e.g., the Vision Transformers) and the LLM with a Vision-toText (V2T) module. However, these models still need to improve understanding and extracting text from images containing rich text information, like documents, webpages, tables, and charts. The main reason is that the visual encoder and V2T module are trained on general image-text pairs and must be specifically optimized to represent the textual and structural information in text-rich images.

To enhance visual document understanding with MLLMs, prior works like mPLUG-DocOwl, Docpedia, and Ureader attempted to design text-reading tasks to strengthen text recognition ability. Still, they must pay more attention to structure comprehension or cover limited domains of text-rich images, such as web pages or documents.

Researchers from Alibaba Group and the Renmin University of China have introduced DocOwl 1.5, a Unified Structure Learning, to boost the performance of MLLMs. Unified Structure Learning comprises structure-aware parsing tasks and multi-grained text localization tasks across five domains: document, webpage, table, chart, and natural image. To better encode structure information, they have designed a simple and effective vision-to-text module H-Reducer, which can not only maintain the layout information but also reduce the length of visual features by merging horizontal adjacent patches through convolution, enabling the LLM to understand high-resolution images more efficiently.

DocOwl 1.5 follows the typical architecture of MLLMs: a visual encoder, a vision-to-text module, and a LLM as the decoder.

High-resolution Image Encoding: High-resolution images ensure the decoder can use rich text information from document images. They utilize a parameter-free

Shape-adaptive Cropping Module to crop a shape-variable high-resolution image into multiple

fixed-size sub-images.

Spatial-aware Vision-to-Text Module: H-Reducer: They have designed a comparatively more appropriate vision-to-text module for Visual Document Understanding, namely H-Reducer, which reduces visual sequence length and keeps spatial information. The H-Reducer comprises a convolution layer to reduce sequence length and a fully connected layer to project visual features to language embedding space. Since most textual information in document images is arranged from left to right, the horizontal text information is usually semantically coherent. Thus, the kernel and stride sizes in the convolution layer are set as 1×4 to ensemble horizontal four visual features. The output channel is set equal to the input channel.

Multimodal Modeling with LLM as the decoder: They applied the Modeality-Adaptive Module (MAM) in LLM to better distinguish visual and textual inputs. During self-attention, MAM utilizes two sets of linear projection layers to perform the key/value projection for visual features and textual features separately. 

Evaluated DocOwl 1.5 on the test across ten challenges, from documents and tables to charts and webpage screenshots. Compared to other models, even those with massive parameter counts, DocOwl 1.5 surpasses them all. DocOwl 1.5 also outperforms CogAgent on InfoVQA and ChartQA and achieves comparable performance on DocVQA. This suggests that unified structure learning with DocStruct4M is more efficient in learning printed text recognition and how to analyze documents. 

Major components that helped researchers get state-of-the-art performance:

Effectiveness of H-Reducer:  Shape-adaptive Cropping Module, the image resolution supported by the MLLM is the product of each crop’s cropping number and basic resolution. With the Abstractor as the vision-to-text module, reducing the cropping number causes an obvious performance decrease (r4 vs r3) on documents. However, with a smaller cropping number, the H-Reducer performs better than the Abstractor (r5 vs. r3), and the H-Reducer is stronger in maintaining rich text information during vision-and-language feature alignment.

Effectiveness of Unified Structure Learning. With only the structure-aware parsing tasks, there is significant improvement across different domains (r10 vs r5). This validates that finetuning the visual encoder and H-Reducer with structure-aware parsing tasks greatly helps MLLMs understand text-rich images.

Effectiveness of the Two-stage Training: For joint training, the model improves significantly on DocVQA as the samples of Unified Structure Learning increase when it is below 1M. However, as the Unified Structure Learning samples are further increased, the improvement of the model becomes subtle, and its performance could be better than that of the one using two-stage training. This shows that the two-stage training could enhance basic text recognition and structure parsing abilities and is more beneficial and efficient for downstream document understanding.

In conclusion, researchers from Alibaba Group and the Renmin University of China have proposed DocOwl 1.5, a Unified Structure Learning across five domains of text-rich images, including both structure-aware parsing tasks and multi-grained text localization tasks. To better maintain structure and spatial information during vision-and-language feature alignment, they designed a simple and effective vision-to-text module named H-Reducer. It mainly utilizes a convolution layer to aggregate horizontally neighboring visual features. DocOwl 1.5 achieves state-of-the-art OCR-free performance on ten visual document understanding benchmarks.


Check out the Paper and Github. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. Join our Telegram Channel, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 39k+ ML SubReddit


YOU MAY ALSO LIKE

Temporal Raises $550M Series E at $12.55B Valuation to Expand Operations – Unite.AI

What Is MSI Mode On Windows PCs And Does It Speed Up Your GPU?

Asjad is an intern consultant at Marktechpost. He is persuing B.Tech in mechanical engineering at the Indian Institute of Technology, Kharagpur. Asjad is a Machine learning and deep learning enthusiast who is always researching the applications of machine learning in healthcare.


🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

Temporal Raises 0M Series E at .55B Valuation to Expand Operations – Unite.AI
AI & Technology

Temporal Raises $550M Series E at $12.55B Valuation to Expand Operations – Unite.AI

September 14, 2026
What Is MSI Mode On Windows PCs And Does It Speed Up Your GPU?
AI & Technology

What Is MSI Mode On Windows PCs And Does It Speed Up Your GPU?

September 14, 2026
How To Block Time-Wasting Apps On iPhone Using Screen Time
AI & Technology

How To Block Time-Wasting Apps On iPhone Using Screen Time

September 14, 2026
What Is Agentic RAG? When AI Plans Its Own Search and Retrieval – Unite.AI
AI & Technology

What Is Agentic RAG? When AI Plans Its Own Search and Retrieval – Unite.AI

September 14, 2026
Next Post
The metrics you can’t afford to ignore: What the best CEOs know

The metrics you can’t afford to ignore: What the best CEOs know

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Your Mortgage Rate Affects Your Total Cost More Than the Purchase Price

Your Mortgage Rate Affects Your Total Cost More Than the Purchase Price

September 9, 2026
Everything You Need to Know About Apple’s iPhone Duo

Everything You Need to Know About Apple’s iPhone Duo

September 12, 2026
BrainChip Holdings Ltd (BCHPY) Q2 2026 Earnings Call Transcript

BrainChip Holdings Ltd (BCHPY) Q2 2026 Earnings Call Transcript

September 11, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!