• bitcoinBitcoin(BTC)$77,254.000.03%
  • ethereumEthereum(ETH)$2,521.210.40%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$725.89-1.00%
  • rippleXRP(XRP)$1.370.28%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$101.790.33%
  • tronTRON(TRX)$0.339711-0.30%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.00-3.07%
  • zcashZcash(ZEC)$1,145.531.39%
  • HyperliquidHyperliquid(HYPE)$79.331.06%
  • dogecoinDogecoin(DOGE)$0.0847810.55%
  • RainRain(RAIN)$0.0157673.73%
  • moneroMonero(XMR)$531.34-0.91%
  • USDSUSDS(USDS)$1.00-0.01%
  • whitebitWhiteBIT Coin(WBT)$80.290.16%
  • chainlinkChainlink(LINK)$11.520.19%
  • leo-tokenLEO Token(LEO)$9.06-0.75%
  • cardanoCardano(ADA)$0.207552-0.01%
  • stellarStellar(XLM)$0.1805390.43%
  • Ethena USDeEthena USDe(USDE)$1.00-0.02%
  • daiDai(DAI)$1.000.01%
  • bitcoin-cashBitcoin Cash(BCH)$226.01-1.55%
  • USD1USD1(USD1)$1.00-0.01%
  • litecoinLitecoin(LTC)$54.200.84%
  • uniswapUniswap(UNI)$6.383.24%
  • CantonCanton(CC)$0.0987710.51%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.370.57%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • hedera-hashgraphHedera(HBAR)$0.0754021.50%
  • avalanche-2Avalanche(AVAX)$7.42-0.38%
  • shiba-inuShiba Inu(SHIB)$0.0000050.65%
  • nearNEAR Protocol(NEAR)$2.35-0.04%
  • suiSui(SUI)$0.720.35%
  • crypto-com-chainCronos(CRO)$0.0595423.67%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • MemeCoreMemeCore(M)$1.18-1.26%
  • tether-goldTether Gold(XAUT)$4,348.90-0.01%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • okbOKB(OKB)$114.320.12%
  • BittensorBittensor(TAO)$237.461.82%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.07%
  • aaveAave(AAVE)$127.251.68%
  • pax-goldPAX Gold(PAXG)$4,355.740.02%
  • AsterAster(ASTER)$0.701.62%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0576854.75%
  • mantleMantle(MNT)$0.55-4.83%
  • polkadotPolkadot(DOT)$1.01-3.18%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

JPMorgan AI Research Introduces DocLLM: A Lightweight Extension to Traditional Large Language Models Tailored for Generative Reasoning Over Documents with Rich Layouts

January 5, 2024
in AI & Technology
Reading Time: 4 mins read
A A
JPMorgan AI Research Introduces DocLLM: A Lightweight Extension to Traditional Large Language Models Tailored for Generative Reasoning Over Documents with Rich Layouts
ShareShareShareShareShare

Enterprise documents like contracts, reports, invoices, and receipts come with intricate layouts. These documents may be automatically interpreted and analyzed, which is useful and can result in the creation of AI-driven solutions. However, there are a number of challenges, as these documents can have rich semantics that lie at the intersection of textual and spatial modalities. The complex layouts of the documents provide crucial visual clues that are necessary for their efficient interpretation.

While Document AI (DocAI) has made significant strides in areas such as question answering, categorization, and extraction, real-world applications continue to face persistent hurdles related to accuracy, reliability, contextual understanding, and generalization to new domains.

To address these issues, a team of researchers from JPMorgan AI Research has introduced DocLLM, a lightweight version of conventional Large Language Models (LLMs) that takes into account both textual semantics and spatial layout and has been specifically created for reasoning over visual documents.

DocLLM is inherently multi-modal since it represents both text semantics and spatial layouts. In contrast to traditional methods, it has been developed in a way that it uses bounding box coordinates acquired using optical character recognition (OCR) to add spatial layout information, hence removing the requirement for a sophisticated visual encoder. This design decision decreases processing times, barely slightly increases model size, and maintains the causal decoder architecture.

The team has shared that for several document intelligence tasks, including form comprehension, table alignment, and visual question responding, just having a spatial layout structure is adequate. By separating spatial information from textual information, the method has extended typical transformers’ self-attention mechanism to capture cross-modal interactions.

Visual documents frequently have fragmented text sections, erratic layouts, and varied information. To address this, the study has suggested changing the pre-training target during the self-supervised pre-training phase. It has recommended infilling to accommodate various text arrangements and cohesive text blocks. With this adjustment, the model can more effectively handle mixed data types, complex layouts, contextual completions, and misaligned text.

DocLLM’s pre-trained knowledge has been fine-tuned on instruction data from many datasets to suit different document intelligence jobs. These tasks include document categorization, visual question answering, natural language inference, and key information extraction. 

Both single- and multi-page documents have been covered by the instruction-tuning data, and layout cues like field separators, titles, and captions can be included to make it easier for readers to understand the papers’ logical structure. For the Llama2-7B model, the changes made by DocLLM have yielded notable performance gains, ranging from 15% to 61%, in four of the five previously unpublished datasets.

The team has summarized their primary contributions as follows.

  1. A typical LLM with a lightweight extension designed especially for visual document interpretation has been introduced,
  1. The study aims to provide a unique attention mechanism that can distinguish between textual and spatial information, enabling the efficient capture of cross-modal alignment between layout and text.
  1. A pre-training goal has been outlined to address the difficulties caused by asymmetrical layouts in visual documents.
  1. A specialized instruction-tuning dataset has been designed for visual document intelligence tasks that should be curated to fine-tune the model effectively.
  1. In-depth trials have been performed, which yielded important insights into how the suggested model behaves and functions while managing visual documents.

Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to join our 35k+ ML SubReddit, 41k+ Facebook Community, Discord Channel, LinkedIn Group, Twitter, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.

If you like our work, you will love our newsletter..


YOU MAY ALSO LIKE

Implementation of Machine Learning Workflows with NVIDIA cuML, RAPIDS, GPU Benchmarking, Explainability, Clustering, and Model Inference

Hyundai Motor Group Puts Data Flywheel Into Full Operation – Unite.AI

Tanya Malhotra is a final year undergrad from the University of Petroleum & Energy Studies, Dehradun, pursuing BTech in Computer Science Engineering with a specialization in Artificial Intelligence and Machine Learning.
She is a Data Science enthusiast with good analytical and critical thinking, along with an ardent interest in acquiring new skills, leading groups, and managing work in an organized manner.


🐝 Get stunning professional headshots effortlessly with Aragon- TRY IT NOW!.


Credit: Source link

ShareTweetSendSharePin

Related Posts

Implementation of Machine Learning Workflows with NVIDIA cuML, RAPIDS, GPU Benchmarking, Explainability, Clustering, and Model Inference
AI & Technology

Implementation of Machine Learning Workflows with NVIDIA cuML, RAPIDS, GPU Benchmarking, Explainability, Clustering, and Model Inference

September 13, 2026
Hyundai Motor Group Puts Data Flywheel Into Full Operation – Unite.AI
AI & Technology

Hyundai Motor Group Puts Data Flywheel Into Full Operation – Unite.AI

September 13, 2026
What Is The Difference Between A Dead Pixel And A Stuck Pixel?
AI & Technology

What Is The Difference Between A Dead Pixel And A Stuck Pixel?

September 13, 2026
Cognition Releases SWE-2: A Kimi K3 Post-Trained Coding Model That Matches Fable 5.1 on FrontierCode at 64% Lower Cost
AI & Technology

Cognition Releases SWE-2: A Kimi K3 Post-Trained Coding Model That Matches Fable 5.1 on FrontierCode at 64% Lower Cost

September 12, 2026
Next Post
Flashback 2008: ‘Joe the Plumber’ poses tax policy question to Obama

Flashback 2008: 'Joe the Plumber’ poses tax policy question to Obama

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Hyundai Motor Group Puts Data Flywheel Into Full Operation – Unite.AI

Hyundai Motor Group Puts Data Flywheel Into Full Operation – Unite.AI

September 13, 2026
Bunching Deductions Into One Year Can Beat the Standard Deduction

Bunching Deductions Into One Year Can Beat the Standard Deduction

September 12, 2026
TikTok rejects Meta ads urging firm to join landmark child safety settlement: report

TikTok rejects Meta ads urging firm to join landmark child safety settlement: report

September 11, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!