• bitcoinBitcoin(BTC)$84,955.000.29%
  • ethereumEthereum(ETH)$2,684.96-0.51%
  • tetherTether(USDT)$1.000.02%
  • binancecoinBNB(BNB)$773.870.34%
  • rippleXRP(XRP)$1.500.16%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$119.901.35%
  • tronTRON(TRX)$0.3363330.46%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.041.54%
  • zcashZcash(ZEC)$1,359.630.09%
  • HyperliquidHyperliquid(HYPE)$88.941.13%
  • dogecoinDogecoin(DOGE)$0.0948940.09%
  • chainlinkChainlink(LINK)$14.16-1.05%
  • moneroMonero(XMR)$540.65-1.72%
  • whitebitWhiteBIT Coin(WBT)$84.590.24%
  • USDSUSDS(USDS)$1.00-0.02%
  • cardanoCardano(ADA)$0.2516300.78%
  • leo-tokenLEO Token(LEO)$8.971.10%
  • RainRain(RAIN)$0.011382-9.15%
  • stellarStellar(XLM)$0.2207750.34%
  • nearNEAR Protocol(NEAR)$4.86-0.16%
  • bitcoin-cashBitcoin Cash(BCH)$310.080.80%
  • uniswapUniswap(UNI)$9.00-1.02%
  • litecoinLitecoin(LTC)$70.444.40%
  • Ethena USDeEthena USDe(USDE)$1.00-0.02%
  • avalanche-2Avalanche(AVAX)$11.041.13%
  • CantonCanton(CC)$0.1213890.12%
  • suiSui(SUI)$1.17-0.48%
  • Blockchain USDBlockchain USD(USDB)$0.871,000.00%
  • daiDai(DAI)$1.00-0.01%
  • hedera-hashgraphHedera(HBAR)$0.102984-0.82%
  • USD1USD1(USD1)$1.000.00%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.52-0.76%
  • BitwayBitway(BTW)$1.451.48%
  • quant-networkQuant(QNT)$245.39-7.61%
  • BittensorBittensor(TAO)$308.671.00%
  • shiba-inuShiba Inu(SHIB)$0.0000060.84%
  • crypto-com-chainCronos(CRO)$0.068539-1.14%
  • tether-goldTether Gold(XAUT)$4,140.39-0.83%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.02%
  • aaveAave(AAVE)$182.578.45%
  • Pump.funPump.fun(PUMP)$0.0058754.28%
  • okbOKB(OKB)$121.24-0.05%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • EthenaEthena(ENA)$0.243251-2.78%
  • OndoOndo(ONDO)$0.501.81%
  • Circle USYCCircle USYC(USYC)$1.140.03%
  • MemeCoreMemeCore(M)$1.053.71%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.00%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Datalab Introduces OmniExtractBench to Fix Bias and Opacity in Extraction Benchmarks

October 2, 2026
in AI & Technology
Reading Time: 16 mins read
A A
Datalab Introduces OmniExtractBench to Fix Bias and Opacity in Extraction Benchmarks
ShareShareShareShareShare

Datalab has released OmniExtractBench, an open benchmark for structured document extraction. It tests how accurately a system fills a JSON schema from a PDF. The benchmark pools 620 documents from 4 existing benchmarks. One deterministic scorer grades all of them and explains each decision.

The release lands while extraction vendors publish their own leaderboards. Datalab argues those leaderboards are hard to compare or audit. OmniExtractBench is its attempt at a shared yardstick.

Is it deployable? Yes, the scorer installs from PyPI as omni-extract-bench (v0.1.7, Python 3.11+, SciPy only) under Apache 2.0. Rerunning vendors requires your own API keys and paid credits.

OmniExtractBench is a structured extraction benchmark built by Datalab. Each task gives a system a PDF and a JSON schema. The system returns JSON, which is scored value by value against a gold file. The code is on GitHub, and the data is on Hugging Face under CC BY 4.0.

The 4 flaws it targets

Datalab’s launch post names 4 recurring problems with existing extraction benchmarks:

  • Bias: documents and scoring can favor the vendor that built the benchmark.
  • Opaque harnesses: a low score may reflect a broken harness, not a weak model.
  • Unclear scoring: readers cannot tell why a given document scored low.
  • Narrow document variety: some suites hold only dense tables, others only clean, unscanned files.

Where the 620 documents come from

Suite Documents Upstream publisher Content
ExtractBench 329 LlamaIndex Forms, filings, decks
Internal 202 Datalab (synthetic) Dense scalar schemas, small documents
LongExtractBench 47 micro1 (commissioned by Reducto) Very large tables
LongArray-Extract 42 Extend Large tables with repeated scalars

Regulatory filing forms are the largest category, at 88 documents. 128 documents are a single page. At the other end, 33 documents over 100 pages hold 40% of all pages. Datalab’s own synthetic suite is the second largest share.

How the scorer works

The scorer flattens prediction and gold JSON into addresses, which are paths to single values. It normalizes each value first, so “03/31/2024” matches “2024-03-31”.

Tables are the hard part. Compared by position, one missed row shifts every row after it. We reran the scorer on a 100-row table missing its first row. Positional comparison scored 0%, while OmniExtractBench scored 99%.

The fix is content-based pairing with the Hungarian algorithm. ExtractBench and LongArray-Extract already align rows this way. OmniExtractBench adds a verdict layer on top.

6 verdicts per value

  • matched: paired, and the values agree.
  • misread: paired, but the values differ.
  • unfound: gold has a value, the prediction does not.
  • fabricated: the schema allows it, gold is silent, the prediction fills it.
  • invented_item: part of a predicted row that pairs with nothing.
  • invented_field: an address the schema never declared.

Accuracy is matched values over all verdicts. Precision divides matched values by predicted values. Recall divides them by gold values. The full rules are in the metric spec.

The null rule

Empty strings, None and whitespace count as omissions, so those addresses are dropped. Strings like “NA” or “-” remain real answers. This blocks a quiet exploit: padding a schema with empty optional fields to earn free matches. In our test, padded null fields added 0 verdicts.

Benchmark Publisher Documents Row alignment Per-value explanation Scorer license Data license
OmniExtractBench Datalab 620, from 4 sources Hungarian, by content Yes, 6 verdict types Apache 2.0 CC BY 4.0
ExtractBench LlamaIndex 370 Hungarian Per-field diffs, HTML report Apache 2.0 Apache 2.0
LongArray-Extract Extend 45, synthetic Hungarian Per-document score Not stated CC BY 4.0
LongExtractBench micro1 225, of which 50 public By row key Not stated MIT CC BY 4.0, labels only

Sources linked on each benchmark name. Checked September 27, 2026.

Who misses fields and who invents values

Datalab scored 10 system configurations on the full corpus. Its accurate mode led at 93.85 accuracy. Datalab balanced (93.48) and Reducto deep_extract v2 (93.47) are effectively tied. Precision and recall then show how each system fails.

  • Balanced: Datalab (both modes) and Reducto keep precision and recall within 0.6 points.
  • Leans to misses: GPT 5.6-sol posts 95.11 precision but 84.99 recall. It loses 11.88% to unfound values. Gemini and Claude show the same pattern, less sharply.
  • Leans to invented values: LlamaExtract has 93.13 recall but 86.57 precision, losing 9.03% to fabricated values. Extend loses 4.01% to invented items.
  • Low on both: Mistral OCR 4.1 and Azure Content Understanding trail on both metrics, with recall lower still.

How each system falls short of 100%, by verdict type. Source: Datalab.

Run it yourself

uv pip install omni-extract-bench
oeb score --pred pred.json --gt gold.json --schema schema.json --verdicts

Install the [benchmark] extra and run oeb benchmark to rerun vendors. Runs are resumable, and each provider needs its own credentials. Datalab also suggests testing its playground on your own documents.

Key Takeaways

  • OmniExtractBench pools 620 documents from LlamaIndex, micro1, Extend and Datalab suites.
  • A deterministic scorer gives every value 1 of 6 auditable verdicts.
  • Content-based row pairing stops one missed row from zeroing a table.
  • Dropping null addresses stops schema padding from inflating scores.
  • Datalab accurate leads at 93.85; Datalab balanced and Reducto tie near 93.5.

FAQ

  • What does OmniExtractBench measure? It measures how accurately a system extracts values from a PDF into a JSON schema, scored per value.
  • Is OmniExtractBench open source? Yes. The scorer is Apache 2.0 on GitHub and PyPI. The dataset is CC BY 4.0 on Hugging Face.
  • Who built OmniExtractBench? Datalab built it. Datalab also maintains the open-source Marker and Surya document tools.

Thanks to the Datalab team for the resources behind this article. Datalab supported and sponsored this content.


Asif Razzaq is the CEO of Marktechpost AI Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.

Credit: Source link

ShareTweetSendSharePin

Related Posts

Soon-To-Be Merged Paramount And Warner Bros. Will Be Known As Skydance Going Forward
AI & Technology

Soon-To-Be Merged Paramount And Warner Bros. Will Be Known As Skydance Going Forward

October 2, 2026
NVIDIA’s Smuggling Problem Is Getting Worse
AI & Technology

NVIDIA’s Smuggling Problem Is Getting Worse

October 2, 2026
Four Principles for Rebuilding Identity Verification – Unite.AI
AI & Technology

Four Principles for Rebuilding Identity Verification – Unite.AI

October 2, 2026
AKASA Brings Autonomous AI to Inpatient Coding and Clinical Documentation – Unite.AI
AI & Technology

AKASA Brings Autonomous AI to Inpatient Coding and Clinical Documentation – Unite.AI

October 2, 2026
Next Post
New video of hero passengers from plane attack

New video of hero passengers from plane attack

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Cleveland Browns defeat Pittsburgh Steelers in a Thursday night thriller – WTAE

Cleveland Browns defeat Pittsburgh Steelers in a Thursday night thriller – WTAE

October 2, 2026
Growing desperation in Indiana as thousands still without power

Growing desperation in Indiana as thousands still without power

September 25, 2026
McDonald’s promises to bring back colorful PlayPlaces in bid to win over frustrated parents

McDonald’s promises to bring back colorful PlayPlaces in bid to win over frustrated parents

September 25, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!