• bitcoinBitcoin(BTC)$64,345.00-1.20%
  • ethereumEthereum(ETH)$1,858.90-2.80%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$560.58-1.10%
  • usd-coinUSDC(USDC)$1.000.00%
  • rippleXRP(XRP)$1.10-2.90%
  • solanaSolana(SOL)$74.39-3.60%
  • tronTRON(TRX)$0.3303241.10%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.043.70%
  • whitebitWhiteBIT Coin(WBT)$56.06-1.60%
  • HyperliquidHyperliquid(HYPE)$58.09-1.60%
  • dogecoinDogecoin(DOGE)$0.068488-4.50%
  • RainRain(RAIN)$0.0142890.50%
  • USDSUSDS(USDS)$1.000.00%
  • leo-tokenLEO Token(LEO)$9.59-1.30%
  • zcashZcash(ZEC)$496.48-3.60%
  • moneroMonero(XMR)$353.731.00%
  • chainlinkChainlink(LINK)$8.36-2.40%
  • stellarStellar(XLM)$0.180265-1.90%
  • cardanoCardano(ADA)$0.164510-5.10%
  • CantonCanton(CC)$0.118917-1.00%
  • daiDai(DAI)$1.000.00%
  • bitcoin-cashBitcoin Cash(BCH)$208.08-3.10%
  • USD1USD1(USD1)$1.00-0.10%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.45-4.00%
  • litecoinLitecoin(LTC)$46.17-1.40%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • hedera-hashgraphHedera(HBAR)$0.070518-3.20%
  • Circle USYCCircle USYC(USYC)$1.130.00%
  • suiSui(SUI)$0.72-5.20%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.056550-2.70%
  • avalanche-2Avalanche(AVAX)$6.16-5.60%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,042.19-0.80%
  • shiba-inuShiba Inu(SHIB)$0.000004-2.00%
  • nearNEAR Protocol(NEAR)$1.83-1.60%
  • uniswapUniswap(UNI)$3.77-0.60%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.00%
  • OndoOndo(ONDO)$0.383716-3.10%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.057573-10.90%
  • BittensorBittensor(TAO)$187.77-3.30%
  • pax-goldPAX Gold(PAXG)$4,038.28-0.90%
  • okbOKB(OKB)$81.81-3.00%
  • AsterAster(ASTER)$0.620.30%
  • HTX DAOHTX DAO(HTX)$0.000002-0.40%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • MemeCoreMemeCore(M)$1.204.40%
  • usddUSDD(USDD)$1.000.00%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

OpenAI Introduces GDPval: A New Evaluation Suite that Measures AI on Real-World Economically Valuable Tasks

September 25, 2025
in AI & Technology
Reading Time: 7 mins read
A A
OpenAI Introduces GDPval: A New Evaluation Suite that Measures AI on Real-World Economically Valuable Tasks
ShareShareShareShareShare

OpenAI introduced GDPval, a new evaluation suite designed to measure how AI models perform on real-world, economically valuable tasks across 44 occupations in nine GDP-dominant U.S. sectors. Unlike academic benchmarks, GDPval centers on authentic deliverables—presentations, spreadsheets, briefs, CAD artifacts, audio/video—graded by occupational experts through blinded pairwise comparisons. OpenAI also released a 220-task “gold” subset and an experimental automated grader hosted at evals.openai.com.

From Benchmarks to Billables: How GDPval Builds Tasks

GDPval aggregates 1,320 tasks sourced from industry professionals averaging 14 years of experience. Tasks map to O*NET work activities and include multi-modal file handling (docs, slides, images, audio, video, spreadsheets, CAD), with up to dozens of reference files per task. The gold subset provides public prompts and references; primary scoring still relies on expert pairwise judgments due to subjectivity and format requirements.

YOU MAY ALSO LIKE

China’s AI Ambitions on Display at Tech Conference

Apple to Launch Leasing Program to Spur Sales

https://openai.com/index/gdpval/

What the Data Says: Model vs. Expert

On the gold subset, frontier models approach expert quality on a substantial fraction of tasks under blind expert review, with model progress trending roughly linearly across releases. Reported model-vs-human win/tie rates near parity for top models, error profiles cluster around instruction-following, formatting, data usage, and hallucinations. Increased reasoning effort and stronger scaffolding (e.g., format checks, artifact rendering for self-inspection) yield predictable gains.

Time–Cost Math: Where AI Pays Off

GDPval runs scenario analyses comparing human-only to model-assisted workflows with expert review. It quantifies (i) human completion time and wage-based cost, (ii) reviewer time/cost, (iii) model latency and API cost, and (iv) empirically observed win rates. Results indicate potential time/cost reductions for many task classes once review overhead is included.

Automated Judging: Useful Proxy, Not Oracle

For the gold subset, an automated pairwise grader shows ~66% agreement with human experts, within ~5 percentage points of human–human agreement (~71%). It’s positioned as an accessibility proxy for rapid iteration, not a replacement for expert review.

https://openai.com/index/gdpval/

Why This Isn’t Yet Another Benchmark

  • Occupational breadth: Spans top GDP sectors and a wide slice of O*NET work activities, not just narrow domains.
  • Deliverable realism: Multi-file, multi-modal inputs/outputs stress structure, formatting, and data handling.
  • Moving ceiling: Uses human preference win rate against expert deliverables, enabling re-baselining as models improve.

Boundary Conditions: Where GDPval Doesn’t Reach

GDPval-v0 targets computer-mediated knowledge work. Physical labor, long-horizon interactivity, and organization-specific tooling are out of scope. Tasks are one-shot and precisely specified; ablations show performance drops with reduced context. Construction and grading are resource-intensive, motivating the automated grader—whose limits are documented—and future expansion.

Fit in the Stack: How GDPval Complements Other Evals

GDPval augments existing OpenAI evals with occupational, multi-modal, file-centric tasks and reports human preference outcomes, time/cost analyses, and ablations on reasoning effort and agent scaffolding. v0 is versioned and expected to broaden coverage and realism over time.

Summary

GDPval formalizes evaluation for economically relevant knowledge work by pairing expert-built tasks with blinded human preference judgments and an accessible automated grader. The framework quantifies model quality and practical time/cost trade-offs while exposing failure modes and the effects of scaffolding and reasoning effort. Scope remains v0—computer-mediated, one-shot tasks with expert review—yet it establishes a reproducible baseline for tracking real-world capability gains across occupations.


Check out the Paper, Technical details, and Dataset on Hugging Face. Feel free to check out our GitHub Page for Tutorials, Codes and Notebooks. Also, feel free to follow us on Twitter and don’t forget to join our 100k+ ML SubReddit and Subscribe to our Newsletter.


Michal Sutter is a data science professional with a Master of Science in Data Science from the University of Padova. With a solid foundation in statistical analysis, machine learning, and data engineering, Michal excels at transforming complex datasets into actionable insights.

🔥[Recommended Read] NVIDIA AI Open-Sources ViPE (Video Pose Engine): A Powerful and Versatile 3D Video Annotation Tool for Spatial AI

Credit: Source link

ShareTweetSendSharePin

Related Posts

China’s AI Ambitions on Display at Tech Conference
AI & Technology

China’s AI Ambitions on Display at Tech Conference

July 24, 2026
Apple to Launch Leasing Program to Spur Sales
AI & Technology

Apple to Launch Leasing Program to Spur Sales

July 24, 2026
Samsung’s Foldables Are Here, and Apple Is Next
AI & Technology

Samsung’s Foldables Are Here, and Apple Is Next

July 24, 2026
How Demand for AI Compute Is Just Scratching the Surface
AI & Technology

How Demand for AI Compute Is Just Scratching the Surface

July 24, 2026
Next Post
Scooter company fleeces Americans out of millions

Scooter company fleeces Americans out of millions

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Morning News NOW Full Episode – July 3

Morning News NOW Full Episode – July 3

July 21, 2026
Number of Americans ‘not in the labor force’ surges to record 105.8M as total exceeds Great Recession, COVID era

Number of Americans ‘not in the labor force’ surges to record 105.8M as total exceeds Great Recession, COVID era

July 22, 2026
AliExpress Hit With Record 9 Million Fine For Selling Counterfeit And Illegal Products

AliExpress Hit With Record $629 Million Fine For Selling Counterfeit And Illegal Products

July 20, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!