• bitcoinBitcoin(BTC)$84,465.00-2.02%
  • ethereumEthereum(ETH)$2,683.87-2.65%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$765.64-2.92%
  • rippleXRP(XRP)$1.50-4.83%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$114.94-3.07%
  • tronTRON(TRX)$0.341864-0.05%
  • zcashZcash(ZEC)$1,501.24-7.77%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.03-0.54%
  • HyperliquidHyperliquid(HYPE)$94.01-2.94%
  • dogecoinDogecoin(DOGE)$0.092575-7.88%
  • moneroMonero(XMR)$551.52-3.60%
  • whitebitWhiteBIT Coin(WBT)$84.78-2.23%
  • USDSUSDS(USDS)$1.000.00%
  • chainlinkChainlink(LINK)$12.35-5.42%
  • cardanoCardano(ADA)$0.238468-6.99%
  • RainRain(RAIN)$0.012244-6.42%
  • leo-tokenLEO Token(LEO)$8.96-0.16%
  • stellarStellar(XLM)$0.202331-6.56%
  • bitcoin-cashBitcoin Cash(BCH)$336.86-1.94%
  • uniswapUniswap(UNI)$9.30-7.76%
  • nearNEAR Protocol(NEAR)$4.30-2.63%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • litecoinLitecoin(LTC)$62.04-2.24%
  • daiDai(DAI)$1.00-0.01%
  • avalanche-2Avalanche(AVAX)$10.32-8.23%
  • USD1USD1(USD1)$1.000.01%
  • CantonCanton(CC)$0.109900-4.63%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.42-4.05%
  • hedera-hashgraphHedera(HBAR)$0.090575-8.80%
  • suiSui(SUI)$0.96-6.64%
  • shiba-inuShiba Inu(SHIB)$0.000006-7.88%
  • BittensorBittensor(TAO)$287.60-9.27%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • crypto-com-chainCronos(CRO)$0.061142-8.75%
  • BitwayBitway(BTW)$1.0417.17%
  • MemeCoreMemeCore(M)$1.22-7.07%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,289.52-1.65%
  • okbOKB(OKB)$118.89-3.41%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.23%
  • mantleMantle(MNT)$0.65-2.51%
  • aaveAave(AAVE)$139.09-5.91%
  • EthenaEthena(ENA)$0.205530-6.19%
  • OndoOndo(ONDO)$0.413498-6.66%
  • AsterAster(ASTER)$0.69-5.21%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

ServiceNow AI Research Releases DRBench, a Realistic Enterprise Deep-Research Benchmark

October 14, 2025
in AI & Technology
Reading Time: 8 mins read
A A
ServiceNow AI Research Releases DRBench, a Realistic Enterprise Deep-Research Benchmark
ShareShareShareShareShare

ServiceNow Research has released DRBench, a benchmark and runnable environment to evaluate “deep research” agents on open-ended enterprise tasks that require synthesizing facts from both public web and private organizational data into properly cited reports. Unlike web-only testbeds, DRBench stages heterogeneous, enterprise-style workflows—files, emails, chat logs, and cloud storage—so agents must retrieve, filter, and attribute insights across multiple applications before writing a coherent research report.

https://arxiv.org/abs/2510.00172

What DRBench contains?

The initial release provides 15 deep research tasks across 10 enterprise domains (e.g., Sales, Cybersecurity, Compliance). Each task specifies a deep research question, a task context (company and persona), and a set of groundtruth insights spanning three classes: public insights (from dated, time-stable URLs), internal relevant insights, and internal distractor insights. The benchmark explicitly embeds these insights within realistic enterprise files and applications, forcing agents to surface the relevant ones while avoiding distractors. The dataset construction pipeline combines LLM generation with human verification and totals 114 groundtruth insights across tasks.

YOU MAY ALSO LIKE

Meta Brings FDA-Cleared Hearing Enhancement To Its Smart Glasses

Microsoft’s New Surface Pro 12 And Surface Laptop 13 Feature Snapdragon X2 Plus Chips

https://arxiv.org/abs/2510.00172

Enterprise environment

A core contribution is the containerized enterprise environment that integrates commonly used services behind authentication and app-specific APIs. DRBench’s Docker image orchestrates: Nextcloud (shared documents, WebDAV), Mattermost (team chat, REST API), Roundcube with SMTP/IMAP (enterprise email), FileBrowser (local filesystem), and a VNC/NoVNC desktop for GUI interaction. Tasks are initialized by distributing data across these services (documents to Nextcloud and FileBrowser, chats to Mattermost channels, threaded emails to the mail system, and provisioned users with consistent credentials). Agents can operate through web interfaces or programmatic APIs exposed by each service. This setup is intentionally “needle-in-a-haystack”: relevant and distractor insights are injected into realistic files (PDF/DOCX/PPTX/XLSX, chats, emails) and padded with plausible but irrelevant content.

Evaluation: what gets scored

DRBench evaluates four axes aligned to analyst workflows: Insight Recall, Distractor Avoidance, Factuality, and Report Quality. Insight Recall decomposes the agent’s report into atomic insights with citations, matches them against groundtruth injected insights using an LLM judge, and scores recall (not precision). Distractor Avoidance penalizes inclusion of injected distractor insights. Factuality and Report Quality assess the correctness and structure/clarity of the final report under a rubric specified in the report.

https://arxiv.org/abs/2510.00172

Baseline agent and research loop

The research team introduces a task-oriented baseline, DRBench Agent (DRBA), designed to operate natively inside the DRBench environment. DRBA is organized into four components: research planning, action planning, a research loop with Adaptive Action Planning (AAP), and report writing. Planning supports two modes: Complex Research Planning (CRP), which specifies investigation areas, expected sources, and success criteria; and Simple Research Planning (SRP), which produces lightweight sub-queries. The research loop iteratively selects tools, processes content (including storage in a vector store), identifies gaps, and continues until completion or a max-iteration budget; the report writer synthesizes findings with citation tracking.

Why this is important for enterprise agents?

Most “deep research” agents look compelling on public-web question sets, but production usage hinges on reliably finding the right internal needles, ignoring plausible internal distractors, and citing both public and private sources under enterprise constraints (login, permissions, UI friction). DRBench’s design directly targets this gap by: (1) grounding tasks in realistic company/persona contexts; (2) distributing evidence across multiple enterprise apps plus the web; and (3) scoring whether the agent actually extracted the intended insights and wrote a coherent, factual report. This combination makes it a practical benchmark for system builders who need end-to-end evaluation rather than single-tool micro-scores.

https://arxiv.org/abs/2510.00172

Key Takeaways

  • DRBench evaluates deep research agents on complex, open-ended enterprise tasks that require combining public web and private company data.
  • The initial release covers 15 tasks across 10 domains, each grounded in realistic user personas and organizational context.
  • Tasks span heterogeneous enterprise artifacts—productivity software, cloud file systems, emails, chat—plus the open web, going beyond web-only setups.
  • Reports are scored for insight recall, factual accuracy, and coherent, well-structured reporting using rubric-based evaluation.
  • Code and benchmark assets are open-sourced on GitHub for reproducible evaluation and extension.

Editorial comments

From an enterprise evaluation standpoint, DRBench is a useful step toward standardized, end-to-end testing of “deep research” agents: the tasks are open-ended, grounded in realistic personas, and require integrating evidence from the public web and a private company knowledge base, then producing a coherent, well-structured report—precisely the workflow most production teams care about. The release also clarifies what’s being measured—recall of relevant insights, factual accuracy, and report quality—while explicitly moving beyond web-only setups that overfit to browsing heuristics. The 15 tasks across 10 domains are modest in scale but sufficient to expose system bottlenecks (retrieval across heterogeneous artifacts, citation discipline, and planning loops).


Check out the Paper and GitHub page. Feel free to check out our GitHub Page for Tutorials, Codes and Notebooks. Also, feel free to follow us on Twitter and don’t forget to join our 100k+ ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.

The post ServiceNow AI Research Releases DRBench, a Realistic Enterprise Deep-Research Benchmark appeared first on MarkTechPost.

Credit: Source link

ShareTweetSendSharePin

Related Posts

Meta Brings FDA-Cleared Hearing Enhancement To Its Smart Glasses
AI & Technology

Meta Brings FDA-Cleared Hearing Enhancement To Its Smart Glasses

September 23, 2026
Microsoft’s New Surface Pro 12 And Surface Laptop 13 Feature Snapdragon X2 Plus Chips
AI & Technology

Microsoft’s New Surface Pro 12 And Surface Laptop 13 Feature Snapdragon X2 Plus Chips

September 23, 2026
Google Releases Gemini 3.8 Flash TTS and Flash-Lite TTS With Prompt-Based Voice Design
AI & Technology

Google Releases Gemini 3.8 Flash TTS and Flash-Lite TTS With Prompt-Based Voice Design

September 23, 2026
NVIDIA Releases Nemotron 3 Diarization: A 100M-Parameter Open-Weight Model That Tracks 8 Speakers in Real Time
AI & Technology

NVIDIA Releases Nemotron 3 Diarization: A 100M-Parameter Open-Weight Model That Tracks 8 Speakers in Real Time

September 23, 2026
Next Post
7 LLM Generation Parameters—What They Do and How to Tune Them?

7 LLM Generation Parameters—What They Do and How to Tune Them?

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Alibaba Qwen Team Releases Qwen3.8-LiveTranslate: A Real-Time Interpretation Model That Cuts Average Lag to 2.3 Seconds Across 60 Languages

Alibaba Qwen Team Releases Qwen3.8-LiveTranslate: A Real-Time Interpretation Model That Cuts Average Lag to 2.3 Seconds Across 60 Languages

September 20, 2026
Police ID suspect in deadly Minneapolis shooting as man facing eviction

Police ID suspect in deadly Minneapolis shooting as man facing eviction

September 18, 2026
Taylor Swift gives K to injured Connecticut hero mom

Taylor Swift gives $50K to injured Connecticut hero mom

September 20, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!