• bitcoinBitcoin(BTC)$77,562.00-1.42%
  • ethereumEthereum(ETH)$2,415.02-2.38%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$687.82-0.79%
  • rippleXRP(XRP)$1.35-2.61%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$100.32-3.28%
  • tronTRON(TRX)$0.321770-3.12%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.054.38%
  • HyperliquidHyperliquid(HYPE)$82.99-1.52%
  • zcashZcash(ZEC)$840.61-1.97%
  • dogecoinDogecoin(DOGE)$0.081498-2.05%
  • RainRain(RAIN)$0.0172312.97%
  • USDSUSDS(USDS)$1.000.00%
  • moneroMonero(XMR)$515.62-0.11%
  • leo-tokenLEO Token(LEO)$9.330.00%
  • whitebitWhiteBIT Coin(WBT)$71.31-1.68%
  • chainlinkChainlink(LINK)$11.23-1.86%
  • cardanoCardano(ADA)$0.197923-1.51%
  • stellarStellar(XLM)$0.175890-1.16%
  • bitcoin-cashBitcoin Cash(BCH)$249.190.54%
  • daiDai(DAI)$1.00-0.02%
  • CantonCanton(CC)$0.115466-6.37%
  • Ethena USDeEthena USDe(USDE)$1.00-0.03%
  • USD1USD1(USD1)$1.00-0.02%
  • litecoinLitecoin(LTC)$49.270.70%
  • uniswapUniswap(UNI)$5.9911.22%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.31-5.76%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • hedera-hashgraphHedera(HBAR)$0.073892-0.66%
  • avalanche-2Avalanche(AVAX)$7.21-0.87%
  • shiba-inuShiba Inu(SHIB)$0.0000050.59%
  • suiSui(SUI)$0.72-1.45%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.02%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • crypto-com-chainCronos(CRO)$0.054813-1.92%
  • tether-goldTether Gold(XAUT)$4,294.92-3.09%
  • nearNEAR Protocol(NEAR)$1.88-4.74%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • MemeCoreMemeCore(M)$1.05-3.28%
  • okbOKB(OKB)$109.87-1.95%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.17%
  • BittensorBittensor(TAO)$220.39-4.42%
  • aaveAave(AAVE)$129.232.60%
  • AsterAster(ASTER)$0.70-0.22%
  • pax-goldPAX Gold(PAXG)$4,301.60-3.13%
  • MorphoMorpho(MORPHO)$2.675.74%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.057284-0.31%
  • mantleMantle(MNT)$0.54-1.29%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

A single AI agent conversation can look perfect and still be broken, leaders from LangChain, Conviva and CoreWeave said at VB Transform 2026

July 20, 2026
in AI & Technology
Reading Time: 4 mins read
A A
A single AI agent conversation can look perfect and still be broken, leaders from LangChain, Conviva and CoreWeave said at VB Transform 2026
ShareShareShareShareShare

A single AI agent conversation can look flawless scored on its own and still point to a broken product. That gap is driving a shift in how enterprises evaluate agents, away from scoring individual traces and toward comparing cohorts of users against a baseline.

YOU MAY ALSO LIKE

This Is The Best Setting And Placement For Your Dolby Atmos Soundbar

Aramco Digital and Avathon Partner on Autonomous Operations AI – Unite.AI

At VB Transform 2026, Harrison Chase, CEO of LangChain; Hui Zhang, CTO and co-founder of Conviva; and Emmanuel Turlay, director of engineering at CoreWeave, described that shift, along with a parallel move toward cheaper, narrower judge models.

Agent-as-judge — judging one AI agent’s output with another — hasn’t replaced LLM-as-judge, which Chase said remains the default. The larger tension, Zhang said, is between automated judging, whether by LLM or agent, and human review.

“You have scalable but ungrounded, whether it’s agents as judge or LLMs as judge, you grade the outcome, you grade the work. It still is very difficult to ground it and then you use humans and that’s just not scalable,” Zhang said. “The whole industry is facing this, which poison you want to pick.”

Evaluation criteria now function as the product spec

That gap — a conversation that scores well but still signals a broken product — is what teams try to close by building an exhaustive evaluation suite before they ship anything. Chase said that doesn’t work.

“We sometimes see teams that have almost eval paralysis,” Chase said. “They’re like, this is an eval set, I can’t launch it. The best teams launch and then iterate.”

Chase framed evaluation criteria as a living specification, not a one-time test suite: a product requirements document — the standard software-development spec for what an application should do. “Evals are like the new PRD,” he said. “They define what your agent should and shouldn’t do.”

Turlay described hitting the same failure from a different angle. “I was trying to reach 100% coverage for my tests, and I still had bugs in production,” he said — a test suite that looked complete but still missed what mattered, the same gap Chase was describing with evals.

Broad, always-on monitoring, he said, catches more real failures than an exhaustive pre-launch test suite. Teams should set up wide online checks first, use those to identify failure classes as they occur, then build a targeted offline evaluation set around the problems that surface.

Why scoring traces one at a time is a mistake

Even a well-built evaluation process can still score the wrong thing. Zhang’s objection is to how most teams run evaluation: sampling traces, whether 50 of them or a full population, scoring each in isolation. That approach misses a signal that only shows up when comparing cohorts of users against a baseline, a method Zhang calls contrastive analysis.

Zhang illustrated it with a retail example: a shopper asks an agent for a running shoe ahead of a half marathon, the agent asks qualifying questions, and the shopper buys a shoe. Scored individually, that interaction looks fine. But the clarification ratio, how many follow-up questions an agent asks before completing a task, came in three times higher than baseline for that shoe category across the full user population. A second metric, how often shoppers finished their purchase outside the conversation, was five times higher than baseline for the same category.

Neither number is visible from a single trace. Both point to a debuggable, category-specific problem. Zhang said the industry also lacks a second data source: what happens before, between and after the conversation, not just the trace itself.

Sizing the judge to the job

Once contrastive analysis flags which category is actually broken, the next problem is what watches for it going forward — and at what cost. Turlay’s rule was to start with the most capable model available to prove a task is solvable, then work down. If it can’t be done with a top-tier model, he said, it won’t work with a smaller one. Once a pattern proves viable, teams can sample a fraction of traffic instead of judging every interaction, and move simpler tasks like binary classification to smaller open source models.

LangChain took that further, fine-tuning its own model to detect when a user believes the agent made a mistake, a signal Chase calls perceived error. “The model we fine-tuned was a Qwen model,” he said, referring to Alibaba’s open source family. Combining hand labeling with distillation, the result performed well. “Same as [Claude]Sonnet, for, depending on how we served it, either 10 to 100x cost reduction,” Chase said.

Not every guardrail needs a model. Chase pointed to Claude Code’s own guardrails as proof: regexes, the common programming technique for finding and validating patterns in code. “A lot of the guardrails they had were just regexes,” he said. “They weren’t small LLMs, they were just regexes.”

LLM-as-judge doesn’t mean human-in-the-loop disappears

The bigger question is whether using LLM as a judge removes the need for a human in the loop.

Turlay pointed to accountability, drawing on his prior work at a self-driving car company. His team compressed data intake and retraining into a two-week cycle for shipping a new model to the car. Even then, someone still had to sign off.

“I felt confident on behalf of the company to say this model should go into the car,” he said. The same logic extends to legal, finance and healthcare. “Before we can remove a human to say, I endorse this and I take responsibility legally for it, it’s going to be a while before agents can do that on their own.”

Zhang agreed a human has to remain the guardian on corner cases, even as automation eventually runs at a scale that beats individual human accuracy — machines can see more at the pattern level.

Chase went further: that human check isn’t just a safety net. “Human in the loop is really important for building trust in how these agentic systems work, and also really important for memory and learning from systems,” he said. “There has to be interactions in order for the system to learn.”

Credit: Source link

ShareTweetSendSharePin

Related Posts

This Is The Best Setting And Placement For Your Dolby Atmos Soundbar
AI & Technology

This Is The Best Setting And Placement For Your Dolby Atmos Soundbar

September 2, 2026
Aramco Digital and Avathon Partner on Autonomous Operations AI – Unite.AI
AI & Technology

Aramco Digital and Avathon Partner on Autonomous Operations AI – Unite.AI

September 1, 2026
Anthropic Releases Claude Fable 5.1 and Claude Mythos 5.1: 52.6% on Terminal-Bench-Science and 75% Cheaper Cache Reads
AI & Technology

Anthropic Releases Claude Fable 5.1 and Claude Mythos 5.1: 52.6% on Terminal-Bench-Science and 75% Cheaper Cache Reads

September 1, 2026
Frontier models can recover up to 65% of facts they can’t directly recall — just by thinking longer
AI & Technology

Frontier models can recover up to 65% of facts they can’t directly recall — just by thinking longer

September 1, 2026
Next Post
California gas suffers setback as judge slaps down oil giants’ claims

California gas suffers setback as judge slaps down oil giants' claims

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
NBA coaching pioneer Don Nelson dies at 86

NBA coaching pioneer Don Nelson dies at 86

August 27, 2026
Violent house explosion jolts families awake in Ohio

Violent house explosion jolts families awake in Ohio

September 1, 2026
Measles’ role in child’s death ignites fight between RFK Jr. and Pennsylvania governor – The Washington Post

Measles’ role in child’s death ignites fight between RFK Jr. and Pennsylvania governor – The Washington Post

August 29, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!