• bitcoinBitcoin(BTC)$85,782.005.58%
  • ethereumEthereum(ETH)$2,746.273.27%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$791.631.29%
  • rippleXRP(XRP)$1.527.87%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$117.925.92%
  • tronTRON(TRX)$0.3476831.34%
  • zcashZcash(ZEC)$1,460.98-3.08%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.011.28%
  • HyperliquidHyperliquid(HYPE)$93.390.42%
  • dogecoinDogecoin(DOGE)$0.09956513.16%
  • moneroMonero(XMR)$586.042.48%
  • whitebitWhiteBIT Coin(WBT)$86.283.88%
  • RainRain(RAIN)$0.013872-2.14%
  • chainlinkChainlink(LINK)$13.043.45%
  • USDSUSDS(USDS)$1.00-0.01%
  • cardanoCardano(ADA)$0.2480497.97%
  • leo-tokenLEO Token(LEO)$8.950.25%
  • stellarStellar(XLM)$0.2145918.65%
  • nearNEAR Protocol(NEAR)$4.568.91%
  • uniswapUniswap(UNI)$9.245.99%
  • bitcoin-cashBitcoin Cash(BCH)$266.764.97%
  • avalanche-2Avalanche(AVAX)$11.19-0.48%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • litecoinLitecoin(LTC)$61.053.87%
  • CantonCanton(CC)$0.1180186.08%
  • daiDai(DAI)$1.000.01%
  • USD1USD1(USD1)$1.00-0.01%
  • suiSui(SUI)$1.0613.90%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.454.71%
  • hedera-hashgraphHedera(HBAR)$0.0926567.86%
  • BittensorBittensor(TAO)$321.7820.93%
  • shiba-inuShiba Inu(SHIB)$0.0000068.82%
  • crypto-com-chainCronos(CRO)$0.0677239.94%
  • MemeCoreMemeCore(M)$1.44-4.59%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • paypal-usdPayPal USD(PYUSD)$1.000.01%
  • tether-goldTether Gold(XAUT)$4,350.73-0.41%
  • okbOKB(OKB)$123.082.87%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.15-0.01%
  • aaveAave(AAVE)$144.374.79%
  • OndoOndo(ONDO)$0.4431184.12%
  • EthenaEthena(ENA)$0.212556-0.13%
  • mantleMantle(MNT)$0.656.22%
  • pepePepe(PEPE)$0.00000525.38%
  • BitwayBitway(BTW)$0.784.63%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Patronus AI secures $17M to tackle AI hallucinations and copyright violations, fuel enterprise adoption

May 22, 2024
in AI & Technology
Reading Time: 5 mins read
A A
Patronus AI secures M to tackle AI hallucinations and copyright violations, fuel enterprise adoption
ShareShareShareShareShare

Join us in returning to NYC on June 5th to collaborate with executive leaders in exploring comprehensive methods for auditing AI models regarding bias, performance, and ethical compliance across diverse organizations. Find out how you can attend here.


As companies race to implement generative AI, concerns about the accuracy and safety of large language models (LLMs) threaten to derail widespread enterprise adoption. Stepping into the fray is Patronus AI, a San Francisco startup that just raised $17 million in Series A funding to automatically detect costly — and potentially dangerous — LLM mistakes at scale.

YOU MAY ALSO LIKE

Why It’s Important To Unplug Your PC During A Power Outage

Why Is Your Laptop Fan So Loud?

The round, which brings Patronus AI’s total funding to $20 million, was led by Glenn Solomon at Notable Capital, with participation from Lightspeed Venture Partners, former DoorDash executive Gokul Rajaram, Factorial Capital, Datadog, and several unnamed tech executives. 

Founded by former Meta machine learning (ML) experts Anand Kannappan and Rebecca Qian, Patronus AI has developed a first-of-its-kind automated evaluation platform that promises to identify errors like hallucinations, copyright infringement and safety violations in LLM outputs. Using proprietary AI, the system scores model performance, stress-tests models with adversarial examples and enables granular benchmarking — all without the manual effort required by most enterprises today.

Exposing the dark side of generative AI: hallucinations, copyright violations and safety risks

“There’s a range of things that our product is actually really good at being able to catch, in terms of mistakes,” said Kannappan, CEO of Patronus AI, in an interview with VentureBeat. “It includes things like hallucinations, and copyright and safety related risks, as well as a lot of enterprise-specific capabilities around things like style and tone of voice of the brand.”

VB Event

The AI Impact Tour: The AI Audit

Join us as we return to NYC on June 5th to engage with top executive leaders, delving into strategies for auditing AI models to ensure fairness, optimal performance, and ethical compliance across diverse organizations. Secure your attendance for this exclusive invite-only event.

Request an invite

The emergence of powerful LLMs like OpenAI’s GPT-4o and Meta’s Llama 3 has set off an arms race in Silicon Valley to capitalize on the technology’s generative abilities. But as hype cycles accelerate, so too have high-profile model failures, from news site CNET publishing error-riddled AI-generated articles to drug discovery startups retracting research papers based on LLM-hallucinated molecules.

These public missteps only scratch the surface of broader issues endemic to the current crop of LLMs, Patronus AI claims. The company’s previously published research, including the “CopyrightCatcher” API released three months ago and the “FinanceBench” benchmark unveiled six months ago, reveals startling deficiencies in leading models’ ability to accurately answer questions grounded in fact.

FinanceBench and CopyrightCatcher: Patronus AI’s groundbreaking research reveals LLM deficiencies

For its “FinanceBench” benchmark, Patronus tasked models like GPT-4 with answering financial queries based on public SEC filings. Shockingly, the best performing model answered only 19% of questions correctly after ingesting an entire annual report. A separate experiment with Patronus’ new “CopyrightCatcher” API found open-source LLMs reproducing copyrighted text verbatim in 44% of outputs.

“Even state-of-the-art models were hallucinating and only got like 90% of responses correct in finance settings,” explained Qian, who serves as CTO. “Our research has shown that open source models had over 20% unsafe responses in many high priority areas of harm. And copyright infringement is a huge risk — large publishers, media companies, or anyone using LLMs needs to be concerned.”

While a handful of other startups like Credo AI, Weights & Biases and Robust Intelligence are building tools for LLM evaluation, Patronus believes its research-first approach leveraging the founders’ deep expertise sets it apart. The core technology is based on training dedicated evaluation models that reliably surface edge cases where a given LLM is likely to fail.

“No other company right now has the research and technology at the level of depth that we have as a company,” Kannappan said. “What’s really unique about how we’ve approached everything is our research-first approach — that’s in the form of training evaluation models, developing new alignment techniques, publishing research papers.”

This strategy has already gained traction with several Fortune 500 companies spanning industries like automotive, education, finance and software using Patronus AI to deploy LLMs “safely within their organizations,” per the startup, though it declined to name specific customers. With the fresh capital, Patronus plans to scale up its research, engineering and sales teams while developing additional industry benchmarks.

If Patronus achieves its vision, rigorous automated evaluation of LLMs could become table stakes for enterprises looking to deploy the technology, in the same way security audits paved the way for widespread cloud adoption. Qian sees a future where testing models with Patronus is as commonplace as unit-testing code.

“Our platform is domain-agnostic and so the evaluation technology that we build can be extended to any domain, whether that’s legal, healthcare or others,” she said. “We want to enable enterprises across every industry to leverage the power of LLMs while having assurance the models are safe and aligned with their specific use case requirements.” 

Still, given the black-box nature of foundation models and near-endless space of possible outputs, conclusively validating an LLM’s performance remains an open challenge. By advancing the state-of-the-art in AI evaluation, Patronus aims to accelerate the path to accountable real-world deployment.

“Measuring LLM performance in an automated way is really difficult and that’s just because there’s such a wide space of behavior, given that these models are generative by nature,” acknowledged Kannappan. “But through a research-driven approach, we’re able to catch mistakes in a very reliable and scalable way that manual testing fundamentally cannot.”

VB Daily

Stay in the know! Get the latest news in your inbox daily

By subscribing, you agree to VentureBeat’s Terms of Service.

Thanks for subscribing. Check out more VB newsletters here.

An error occured.

Credit: Source link

ShareTweetSendSharePin

Related Posts

Why It’s Important To Unplug Your PC During A Power Outage
AI & Technology

Why It’s Important To Unplug Your PC During A Power Outage

September 22, 2026
Why Is Your Laptop Fan So Loud?
AI & Technology

Why Is Your Laptop Fan So Loud?

September 22, 2026
AWS Strands Agents Team Releases Strands Harness: An Open-Source Agent Harness With 28% Lower Token Cost at Comparable Accuracy
AI & Technology

AWS Strands Agents Team Releases Strands Harness: An Open-Source Agent Harness With 28% Lower Token Cost at Comparable Accuracy

September 21, 2026
Bungie Leaders Now Say The Studio’s ‘Not Done With Destiny’
AI & Technology

Bungie Leaders Now Say The Studio’s ‘Not Done With Destiny’

September 21, 2026
Next Post
RFK Jr. says he invested K in GameStop after meme stock rally

RFK Jr. says he invested $24K in GameStop after meme stock rally

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Minneapolis police were ‘familiar’ with mass shooting suspect

Minneapolis police were ‘familiar’ with mass shooting suspect

September 18, 2026
Stay Tuned NOW Streaming Behind The Scenes! – Aug 31

Stay Tuned NOW Streaming Behind The Scenes! – Aug 31

September 20, 2026
A Toaster With A Vision

A Toaster With A Vision

September 16, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!