• bitcoinBitcoin(BTC)$81,746.00-1.76%
  • ethereumEthereum(ETH)$2,476.38-3.66%
  • tetherTether(USDT)$1.00-0.03%
  • binancecoinBNB(BNB)$735.97-4.71%
  • rippleXRP(XRP)$1.38-2.49%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$109.46-5.76%
  • tronTRON(TRX)$0.332480-0.93%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.031.18%
  • zcashZcash(ZEC)$1,189.76-10.39%
  • HyperliquidHyperliquid(HYPE)$84.22-4.57%
  • dogecoinDogecoin(DOGE)$0.084154-5.43%
  • moneroMonero(XMR)$541.31-2.27%
  • USDSUSDS(USDS)$1.000.00%
  • whitebitWhiteBIT Coin(WBT)$80.41-2.93%
  • chainlinkChainlink(LINK)$12.75-4.17%
  • cardanoCardano(ADA)$0.232462-9.22%
  • leo-tokenLEO Token(LEO)$8.900.01%
  • RainRain(RAIN)$0.010277-6.54%
  • stellarStellar(XLM)$0.192500-3.79%
  • nearNEAR Protocol(NEAR)$4.51-15.46%
  • bitcoin-cashBitcoin Cash(BCH)$273.86-9.16%
  • litecoinLitecoin(LTC)$63.20-4.27%
  • Ethena USDeEthena USDe(USDE)$1.00-0.02%
  • CantonCanton(CC)$0.119115-0.22%
  • daiDai(DAI)$1.000.02%
  • uniswapUniswap(UNI)$7.16-9.76%
  • avalanche-2Avalanche(AVAX)$10.06-9.32%
  • suiSui(SUI)$1.04-7.80%
  • USD1USD1(USD1)$1.00-0.01%
  • hedera-hashgraphHedera(HBAR)$0.090316-2.64%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.38-4.07%
  • BitwayBitway(BTW)$1.395.97%
  • quant-networkQuant(QNT)$238.82-5.64%
  • tether-goldTether Gold(XAUT)$4,137.860.65%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • shiba-inuShiba Inu(SHIB)$0.000005-2.49%
  • BittensorBittensor(TAO)$269.38-7.39%
  • crypto-com-chainCronos(CRO)$0.060463-4.20%
  • paypal-usdPayPal USD(PYUSD)$1.000.01%
  • EthenaEthena(ENA)$0.207488-8.57%
  • okbOKB(OKB)$125.41-5.20%
  • Pump.funPump.fun(PUMP)$0.005564-10.67%
  • aaveAave(AAVE)$166.60-4.18%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • MemeCoreMemeCore(M)$1.062.16%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.15-0.08%
  • OndoOndo(ONDO)$0.4733700.13%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Arena Secures $200M Series B at $3.1B Valuation to Advance AI Evaluation – Unite.AI

October 8, 2026
in AI & Technology
Reading Time: 4 mins read
A A
Arena Secures 0M Series B at .1B Valuation to Advance AI Evaluation – Unite.AI
ShareShareShareShareShare

AI evaluation company Arena announced a $200 million Series B at a $3.1 billion valuation on October 8, 2026, in a round co-led by Lightspeed Venture Partners and Khosla Ventures, and released a preview of its Arena Alignment Index alongside the financing.

YOU MAY ALSO LIKE

Apple Reportedly Plans Touchscreen MacBook And iPad Mini Launch For This Month

Anthropic Bans ‘Sustained And Needless Abusive Or Cruel Behavior’ Toward Its AI Models

Series B Investors and Stated Purpose

Salesforce Ventures, 01 Advisors, Dell Technologies Capital and Endeavor Catalyst participated in the round, and existing investors a16z, Felicis, AMP PBC, QuantumLight and The House Fund also supported it, the company said.

Arena said the funding continues its mission to measure and advance the AI frontier for real-world use, framing the round as a vote of confidence in the idea that as AI grows more powerful, the world needs an independent, data-driven approach to measure both the capabilities of AI and whether it can be trusted and used safely. The company said agents now write code, run analyses and take actions on people’s behalf rather than simply answering questions, often in areas where the person cannot easily check the work, and argued that static benchmarks break down once models recognize they are being tested. Arena also said it has exceeded $100 million in annualized revenue.

Reported Growth and Earlier Rounds

Arena reported 7 million sessions in Agent Arena in less than five months since its launch, and 350 million sessions across the entire Arena platform. The company said it has collected 62 million votes across text, vision, code, search, video and image modalities, and that it draws tens of millions of monthly visitors from more than 150 countries. It also reported more than 1,000 new model evaluations and 375,000 data points open-sourced to the community, including its leaderboard methodology.

The Series B follows a $150 million Series A the company announced in January 2026, when it still operated under the LMArena name. Felicis and UC Investments (University of California) led that round, with participation from Andreessen Horowitz, The House Fund, LDVP, Kleiner Perkins, Lightspeed Venture Partners and Laude Ventures. The company had announced a $100 million seed round in May 2025.

At the time of the Series A, the company reported 50 million votes, more than 400 new model evaluations and 145,000 open-source battle data points, and said its first evaluation product had launched in September 2025. Arena started as a research experiment, according to the company, which said it began with human preference evaluations and later moved into measuring factuality after finding that the response people prefer is not always the correct one.

Alignment Index Signals and Methodology

The Alignment Index, which Arena describes as a measurement of how AI can deviate from human values in the real world, begins with three signals the company says can be verified against an actual agent trace from real human use: Unauthorized Action, where the model acts beyond what the user asked; False Attribution, where the model attributes a statement, intention or fact to the user that is contradicted by user-provided evidence; and Deceptive Completion, where the model reports a task as complete when it is not.

Arena said the signal definitions are inspired by definitions OpenAI and Anthropic have published in their system cards, so they can serve as an independent assessment of model safety and alignment. The company positions the three signals as complementary to its Agent Arena leaderboard, which ranks agent capabilities.

In a companion research post, Arena said the preview compares 27 models across 90,000 real-world agent sessions drawn from Agent Arena. For each signal, the company wrote rubrics describing recurring failure modes and refined them over repeated rounds of judging and human review. An LLM judge then applied the rubrics to sampled sessions for each model, flagging a session only when it could point to a specific claim or action with supporting evidence, and the reported rates were adjusted for conversation length.

To compute the index, Arena transforms each signal’s flagged rate using one minus the square root of that rate, an approach it says keeps improvements near full alignment visible, then combines the three scores with 50% weight on Unauthorized Action and 25% each on False Attribution and Deceptive Completion. Higher values indicate a safer and better-aligned model, according to the company.

Initial Findings and Next Steps

OpenAI’s GPT-6.1 Sol leads the published preview at 87.9, followed by Anthropic’s Claude Opus 5.5 at 83.2 and SpaceXAI’s Grok 4.7 at 82.7. Arena reported that OpenAI models hold the top five positions among the 27 models scored, with four of them at about 88 points.

Arena reported that about 2% of Claude Opus 5 sessions included an unauthorized action, and that 53.5% of those cases involved deleting or cleaning up the user’s files or earlier work without permission; in Claude Opus 5.5, the cleanup share fell to 20.0%. Deceptive completions affected an average of 10% of sessions, rising to 48.0% in code debugging, the highest rate of any task category, while unauthorized-action detections peaked in code explanation at 6.3% and false attribution was highest in professional writing at 13.7%.

Detection rates rose with conversation length across all three signals: in sessions with 20 or more user messages, deceptive completion was flagged in 45.4% of sessions and unauthorized action in 12.4%. Arena also reported that the newest models in the GPT, Claude, Gemini Flash and Grok lineages show lower detection rates than their predecessors on most signals, noting GPT-6.1 Sol’s slightly higher false attribution than GPT-6 Sol and Claude Opus 5’s slightly higher unauthorized action than Claude Opus 4.8 as exceptions.

Arena said it will add more safety-related signals to the index, starting with how well models refuse harmful prompts, and will expand it to new models and real-world settings while continuing to update the leaderboard signal by signal.

Credit: Source link

ShareTweetSendSharePin

Related Posts

Apple Reportedly Plans Touchscreen MacBook And iPad Mini Launch For This Month
AI & Technology

Apple Reportedly Plans Touchscreen MacBook And iPad Mini Launch For This Month

October 8, 2026
Anthropic Bans ‘Sustained And Needless Abusive Or Cruel Behavior’ Toward Its AI Models
AI & Technology

Anthropic Bans ‘Sustained And Needless Abusive Or Cruel Behavior’ Toward Its AI Models

October 8, 2026
Anthropic Launches Cyber Mission for Critical Infrastructure, Open Source – Unite.AI
AI & Technology

Anthropic Launches Cyber Mission for Critical Infrastructure, Open Source – Unite.AI

October 8, 2026
The Atari 800XL Hybrid Gaming Console Is Back After Decades In Obscurity
AI & Technology

The Atari 800XL Hybrid Gaming Console Is Back After Decades In Obscurity

October 8, 2026
Next Post
The Atari 800XL Hybrid Gaming Console Is Back After Decades In Obscurity

The Atari 800XL Hybrid Gaming Console Is Back After Decades In Obscurity

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
SpaceX Launches Crew for NASA on Retiring Dragon Capsule

SpaceX Launches Crew for NASA on Retiring Dragon Capsule

October 2, 2026
What Are The Benefits Of Having A Second Monitor?

What Are The Benefits Of Having A Second Monitor?

October 5, 2026
Trump says AI tech leaders agree to safety accord

Trump says AI tech leaders agree to safety accord

October 4, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!