• bitcoinBitcoin(BTC)$84,365.000.38%
  • ethereumEthereum(ETH)$2,694.170.17%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$773.35-0.43%
  • rippleXRP(XRP)$1.53-2.49%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$121.29-0.57%
  • tronTRON(TRX)$0.334201-1.12%
  • zcashZcash(ZEC)$1,650.176.32%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.062.92%
  • HyperliquidHyperliquid(HYPE)$92.480.21%
  • dogecoinDogecoin(DOGE)$0.096604-2.16%
  • chainlinkChainlink(LINK)$14.121.26%
  • moneroMonero(XMR)$558.580.35%
  • whitebitWhiteBIT Coin(WBT)$84.100.26%
  • USDSUSDS(USDS)$1.000.01%
  • cardanoCardano(ADA)$0.253505-1.67%
  • RainRain(RAIN)$0.01288313.30%
  • leo-tokenLEO Token(LEO)$8.961.84%
  • stellarStellar(XLM)$0.215956-1.64%
  • bitcoin-cashBitcoin Cash(BCH)$335.76-1.81%
  • nearNEAR Protocol(NEAR)$5.041.67%
  • uniswapUniswap(UNI)$9.701.07%
  • litecoinLitecoin(LTC)$72.510.48%
  • CantonCanton(CC)$0.1353874.51%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • avalanche-2Avalanche(AVAX)$10.831.85%
  • suiSui(SUI)$1.16-2.29%
  • daiDai(DAI)$1.000.00%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.609.94%
  • USD1USD1(USD1)$1.000.01%
  • hedera-hashgraphHedera(HBAR)$0.093330-2.15%
  • BittensorBittensor(TAO)$320.411.13%
  • shiba-inuShiba Inu(SHIB)$0.000006-0.36%
  • crypto-com-chainCronos(CRO)$0.0679122.48%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • MemeCoreMemeCore(M)$1.220.74%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • EthenaEthena(ENA)$0.2695900.30%
  • BitwayBitway(BTW)$1.00-24.00%
  • tether-goldTether Gold(XAUT)$4,279.37-0.11%
  • OndoOndo(ONDO)$0.54-2.02%
  • okbOKB(OKB)$120.99-0.08%
  • Ripple USDRipple USD(RLUSD)$1.000.03%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • aaveAave(AAVE)$155.781.48%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • mantleMantle(MNT)$0.693.12%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.15-0.06%
  • quant-networkQuant(QNT)$155.2456.30%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

From Prediction to Reasoning: Evaluating o1’s Impact on LLM Probabilistic Biases

October 8, 2024
in AI & Technology
Reading Time: 4 mins read
A A
From Prediction to Reasoning: Evaluating o1’s Impact on LLM Probabilistic Biases
ShareShareShareShareShare

Large language models (LLMs) have gained significant attention in recent years, but understanding their capabilities and limitations remains a challenge. Researchers are trying to develop methodologies to reason about the strengths and weaknesses of AI systems, particularly LLMs. The current approaches often lack a systematic framework for predicting and analyzing these systems’ behaviours. This has led to difficulties in anticipating how LLMs will perform various tasks, especially those that differ from their primary training objective. The challenge lies in bridging the gap between the AI system’s training process and its observed performance on diverse tasks, necessitating a more comprehensive analytical approach.

In this study, researchers from the Wu Tsai Institute, Yale University, OpenAI, Princeton University, Roundtable, and Princeton University have focused on analyzing OpenAI’s new system, o1, which was explicitly optimized for reasoning tasks, to determine if it exhibits the same “embers of autoregression” observed in previous LLMs. The researchers apply the teleological perspective, which considers the pressures shaping AI systems, to predict and evaluate o1’s performance. This approach examines whether o1’s departure from pure next-word prediction training mitigates limitations associated with that objective. The study compares o1’s performance to other LLMs on various tasks, assessing its sensitivity to output probability and task frequency. In addition to that, the researchers introduce a robust metric—token count during answer generation—to quantify task difficulty. This comprehensive analysis aims to reveal whether o1 represents a significant advancement or still retains behavioural patterns linked to next-word prediction training.

YOU MAY ALSO LIKE

Your Old GPU Could Be Worth More Than You Think

Sarvam AI Releases Saaras V4: A Speech-to-Text Model for All 22 Indian Languages and Global English

The study’s results reveal that o1, while showing significant improvements over previous LLMs, still exhibits sensitivity to output probability and task frequency. Across four tasks (shift ciphers, Pig Latin, article swapping, and reversal), o1 demonstrated higher accuracy on examples with high-probability outputs compared to low-probability ones. For instance, in the shift cipher task, o1’s accuracy ranged from 47% for low-probability cases to 92% for high-probability cases. In addition to that,, o1 consumed more tokens when processing low-probability examples, further indicating increased difficulty. Regarding task frequency, o1 initially showed similar performance on common and rare task variants, outperforming other LLMs on rare variants. However, when tested on more challenging versions of sorting and shift cipher tasks, o1 displayed better performance on common variants, suggesting that task frequency effects become apparent when the model is pushed to its limits.

The researchers conclude that o1, despite its significant improvements over previous LLMs, still exhibits sensitivity to output probability and task frequency. This aligns with the teleological perspective, which considers all optimization processes applied to an AI system. O1’s strong performance on algorithmic tasks reflects its explicit optimization for reasoning. However, the observed behavioural patterns suggest that o1 likely underwent substantial next-word prediction training as well. The researchers propose two potential sources for o1’s probability sensitivity: biases in text generation inherent to systems optimized for statistical prediction, and biases in the development of chains of thought favoring high-probability scenarios. To overcome these limitations, the researchers suggest incorporating model components that do not rely on probabilistic judgments, such as modules executing Python code. Ultimately, while o1 represents a significant advancement in AI capabilities, it still retains traces of its autoregressive training, demonstrating that the path to AGI continues to be influenced by the foundational techniques used in language model development.


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter and join our Telegram Channel and LinkedIn Group. If you like our work, you will love our newsletter.. Don’t Forget to join our 50k+ ML SubReddit

Interested in promoting your company, product, service, or event to over 1 Million AI developers and researchers? Let’s collaborate!


Asjad is an intern consultant at Marktechpost. He is persuing B.Tech in mechanical engineering at the Indian Institute of Technology, Kharagpur. Asjad is a Machine learning and deep learning enthusiast who is always researching the applications of machine learning in healthcare.


Credit: Source link

ShareTweetSendSharePin

Related Posts

Your Old GPU Could Be Worth More Than You Think
AI & Technology

Your Old GPU Could Be Worth More Than You Think

September 26, 2026
Sarvam AI Releases Saaras V4: A Speech-to-Text Model for All 22 Indian Languages and Global English
AI & Technology

Sarvam AI Releases Saaras V4: A Speech-to-Text Model for All 22 Indian Languages and Global English

September 26, 2026
Supersonic Labs Releases Julia 1: A 144.3M-Parameter Open Decision Model That Runs on a CPU
AI & Technology

Supersonic Labs Releases Julia 1: A 144.3M-Parameter Open Decision Model That Runs on a CPU

September 26, 2026
This External GPU Uses Wi-Fi To Transform Any Device Into A Gaming Rig
AI & Technology

This External GPU Uses Wi-Fi To Transform Any Device Into A Gaming Rig

September 26, 2026
Next Post
CBIZ’s Rathbun: Tech Job Losses Are a Correction

CBIZ's Rathbun: Tech Job Losses Are a Correction

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
UEFA preparing criminal complaint against FIFA president

UEFA preparing criminal complaint against FIFA president

September 22, 2026
Stop Waiting For The Dip — Rich Ross On What To Buy Now

Stop Waiting For The Dip — Rich Ross On What To Buy Now

September 23, 2026
Big Tech vs Mid-Caps? Josh Wein Goes Rapid-Fire on Stocks

Big Tech vs Mid-Caps? Josh Wein Goes Rapid-Fire on Stocks

September 21, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!