• bitcoinBitcoin(BTC)$84,321.000.31%
  • ethereumEthereum(ETH)$2,692.42-0.04%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$773.03-0.42%
  • rippleXRP(XRP)$1.53-2.68%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$121.48-0.53%
  • tronTRON(TRX)$0.334153-1.12%
  • zcashZcash(ZEC)$1,659.746.60%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.062.99%
  • HyperliquidHyperliquid(HYPE)$92.06-0.64%
  • dogecoinDogecoin(DOGE)$0.096592-2.62%
  • chainlinkChainlink(LINK)$14.131.56%
  • moneroMonero(XMR)$558.940.18%
  • whitebitWhiteBIT Coin(WBT)$84.110.17%
  • USDSUSDS(USDS)$1.000.00%
  • cardanoCardano(ADA)$0.253121-2.02%
  • RainRain(RAIN)$0.01288910.67%
  • leo-tokenLEO Token(LEO)$8.961.27%
  • stellarStellar(XLM)$0.216876-1.41%
  • bitcoin-cashBitcoin Cash(BCH)$334.88-2.72%
  • nearNEAR Protocol(NEAR)$4.980.07%
  • uniswapUniswap(UNI)$9.731.45%
  • litecoinLitecoin(LTC)$72.27-0.33%
  • CantonCanton(CC)$0.1362425.12%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • avalanche-2Avalanche(AVAX)$10.862.31%
  • suiSui(SUI)$1.16-3.73%
  • daiDai(DAI)$1.000.01%
  • USD1USD1(USD1)$1.000.01%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.577.80%
  • hedera-hashgraphHedera(HBAR)$0.093432-1.93%
  • BittensorBittensor(TAO)$320.930.16%
  • shiba-inuShiba Inu(SHIB)$0.000006-0.02%
  • crypto-com-chainCronos(CRO)$0.0682723.88%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • MemeCoreMemeCore(M)$1.220.99%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • EthenaEthena(ENA)$0.2702011.14%
  • BitwayBitway(BTW)$1.00-24.28%
  • tether-goldTether Gold(XAUT)$4,279.82-0.12%
  • OndoOndo(ONDO)$0.54-3.01%
  • okbOKB(OKB)$120.950.15%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • aaveAave(AAVE)$155.311.11%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.15-0.06%
  • mantleMantle(MNT)$0.692.60%
  • quant-networkQuant(QNT)$147.1047.82%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

From Prediction to Reasoning: Evaluating o1’s Impact on LLM Probabilistic Biases

October 8, 2024
in AI & Technology
Reading Time: 4 mins read
A A
From Prediction to Reasoning: Evaluating o1’s Impact on LLM Probabilistic Biases
ShareShareShareShareShare

Large language models (LLMs) have gained significant attention in recent years, but understanding their capabilities and limitations remains a challenge. Researchers are trying to develop methodologies to reason about the strengths and weaknesses of AI systems, particularly LLMs. The current approaches often lack a systematic framework for predicting and analyzing these systems’ behaviours. This has led to difficulties in anticipating how LLMs will perform various tasks, especially those that differ from their primary training objective. The challenge lies in bridging the gap between the AI system’s training process and its observed performance on diverse tasks, necessitating a more comprehensive analytical approach.

In this study, researchers from the Wu Tsai Institute, Yale University, OpenAI, Princeton University, Roundtable, and Princeton University have focused on analyzing OpenAI’s new system, o1, which was explicitly optimized for reasoning tasks, to determine if it exhibits the same “embers of autoregression” observed in previous LLMs. The researchers apply the teleological perspective, which considers the pressures shaping AI systems, to predict and evaluate o1’s performance. This approach examines whether o1’s departure from pure next-word prediction training mitigates limitations associated with that objective. The study compares o1’s performance to other LLMs on various tasks, assessing its sensitivity to output probability and task frequency. In addition to that, the researchers introduce a robust metric—token count during answer generation—to quantify task difficulty. This comprehensive analysis aims to reveal whether o1 represents a significant advancement or still retains behavioural patterns linked to next-word prediction training.

YOU MAY ALSO LIKE

Supersonic Labs Releases Julia 1: A 144.3M-Parameter Open Decision Model That Runs on a CPU

This External GPU Uses Wi-Fi To Transform Any Device Into A Gaming Rig

The study’s results reveal that o1, while showing significant improvements over previous LLMs, still exhibits sensitivity to output probability and task frequency. Across four tasks (shift ciphers, Pig Latin, article swapping, and reversal), o1 demonstrated higher accuracy on examples with high-probability outputs compared to low-probability ones. For instance, in the shift cipher task, o1’s accuracy ranged from 47% for low-probability cases to 92% for high-probability cases. In addition to that,, o1 consumed more tokens when processing low-probability examples, further indicating increased difficulty. Regarding task frequency, o1 initially showed similar performance on common and rare task variants, outperforming other LLMs on rare variants. However, when tested on more challenging versions of sorting and shift cipher tasks, o1 displayed better performance on common variants, suggesting that task frequency effects become apparent when the model is pushed to its limits.

The researchers conclude that o1, despite its significant improvements over previous LLMs, still exhibits sensitivity to output probability and task frequency. This aligns with the teleological perspective, which considers all optimization processes applied to an AI system. O1’s strong performance on algorithmic tasks reflects its explicit optimization for reasoning. However, the observed behavioural patterns suggest that o1 likely underwent substantial next-word prediction training as well. The researchers propose two potential sources for o1’s probability sensitivity: biases in text generation inherent to systems optimized for statistical prediction, and biases in the development of chains of thought favoring high-probability scenarios. To overcome these limitations, the researchers suggest incorporating model components that do not rely on probabilistic judgments, such as modules executing Python code. Ultimately, while o1 represents a significant advancement in AI capabilities, it still retains traces of its autoregressive training, demonstrating that the path to AGI continues to be influenced by the foundational techniques used in language model development.


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter and join our Telegram Channel and LinkedIn Group. If you like our work, you will love our newsletter.. Don’t Forget to join our 50k+ ML SubReddit

Interested in promoting your company, product, service, or event to over 1 Million AI developers and researchers? Let’s collaborate!


Asjad is an intern consultant at Marktechpost. He is persuing B.Tech in mechanical engineering at the Indian Institute of Technology, Kharagpur. Asjad is a Machine learning and deep learning enthusiast who is always researching the applications of machine learning in healthcare.


Credit: Source link

ShareTweetSendSharePin

Related Posts

Supersonic Labs Releases Julia 1: A 144.3M-Parameter Open Decision Model That Runs on a CPU
AI & Technology

Supersonic Labs Releases Julia 1: A 144.3M-Parameter Open Decision Model That Runs on a CPU

September 26, 2026
This External GPU Uses Wi-Fi To Transform Any Device Into A Gaming Rig
AI & Technology

This External GPU Uses Wi-Fi To Transform Any Device Into A Gaming Rig

September 26, 2026
You Can Use Your Old Laptop To Make A Smart Home Hub
AI & Technology

You Can Use Your Old Laptop To Make A Smart Home Hub

September 26, 2026
TikTok Will Pay Alabama 0 Million To Settle Social Media Addiction Lawsuit
AI & Technology

TikTok Will Pay Alabama $100 Million To Settle Social Media Addiction Lawsuit

September 26, 2026
Next Post
CBIZ’s Rathbun: Tech Job Losses Are a Correction

CBIZ's Rathbun: Tech Job Losses Are a Correction

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Federal judge orders DHS to pause construction of border wall in Big Bend National Park

Federal judge orders DHS to pause construction of border wall in Big Bend National Park

September 21, 2026
Humanoid robot beats Usain Bolt’s 100-meter record

Humanoid robot beats Usain Bolt’s 100-meter record

September 25, 2026
Trump welcomes Xi for state dinner as tech bosses and other guests arrive at White House – BBC

Trump welcomes Xi for state dinner as tech bosses and other guests arrive at White House – BBC

September 25, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!