• bitcoinBitcoin(BTC)$77,351.000.16%
  • ethereumEthereum(ETH)$2,539.053.18%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$726.601.69%
  • rippleXRP(XRP)$1.360.71%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$102.812.78%
  • tronTRON(TRX)$0.338410-0.19%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.03-0.02%
  • zcashZcash(ZEC)$1,186.854.77%
  • HyperliquidHyperliquid(HYPE)$80.790.44%
  • dogecoinDogecoin(DOGE)$0.0845750.51%
  • RainRain(RAIN)$0.015594-1.62%
  • moneroMonero(XMR)$520.552.03%
  • USDSUSDS(USDS)$1.000.01%
  • whitebitWhiteBIT Coin(WBT)$80.420.62%
  • chainlinkChainlink(LINK)$11.630.34%
  • leo-tokenLEO Token(LEO)$9.15-0.40%
  • cardanoCardano(ADA)$0.206740-1.11%
  • stellarStellar(XLM)$0.1792590.99%
  • Ethena USDeEthena USDe(USDE)$1.000.02%
  • bitcoin-cashBitcoin Cash(BCH)$229.571.49%
  • daiDai(DAI)$1.000.02%
  • USD1USD1(USD1)$1.000.04%
  • litecoinLitecoin(LTC)$53.792.70%
  • CantonCanton(CC)$0.097921-1.12%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.370.74%
  • uniswapUniswap(UNI)$6.080.62%
  • avalanche-2Avalanche(AVAX)$7.48-1.55%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • hedera-hashgraphHedera(HBAR)$0.074575-1.59%
  • nearNEAR Protocol(NEAR)$2.50-1.13%
  • shiba-inuShiba Inu(SHIB)$0.0000051.58%
  • suiSui(SUI)$0.73-1.39%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.056343-0.35%
  • MemeCoreMemeCore(M)$1.192.54%
  • tether-goldTether Gold(XAUT)$4,346.550.45%
  • Circle USYCCircle USYC(USYC)$1.140.03%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • okbOKB(OKB)$113.451.94%
  • BittensorBittensor(TAO)$236.18-1.79%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.14%
  • aaveAave(AAVE)$125.602.38%
  • mantleMantle(MNT)$0.581.85%
  • pax-goldPAX Gold(PAXG)$4,353.730.59%
  • AsterAster(ASTER)$0.68-3.36%
  • polkadotPolkadot(DOT)$1.05-5.93%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.054531-3.36%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Emergent Ability Unveiled: Can Only Mature AI Like GPT-4 Can Self-Improve? Exploring the Implications of Autonomous Growth in Language Models

June 22, 2023
in AI & Technology
Reading Time: 5 mins read
A A
Emergent Ability Unveiled: Can Only Mature AI Like GPT-4 Can Self-Improve? Exploring the Implications of Autonomous Growth in Language Models
ShareShareShareShareShare

Researchers investigate if, similar to AlphaGo Zero, where AI agents develop themselves by repeatedly engaging in competitive games with clearly laid out rules, many Large Language Models (LLMs) may enhance one another in a negotiating game with little to no human interaction. The results of this study will have far-reaching effects. In contrast to today’s data-hungry LLM training, powerful agents may be built with few human annotations if the agents can progress independently. It also suggests powerful agents with little human supervision, which is problematic. In this study, researchers from the University of Edinburgh and  Allen Institute for AI invite two language models a customer and a seller to haggle over a purchase. 

Figure 1: Setting for our negotiating game. They invite two LLM agents to play the vendor and the buyer in a game of haggling. Their objectives are to sell or purchase the product for more or less money. They ask a third LLM, an AI critic, to give the player we want to get better with after a round. After that, they urge the player to adjust their bargaining tactics in light of the criticism. They continue doing this over several rounds to see whether the models can get better and better. 

The customer wants to pay less for the product, but the seller is requested to sell it for a greater price (Fig. 1). They ask a third language model to take the role of the critic and provide comments to a player once a bargain has been reached. Then, utilizing AI input from the critic LLM, they play the game again and encourage the player to refine their approach. They select the bargaining game because it has explicit rules in print and a specific, quantifiable goal (a lower/higher contract price) for tactical negotiating. Although the game initially appears simple, it calls for non-trivial language model abilities because the model must be able to:

  1. Clearly understand and strictly adhere to the textual rules of the negotiation game.
  2. Correspond to the textual feedback provided by the critic LM and improve based on it iteratively.
  3. Reflect on the strategy and feedback over the long term and improve over multiple rounds. 

In their experiments, only the models get-3.5-turbo, get-4, and Claude-v1.3 meet the requirements of being capable of understanding negotiation rules and strategies and being well-aligned with AI instructions. As a result, not all of the models they considered exhibited all of these abilities (Fig. 2). In the first studies, they also tested more complex textual games, such as board games and text-based role-playing games, but it proved more difficult for the agents to comprehend and adhere to the rules. Their method is known as ICL-AIF (In-Context Learning from AI Feedback). 

🚀 JOIN the fastest ML Subreddit Community
Figure 2: Models are divided into multiple tiers based on the abilities that are necessary in our game (C2 – negotiation, C3 – AI feedback, and C4 – ongoing improvements). Our research reveals that only robust and well-aligned models, such as gpt-4 and claude-v1.3, can benefit from iterative AI input and constantly develop

They leverage the AI critic’s comments and the prior dialogue history rounds as in-context demonstrations. This turns the player’s real development in the previous rounds and the critic’s ideas for changes into the few-shot cues for the subsequent round of bargaining. For two reasons, they use in-context learning: (1) fine-tuning large language models with reinforcement learning is prohibitively expensive, and (2) in-context learning has recently been shown to be closely related to gradient descent, making the conclusions they draw fairly likely to generalize when one fine-tunes the model (if resources permit). 

The reward in Reinforcement Learning from Human Feedback (RLHF) is typically a scalar, but in their ICL-AIF, the feedback is provided in natural language. This is a noteworthy distinction between the two approaches. Instead of relying on human interaction after each round, they examine AI feedback since it is more scalable and can help models progress independently. 

When given feedback while taking on different responsibilities, models respond differently. Improving buyer role models can be more difficult than vendor role models. Even while it is conceivable for powerful agents like get-4 to constantly develop meaningfully utilizing past knowledge and online iterative AI feedback, trying to sell something for more money (or purchase something for less) runs the risk of not making a transaction at all. They also prove that the model can engage in less verbose but more deliberate (and ultimately more successful) bargaining. Overall, they anticipate their work will be an important step towards enhancing language models’ bargaining in a gaming environment with AI feedback. The code is available on GitHub.


Check Out The Paper and Github Link. Don’t forget to join our 24k+ ML SubReddit, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more. If you have any questions regarding the above article or if we missed anything, feel free to email us at [email protected]


Featured Tools From AI Tools Club

🚀 Check Out 100’s AI Tools in AI Tools Club


YOU MAY ALSO LIKE

Lenovo’s Googlebook 15 Seems Decidedly Premium Based On A New Leak

New Images Show A Detailed View Of Meta’s Upcoming Mixed Reality Headset

Aneesh Tickoo is a consulting intern at MarktechPost. He is currently pursuing his undergraduate degree in Data Science and Artificial Intelligence from the Indian Institute of Technology(IIT), Bhilai. He spends most of his time working on projects aimed at harnessing the power of machine learning. His research interest is image processing and is passionate about building solutions around it. He loves to connect with people and collaborate on interesting projects.


Credit: Source link

ShareTweetSendSharePin

Related Posts

Lenovo’s Googlebook 15 Seems Decidedly Premium Based On A New Leak
AI & Technology

Lenovo’s Googlebook 15 Seems Decidedly Premium Based On A New Leak

September 11, 2026
New Images Show A Detailed View Of Meta’s Upcoming Mixed Reality Headset
AI & Technology

New Images Show A Detailed View Of Meta’s Upcoming Mixed Reality Headset

September 11, 2026
Dzmitry Lazerka, Co-Founder of VictoriaMetrics – Interview Series – Unite.AI
AI & Technology

Dzmitry Lazerka, Co-Founder of VictoriaMetrics – Interview Series – Unite.AI

September 11, 2026
Where Should Apple Go After The iPhone Duo? Bring On Smaller And Larger Foldables
AI & Technology

Where Should Apple Go After The iPhone Duo? Bring On Smaller And Larger Foldables

September 11, 2026
Next Post
What to Watch Thursday: Casino Operator Releases Quarterly Earnings

What to Watch Thursday: Casino Operator Releases Quarterly Earnings

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Heroic scenes on highway after car drives off overpass

Heroic scenes on highway after car drives off overpass

September 5, 2026
Sierra Open-Sources Hyper-τ-Bench, a Benchmark for Agent Construction

Sierra Open-Sources Hyper-τ-Bench, a Benchmark for Agent Construction

September 8, 2026
How To Find Your MacBook’s Diagnostic Menu

How To Find Your MacBook’s Diagnostic Menu

September 5, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!