• bitcoinBitcoin(BTC)$84,372.000.52%
  • ethereumEthereum(ETH)$2,695.970.31%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$771.46-0.36%
  • rippleXRP(XRP)$1.52-2.79%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$120.59-0.03%
  • tronTRON(TRX)$0.332980-1.36%
  • zcashZcash(ZEC)$1,640.567.03%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.063.41%
  • HyperliquidHyperliquid(HYPE)$93.221.46%
  • dogecoinDogecoin(DOGE)$0.096099-1.82%
  • chainlinkChainlink(LINK)$14.100.45%
  • moneroMonero(XMR)$560.741.07%
  • whitebitWhiteBIT Coin(WBT)$84.090.33%
  • USDSUSDS(USDS)$1.000.00%
  • cardanoCardano(ADA)$0.251756-1.66%
  • RainRain(RAIN)$0.01277417.95%
  • leo-tokenLEO Token(LEO)$9.041.72%
  • stellarStellar(XLM)$0.214476-2.00%
  • bitcoin-cashBitcoin Cash(BCH)$332.07-1.58%
  • nearNEAR Protocol(NEAR)$5.043.22%
  • uniswapUniswap(UNI)$9.721.88%
  • litecoinLitecoin(LTC)$71.37-0.52%
  • CantonCanton(CC)$0.1357252.32%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • suiSui(SUI)$1.170.50%
  • avalanche-2Avalanche(AVAX)$10.740.87%
  • daiDai(DAI)$1.00-0.02%
  • USD1USD1(USD1)$1.000.01%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.557.44%
  • hedera-hashgraphHedera(HBAR)$0.093270-1.02%
  • BittensorBittensor(TAO)$320.552.78%
  • shiba-inuShiba Inu(SHIB)$0.000006-0.59%
  • crypto-com-chainCronos(CRO)$0.066308-0.18%
  • Global DollarGlobal Dollar(USDG)$1.00-0.02%
  • BitwayBitway(BTW)$1.0415.86%
  • MemeCoreMemeCore(M)$1.230.60%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • EthenaEthena(ENA)$0.266681-0.52%
  • tether-goldTether Gold(XAUT)$4,278.13-0.11%
  • OndoOndo(ONDO)$0.53-1.10%
  • okbOKB(OKB)$120.97-0.62%
  • quant-networkQuant(QNT)$172.8871.47%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • aaveAave(AAVE)$154.900.26%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.15-0.08%
  • mantleMantle(MNT)$0.680.18%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Do Models like GPT-4 Behave Safely When Given the Ability to Act?: This AI Paper Introduces MACHIAVELLI Benchmark to Improve Machine Ethics and Build Safer Adaptive Agents

July 18, 2023
in AI & Technology
Reading Time: 4 mins read
A A
Do Models like GPT-4 Behave Safely When Given the Ability to Act?: This AI Paper Introduces MACHIAVELLI Benchmark to Improve Machine Ethics and Build Safer Adaptive Agents
ShareShareShareShareShare

Natural language processing is one area where AI systems are making rapid strides, and it is important that the models need to be rigorously tested and guided toward safer behavior to reduce deployment risks. Prior evaluation metrics for such sophisticated systems focused on measuring language comprehension or reasoning in vacuums. But now, models are being taught for actual, interactive work. This means that benchmarks need to evaluate how models perform in social settings.

Interactive agents can be put through their paces in text-based games. Agents need planning abilities and the ability to grasp the natural language to progress in these games. Agents’ immoral tendencies should be considered alongside their technical talents while setting benchmarks.

A new work by the University of California, Center For AI Safety, Carnegie Mellon University, and Yale University proposes the Measuring Agents’ Competence & Harmfulness In A Vast Environment of Long-horizon Language Interactions (MACHIAVELLI) benchmark. MACHIAVELLI is an advancement in evaluating an agent’s capacity for planning in naturalistic social settings. The setting is inspired by text-based Choose Your Own Adventure games available at choiceofgames.com, which actual humans developed. These games feature high-level decisions while giving agents realistic objectives while abstracting away low-level environment interactions.

🚀 Automate labeling to save time with smart tools & model predictions

The environment reports the degree to which agent acts are dishonest, lower utility, and seek power, among other behavioral qualities, to keep tabs on unethical behavior. The team achieves this by following the below-mentioned steps:

  1. Operationalizing these behaviors as mathematical formulas
  2. Densely annotating social notions in the games, such as characters’ wellbeing
  3. Using the annotations and formulas to produce a numerical score for each behavior. 

They demonstrate empirically that GPT-4 (OpenAI, 2023) is more effective at collecting annotations than human annotators.

Artificial intelligence agents face the same internal conflict as humans do. Like language models trained for next-token prediction often produce toxic text, artificial agents trained for goal optimization often exhibit immoral and power-seeking behaviors. Amorally trained agents may develop Machiavellian strategies for maximizing their rewards at the expense of others and the environment. By encouraging agents to act morally, this trade-off can be improved.

The team discovers that moral training (nudging the agent to be more ethical) decreases the incidence of harmful activity for language-model agents. Furthermore, behavioral regularization restricts undesirable behavior in both agents without substantially decreasing reward. This work contributes to the development of trustworthy sequential decision-makers.

The researchers try techniques like an artificial conscience and ethics prompts to control agents. Agents can be guided to display less Machiavellian behavior, although much progress remains possible. They advocate for more research into these trade-offs and emphasize expanding the Pareto frontier rather than chasing after limited rewards.


Check out the Paper. All Credit For This Research Goes To the Researchers on This Project. Also, don’t forget to join our 18k+ ML SubReddit, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.

🚀 Check Out 100’s AI Tools in AI Tools Club


YOU MAY ALSO LIKE

Why We Won’t Know How Visible The iPhone Duo’s Crease Is For A Long Time

Your Old GPU Could Be Worth More Than You Think

Tanushree Shenwai is a consulting intern at MarktechPost. She is currently pursuing her B.Tech from the Indian Institute of Technology(IIT), Bhubaneswar. She is a Data Science enthusiast and has a keen interest in the scope of application of artificial intelligence in various fields. She is passionate about exploring the new advancements in technologies and their real-life application.


🔥 StoryBird.ai just dropped some amazing features. Generate an illustrated story from a prompt. Check it out here. (Sponsored)

Credit: Source link

ShareTweetSendSharePin

Related Posts

Why We Won’t Know How Visible The iPhone Duo’s Crease Is For A Long Time
AI & Technology

Why We Won’t Know How Visible The iPhone Duo’s Crease Is For A Long Time

September 27, 2026
Your Old GPU Could Be Worth More Than You Think
AI & Technology

Your Old GPU Could Be Worth More Than You Think

September 26, 2026
Sarvam AI Releases Saaras V4: A Speech-to-Text Model for All 22 Indian Languages and Global English
AI & Technology

Sarvam AI Releases Saaras V4: A Speech-to-Text Model for All 22 Indian Languages and Global English

September 26, 2026
Supersonic Labs Releases Julia 1: A 144.3M-Parameter Open Decision Model That Runs on a CPU
AI & Technology

Supersonic Labs Releases Julia 1: A 144.3M-Parameter Open Decision Model That Runs on a CPU

September 26, 2026
Next Post
Cable TV Users Can Choose Channels With New Verizon Service

Cable TV Users Can Choose Channels With New Verizon Service

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Titanic-themed slide for kids makes headlines

Titanic-themed slide for kids makes headlines

September 27, 2026
Officials describe Nevada’s Hawk Fire as human-caused

Officials describe Nevada’s Hawk Fire as human-caused

September 25, 2026
Wall Street Brunch: U.S.-China Summit In Spotlight (NYSEARCA:SPY)

Wall Street Brunch: U.S.-China Summit In Spotlight (NYSEARCA:SPY)

September 20, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!