• bitcoinBitcoin(BTC)$76,896.000.92%
  • ethereumEthereum(ETH)$2,460.711.59%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$749.893.70%
  • rippleXRP(XRP)$1.310.59%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$102.553.44%
  • tronTRON(TRX)$0.335546-0.08%
  • zcashZcash(ZEC)$1,487.768.26%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.030.08%
  • HyperliquidHyperliquid(HYPE)$86.439.45%
  • dogecoinDogecoin(DOGE)$0.0827422.42%
  • moneroMonero(XMR)$518.583.41%
  • USDSUSDS(USDS)$1.000.01%
  • whitebitWhiteBIT Coin(WBT)$79.121.06%
  • RainRain(RAIN)$0.012687-1.89%
  • chainlinkChainlink(LINK)$11.554.05%
  • leo-tokenLEO Token(LEO)$8.89-0.53%
  • cardanoCardano(ADA)$0.21641610.87%
  • stellarStellar(XLM)$0.1849961.04%
  • uniswapUniswap(UNI)$7.9218.40%
  • bitcoin-cashBitcoin Cash(BCH)$237.978.22%
  • Ethena USDeEthena USDe(USDE)$1.00-0.02%
  • daiDai(DAI)$1.000.00%
  • USD1USD1(USD1)$1.00-0.01%
  • CantonCanton(CC)$0.10871412.90%
  • nearNEAR Protocol(NEAR)$3.2825.32%
  • litecoinLitecoin(LTC)$54.364.71%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.342.18%
  • avalanche-2Avalanche(AVAX)$7.702.82%
  • hedera-hashgraphHedera(HBAR)$0.0760472.75%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • suiSui(SUI)$0.777.03%
  • shiba-inuShiba Inu(SHIB)$0.0000056.32%
  • crypto-com-chainCronos(CRO)$0.0580702.06%
  • MemeCoreMemeCore(M)$1.2613.59%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • tether-goldTether Gold(XAUT)$4,352.811.54%
  • BittensorBittensor(TAO)$237.416.37%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • okbOKB(OKB)$112.621.56%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.15-0.03%
  • aaveAave(AAVE)$130.898.39%
  • AsterAster(ASTER)$0.744.20%
  • Pump.funPump.fun(PUMP)$0.0040868.43%
  • polkadotPolkadot(DOT)$1.1210.58%
  • BitwayBitway(BTW)$0.70-4.80%
  • pax-goldPAX Gold(PAXG)$4,350.751.48%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

CodeEditorBench: A Machine Learning System for Evaluating the Effectiveness of Large Language Models (LLMs) in Code Editing Activities

April 9, 2024
in AI & Technology
Reading Time: 4 mins read
A A
CodeEditorBench: A Machine Learning System for Evaluating the Effectiveness of Large Language Models (LLMs) in Code Editing Activities
ShareShareShareShareShare

Coding-related jobs have led to the rapid advancement of Large Language Models (LLMs), with a focus on code editing. LLMs created specifically for coding jobs are applied to a variety of activities, including code optimisation and repair. As programming tools, they are becoming more and more popular, but most evaluation techniques concentrate on code production, ignoring the crucial role that code editing plays in software development.

In recent research, a team of researchers from the Multimodal Art Projection Research Community, University of Waterloo, HKUST, University of Manchester, Tongji University, and Vector Institute has introduced CodeEditorBench, an assessment system that has been designed to evaluate LLMs’ effectiveness in a range of code editing activities, such as requirement switching, debugging, translating, and polishing. 

In contrast to other benchmarks that primarily concentrate on code creation, CodeEditorBench emphasises real-world applications and pragmatic elements of software development. The team has selected a variety of coding scenarios and challenges from five distinct sources, covering a broad spectrum of programming languages, degrees of difficulty, and editing assignments. By doing this, they have made sure that the evaluation takes into account the variety and complexity of difficulties found in actual coding environments.

The team has found some intriguing trends in their review, which included 19 distinct LLMs. In the CodeEditorBench framework, closed-source models, specifically, Gemini-Ultra and GPT-4 have demonstrated better performance than open-source models. This emphasises how important model architecture and training data are to deciding performance, particularly when varying prompt sensitivity and problem categories. 

The team has summarized their primary contributions as follows.

  1. The goal of CodeEditorBench is to offer a uniform approach for evaluating LLMs. Tools for additional analyses, training, and visualisation have been included in this framework. To promote more research into LLM features, the team has shared that all evaluation-related data will be openly accessible. To improve the assessment’s comprehensiveness, more evaluation measures will be added in the future. 
  1. The main aim is to map the current state of LLMs. OpenCIDS-33B is the most effective base model available to the public, followed by OpenCI-DS-6.7B and DS-33B-INST. Models like Gemini, GPT, and GLM that are not publicly accessible usually perform better than those that are. OpenCIDS-33B and DS-33B-INST, two instruction-tuned models with over 30 billion parameters, close this performance difference. 
  1. The goal of CodeEditorBench is to draw attention to the shortcomings of LLMs, especially when it comes to rewriting and revising code. Though it performs admirably in three of the four categories, GPT4’s code-polishing abilities are noticeably lacking. In a similar vein, Gemini Ultra is not up to the challenge of changing code requirements. The team has recognized these constraints to tackle these particular issues in LLM training and development.

In conclusion, CodeEditorBench’s main objective is to spur advances in LLMs by providing a strong platform for thoroughly assessing code editing capabilities.


Check out the Paper, Project, and Github. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. Join our Telegram Channel, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 40k+ ML SubReddit

[1/n]
🚀🚀🚀 Excited to share our latest work: “CodeEditorBench:Evaluating Code Editing Capability of Large Language Models”! https://t.co/GckeztzIbT

### 🧐 Highlights of the CodeEditorBench:
> 8K meticulously collected code editing questions from five sources: namely… pic.twitter.com/BUaN6v99BM

— Ge Zhang (@GeZhang86038849) April 5, 2024


YOU MAY ALSO LIKE

eGPUs Do Work, But They Come With Some Notable Limitations

Anthropic Says Claude Leads 26% of Its AI Research and Development – Unite.AI

Tanya Malhotra is a final year undergrad from the University of Petroleum & Energy Studies, Dehradun, pursuing BTech in Computer Science Engineering with a specialization in Artificial Intelligence and Machine Learning.
She is a Data Science enthusiast with good analytical and critical thinking, along with an ardent interest in acquiring new skills, leading groups, and managing work in an organized manner.


🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

eGPUs Do Work, But They Come With Some Notable Limitations
AI & Technology

eGPUs Do Work, But They Come With Some Notable Limitations

September 17, 2026
Anthropic Says Claude Leads 26% of Its AI Research and Development – Unite.AI
AI & Technology

Anthropic Says Claude Leads 26% of Its AI Research and Development – Unite.AI

September 17, 2026
FAA Says Laser Strikes On Aircraft Fell For The Third Consecutive Year
AI & Technology

FAA Says Laser Strikes On Aircraft Fell For The Third Consecutive Year

September 17, 2026
Microsoft Open-Sources TauGrid: A Kubernetes-Native Stack for GPU AI Workloads
AI & Technology

Microsoft Open-Sources TauGrid: A Kubernetes-Native Stack for GPU AI Workloads

September 17, 2026
Next Post
Buckingham Palace issues plea for privacy as Princess Kate treated for cancer

Buckingham Palace issues plea for privacy as Princess Kate treated for cancer

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Fox News parts ways with longtime host Maria Bartiromo

Fox News parts ways with longtime host Maria Bartiromo

September 18, 2026
Oracle’s Cloud Growth; Debate Around AI Risks

Oracle’s Cloud Growth; Debate Around AI Risks

September 12, 2026
OpenAI Launches Misalignment Reporting Framework With Six Incident Reports – Unite.AI

OpenAI Launches Misalignment Reporting Framework With Six Incident Reports – Unite.AI

September 16, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!