• bitcoinBitcoin(BTC)$76,720.00-0.76%
  • ethereumEthereum(ETH)$2,481.92-1.89%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$715.73-2.60%
  • rippleXRP(XRP)$1.34-1.79%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$99.79-2.07%
  • tronTRON(TRX)$0.3399600.11%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.00-1.59%
  • zcashZcash(ZEC)$1,080.71-5.85%
  • HyperliquidHyperliquid(HYPE)$77.91-1.67%
  • dogecoinDogecoin(DOGE)$0.083517-1.51%
  • RainRain(RAIN)$0.0153651.68%
  • moneroMonero(XMR)$537.870.30%
  • USDSUSDS(USDS)$1.00-0.01%
  • whitebitWhiteBIT Coin(WBT)$79.58-0.89%
  • chainlinkChainlink(LINK)$11.34-1.71%
  • leo-tokenLEO Token(LEO)$9.06-0.58%
  • cardanoCardano(ADA)$0.204949-1.68%
  • stellarStellar(XLM)$0.178056-1.56%
  • Ethena USDeEthena USDe(USDE)$1.00-0.02%
  • daiDai(DAI)$1.000.01%
  • bitcoin-cashBitcoin Cash(BCH)$222.82-3.76%
  • USD1USD1(USD1)$1.00-0.01%
  • litecoinLitecoin(LTC)$53.69-0.66%
  • uniswapUniswap(UNI)$6.23-1.61%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.35-1.64%
  • CantonCanton(CC)$0.095147-4.01%
  • Global DollarGlobal Dollar(USDG)$1.00-0.03%
  • hedera-hashgraphHedera(HBAR)$0.0751571.04%
  • avalanche-2Avalanche(AVAX)$7.33-1.53%
  • shiba-inuShiba Inu(SHIB)$0.000005-1.46%
  • nearNEAR Protocol(NEAR)$2.32-2.00%
  • suiSui(SUI)$0.71-1.84%
  • crypto-com-chainCronos(CRO)$0.0583011.26%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,346.65-0.08%
  • MemeCoreMemeCore(M)$1.15-2.42%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • okbOKB(OKB)$113.74-0.17%
  • BittensorBittensor(TAO)$233.33-0.77%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.09%
  • aaveAave(AAVE)$124.52-1.40%
  • pax-goldPAX Gold(PAXG)$4,353.21-0.06%
  • AsterAster(ASTER)$0.690.46%
  • mantleMantle(MNT)$0.55-3.94%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0570180.58%
  • polkadotPolkadot(DOT)$1.00-3.57%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Meet Astraios: An AI Model Suite Consisting of 28 Instruction-Tuned OctoCoder Across Scales and PEFT Methods

January 6, 2024
in AI & Technology
Reading Time: 4 mins read
A A
Meet Astraios: An AI Model Suite Consisting of 28 Instruction-Tuned OctoCoder Across Scales and PEFT Methods
ShareShareShareShareShare

Recent research highlights the success of Large Language Models (LLMs) trained on Code, excelling at diverse software engineering tasks. These models fall into three primary paradigms: (i) Code LLMs specialized in code completion, (ii) Task-specific Code LLMs fine-tuned for individual tasks, and (iii) Instruction-tuned Code LLMs adept at adhering to human instructions and demonstrating robustness in handling new tasks. Recent instruction-tuned Code LLMs such as WizardCoder and OctoCoder have notably achieved cutting-edge performance across various tasks without requiring task-specific fine-tuning.

To delve deeper into the identified opportunities, Monash University and ServiceNow Research researchers introduce ASTRAIOS, a collection comprising 28 instruction-tuned Code LLMs. These models undergo fine-tuning using seven tuning methods based on the base models of StarCoder, specifically, models sized at 1B, 3B, 7B, and 16 B. They conduct instruction tuning on these models using the CommitPackFT dataset from OctoPack to ensure a balanced enhancement of their downstream capabilities.

They employ PEFT configurations aligned with Hugging Face’s recommended practices and integrate selected PEFT methods from recent frameworks. They initially scrutinize the scalability of different tuning methods by evaluating cross-entropy loss during instruction tuning. This assessment specifically focuses on assessing model size and training time scales.

Their primary evaluation revolves around five representative code-related tasks: clone detection, defect detection, code synthesis, code repair, and code explanation. Additionally, they conduct further analysis of the tuning methods, examining model robustness and code security. This evaluation involves assessing the models’ ability to generate Code based on perturbed examples and determining the potential vulnerabilities in the generated Code.

Larger PEFT Code LLMs excel in code generation tasks but do not demonstrate similar advantages in code comprehension tasks like clone detection and defect detection. As model size increases, task performance in generation improves but raises concerns regarding susceptibility to adversarial examples and a bias toward insecure Code. 

Their study delves into the relationship among updated parameters, cross-entropy loss, and task performance. They ascertain that the final loss of smaller PEFT models can be used to predict that of larger ones. Moreover, a strong correlation exists between the last loss and overall performance in downstream tasks.

The correlation between model loss and updated parameters is inconsistent across different model sizes in our analysis. However, a noteworthy discovery is the uniformity in relative loss performance across various model sizes when comparing other tuning methods. This consistency implies that the enhancements attained by each tuning method are comparable, irrespective of the model’s scale. Consequently, the loss observed in smaller models tuned using different methods can serve as a valuable indicator for predicting the performance of larger models.


Check out the Paper and Github. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. Join our 35k+ ML SubReddit, 41k+ Facebook Community, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..


YOU MAY ALSO LIKE

AWS Introduces Pizza Bot: An Open Source Inbox for Background AI Agents

Implementation of Machine Learning Workflows with NVIDIA cuML, RAPIDS, GPU Benchmarking, Explainability, Clustering, and Model Inference

Arshad is an intern at MarktechPost. He is currently pursuing his Int. MSc Physics from the Indian Institute of Technology Kharagpur. Understanding things to the fundamental level leads to new discoveries which lead to advancement in technology. He is passionate about understanding the nature fundamentally with the help of tools like mathematical models, ML models and AI.


🐝 Get stunning professional headshots effortlessly with Aragon- TRY IT NOW!.


Credit: Source link

ShareTweetSendSharePin

Related Posts

AWS Introduces Pizza Bot: An Open Source Inbox for Background AI Agents
AI & Technology

AWS Introduces Pizza Bot: An Open Source Inbox for Background AI Agents

September 13, 2026
Implementation of Machine Learning Workflows with NVIDIA cuML, RAPIDS, GPU Benchmarking, Explainability, Clustering, and Model Inference
AI & Technology

Implementation of Machine Learning Workflows with NVIDIA cuML, RAPIDS, GPU Benchmarking, Explainability, Clustering, and Model Inference

September 13, 2026
Hyundai Motor Group Puts Data Flywheel Into Full Operation – Unite.AI
AI & Technology

Hyundai Motor Group Puts Data Flywheel Into Full Operation – Unite.AI

September 13, 2026
What Is The Difference Between A Dead Pixel And A Stuck Pixel?
AI & Technology

What Is The Difference Between A Dead Pixel And A Stuck Pixel?

September 13, 2026
Next Post
Pfizer Stock: Paying Too Much To Boost Revenues (NYSE:PFE)

Pfizer Stock: Paying Too Much To Boost Revenues (NYSE:PFE)

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
What Trump and Iran are signaling about war as American death toll rises

What Trump and Iran are signaling about war as American death toll rises

September 7, 2026
An Attractive ‘Mid-Size’ Foldable With Powerful Specs

An Attractive ‘Mid-Size’ Foldable With Powerful Specs

September 8, 2026
Baseten Adds DeepSeek-V4.1-Flash to Model APIs With 1M-Token Context – Unite.AI

Baseten Adds DeepSeek-V4.1-Flash to Model APIs With 1M-Token Context – Unite.AI

September 11, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!