• bitcoinBitcoin(BTC)$76,708.00-0.80%
  • ethereumEthereum(ETH)$2,477.26-2.15%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$714.97-3.20%
  • rippleXRP(XRP)$1.34-2.34%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$99.60-2.36%
  • tronTRON(TRX)$0.3401430.17%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.00-1.59%
  • zcashZcash(ZEC)$1,093.42-4.99%
  • HyperliquidHyperliquid(HYPE)$77.25-3.13%
  • dogecoinDogecoin(DOGE)$0.083390-1.81%
  • RainRain(RAIN)$0.0153511.45%
  • moneroMonero(XMR)$537.360.66%
  • USDSUSDS(USDS)$1.00-0.02%
  • whitebitWhiteBIT Coin(WBT)$79.58-0.98%
  • chainlinkChainlink(LINK)$11.32-1.88%
  • leo-tokenLEO Token(LEO)$9.05-0.61%
  • cardanoCardano(ADA)$0.203903-2.11%
  • stellarStellar(XLM)$0.177826-2.15%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • daiDai(DAI)$1.000.01%
  • bitcoin-cashBitcoin Cash(BCH)$222.71-3.67%
  • USD1USD1(USD1)$1.00-0.01%
  • litecoinLitecoin(LTC)$53.58-1.04%
  • uniswapUniswap(UNI)$6.20-2.45%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.35-1.67%
  • CantonCanton(CC)$0.094664-4.25%
  • Global DollarGlobal Dollar(USDG)$1.00-0.02%
  • hedera-hashgraphHedera(HBAR)$0.0750860.85%
  • avalanche-2Avalanche(AVAX)$7.31-1.71%
  • shiba-inuShiba Inu(SHIB)$0.000005-1.90%
  • nearNEAR Protocol(NEAR)$2.30-2.90%
  • suiSui(SUI)$0.71-2.21%
  • crypto-com-chainCronos(CRO)$0.0583791.14%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,344.48-0.12%
  • MemeCoreMemeCore(M)$1.15-2.60%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • okbOKB(OKB)$113.16-0.74%
  • BittensorBittensor(TAO)$232.92-1.48%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.05%
  • aaveAave(AAVE)$124.38-1.84%
  • pax-goldPAX Gold(PAXG)$4,350.39-0.11%
  • AsterAster(ASTER)$0.69-0.02%
  • mantleMantle(MNT)$0.56-3.18%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.056750-0.54%
  • polkadotPolkadot(DOT)$1.00-3.80%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Can We Transfer the Capabilities of LLMs like LLaMA from English to Non-English Languages? A Deep Dive into Multilingual Model Proficiency

January 6, 2024
in AI & Technology
Reading Time: 4 mins read
A A
Can We Transfer the Capabilities of LLMs like LLaMA from English to Non-English Languages? A Deep Dive into Multilingual Model Proficiency
ShareShareShareShareShare

Significant achievements have been made in LLMs, exemplified by ChatGPT, excelling in complex language processing tasks. But most mainstream LLMs like LLaMA are pre-trained on English-dominant corpus. Another example is LaMDA, proposed by Google, which is pre-trained on text containing over 90% English. This limits the performance of LLMs in other non-English languages, which is a matter of concern for non-English users.

Recent strides in LLMs like ChatGPT, PaLM, and LLaMA showcase advanced reasoning, planning, and experiential learning capabilities. While many LLMs comprehend diverse languages, imbalanced language resources pose challenges. BLOOM’s pretraining on 46 languages lacks diversity, and LLaMA faces difficulties with non-English languages. Investigations into vocabulary extension and transfer processes reveal efficient language transfer at minimal cost. 

The researchers at the School of Computer Science, Fudan University, have focused on effectively transferring language generation capabilities and following instructions in non-English. To address this, they have analyzed the impact of key factors such as vocabulary extension, further pretraining, and instruction tuning on transfer. Evaluation involves four standardized benchmarks. 

The research explores transferring language generation and instruction-following capabilities to non-English languages using LLaMA. Due to its rich linguistic resources, it employs Chinese as the starting point, extending findings to over ten low-resource languages. Models include LLaMA, LLaMA2, Chinese LLaMA, Chinese LLaMA2, and Open Chinese LLaMA, each with different pretraining scales. Evaluation involves benchmarks like LLM-Eval, C-Eval, MMLU, AGI-Eval, and GAOKAO-Bench. Response quality is assessed based on accuracy, fluency, informativeness, logical coherence, and harmlessness. The study achieves state-of-the-art performance with minimal pretraining data, providing insights for non-English LLM development.

The study investigates language transfer to non-English languages using LLaMA, focusing on vocabulary extension, training scale impact, and multilingual proficiency. Surprisingly, extending the vocabulary diminishes performance in Chinese. While increased pretraining scale initially improves response quality, it plateaus, emphasizing language generation over knowledge acquisition. English proficiency suffers with exclusive Chinese training. Evaluations across 13 low-resource languages show SFT data boost response quality, with Arabic, Indonesian, and Vietnamese excelling. Code-switching samples suggest LLaMA learns cross-lingual semantic alignment during pretraining, enhancing transferability. The study emphasizes nuanced approaches for effective non-English LLM development.  

Table 1: Evaluation results of model response quality for 13 low-resource languages on the LLM-Eval. ACC., F., LC., H., INFO., and AVG. Respectively denote accuracy, fluency, logical coherence, harmlessness, informativeness, and average.
Researchers have focused on effectively transferring language generation capabilities and following instructions to a non-English language. Specifically, they have conducted a comprehensive empirical study to analyze the necessity of vocabulary extension and the required training scale for effective transfer. They found that vocabulary extension is unnecessary and that comparable transfer performance to state-of-the-art models can be achieved with less than 1% of the further pretraining data. Similar results are observed from the extension experiments on the 13 low-resource languages. 


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. Join our 35k+ ML SubReddit, 41k+ Facebook Community, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..


YOU MAY ALSO LIKE

AWS Introduces Pizza Bot: An Open Source Inbox for Background AI Agents

Implementation of Machine Learning Workflows with NVIDIA cuML, RAPIDS, GPU Benchmarking, Explainability, Clustering, and Model Inference

Nikhil is an intern consultant at Marktechpost. He is pursuing an integrated dual degree in Materials at the Indian Institute of Technology, Kharagpur. Nikhil is an AI/ML enthusiast who is always researching applications in fields like biomaterials and biomedical science. With a strong background in Material Science, he is exploring new advancements and creating opportunities to contribute.


🐝 Get stunning professional headshots effortlessly with Aragon- TRY IT NOW!.


Credit: Source link

ShareTweetSendSharePin

Related Posts

AWS Introduces Pizza Bot: An Open Source Inbox for Background AI Agents
AI & Technology

AWS Introduces Pizza Bot: An Open Source Inbox for Background AI Agents

September 13, 2026
Implementation of Machine Learning Workflows with NVIDIA cuML, RAPIDS, GPU Benchmarking, Explainability, Clustering, and Model Inference
AI & Technology

Implementation of Machine Learning Workflows with NVIDIA cuML, RAPIDS, GPU Benchmarking, Explainability, Clustering, and Model Inference

September 13, 2026
Hyundai Motor Group Puts Data Flywheel Into Full Operation – Unite.AI
AI & Technology

Hyundai Motor Group Puts Data Flywheel Into Full Operation – Unite.AI

September 13, 2026
What Is The Difference Between A Dead Pixel And A Stuck Pixel?
AI & Technology

What Is The Difference Between A Dead Pixel And A Stuck Pixel?

September 13, 2026
Next Post
Meet Fusilli: A Python Library for Multi-Modal Data Fusion in Machine Learning

Meet Fusilli: A Python Library for Multi-Modal Data Fusion in Machine Learning

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Husband Only Contributes 17% To Our Household

Husband Only Contributes 17% To Our Household

September 12, 2026
Aya Gold & Silver: Updated PEA Brings Good News To An Already Solid Growth Stock (AYA)

Aya Gold & Silver: Updated PEA Brings Good News To An Already Solid Growth Stock (AYA)

September 11, 2026
Your Largest Bottleneck May Be Your Most Self-Assured AI Champion – Unite.AI

Your Largest Bottleneck May Be Your Most Self-Assured AI Champion – Unite.AI

September 8, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!