• bitcoinBitcoin(BTC)$85,969.001.15%
  • ethereumEthereum(ETH)$2,752.900.91%
  • tetherTether(USDT)$1.000.01%
  • binancecoinBNB(BNB)$789.630.20%
  • rippleXRP(XRP)$1.565.23%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$117.460.28%
  • tronTRON(TRX)$0.344488-0.04%
  • zcashZcash(ZEC)$1,526.64-0.81%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.01-0.03%
  • HyperliquidHyperliquid(HYPE)$95.290.61%
  • dogecoinDogecoin(DOGE)$0.1003637.81%
  • moneroMonero(XMR)$575.010.73%
  • whitebitWhiteBIT Coin(WBT)$86.520.91%
  • chainlinkChainlink(LINK)$13.020.34%
  • USDSUSDS(USDS)$1.000.00%
  • RainRain(RAIN)$0.013472-4.75%
  • cardanoCardano(ADA)$0.2512223.60%
  • leo-tokenLEO Token(LEO)$8.970.66%
  • stellarStellar(XLM)$0.2144093.17%
  • bitcoin-cashBitcoin Cash(BCH)$315.4917.55%
  • nearNEAR Protocol(NEAR)$4.528.56%
  • uniswapUniswap(UNI)$9.111.61%
  • Ethena USDeEthena USDe(USDE)$1.00-0.02%
  • avalanche-2Avalanche(AVAX)$11.01-2.10%
  • litecoinLitecoin(LTC)$61.46-1.36%
  • CantonCanton(CC)$0.1161031.01%
  • daiDai(DAI)$1.000.02%
  • USD1USD1(USD1)$1.00-0.02%
  • hedera-hashgraphHedera(HBAR)$0.0960725.70%
  • suiSui(SUI)$1.02-1.68%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.430.91%
  • BittensorBittensor(TAO)$317.8911.90%
  • shiba-inuShiba Inu(SHIB)$0.0000065.85%
  • crypto-com-chainCronos(CRO)$0.0668055.85%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • MemeCoreMemeCore(M)$1.33-9.87%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,336.11-0.42%
  • okbOKB(OKB)$122.831.06%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.000.02%
  • BitwayBitway(BTW)$0.87-2.64%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.15-0.04%
  • aaveAave(AAVE)$144.31-0.92%
  • mantleMantle(MNT)$0.664.31%
  • EthenaEthena(ENA)$0.211307-4.92%
  • Pump.funPump.fun(PUMP)$0.0045353.59%
  • OndoOndo(ONDO)$0.433870-3.00%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

aiXplain Researchers Develop Innovative Approaches for Arabic Prompt Instruction Following with LLMs

August 17, 2024
in AI & Technology
Reading Time: 6 mins read
A A
aiXplain Researchers Develop Innovative Approaches for Arabic Prompt Instruction Following with LLMs
ShareShareShareShareShare

Large language models require large datasets of prompts paired with particular user requests and correct responses for training purposes. LLMs require this for human-like text understanding and generation as the answers to various questions. Conversely, unlike other languages, mainly Arabic, immense efforts have been made to develop such datasets in English. This imbalance in data availability between languages severely restricts the applicability of LLMs to non-English-speaking regions and, therefore, denotes a critical need in the NLP domain.

The recent research challenge this paper addresses is the need for good-quality Arabic prompts datasets to train LLMs to perform well in Arabic. These issues must be addressed so LLMs can effectively understand and generate Arabic text. Therefore, they would be contributing less usefulness to the Arabic-speaking users. This is quite relevant because Arabic is spoken by one of the largest numbers in the world. Yet, it lacks sufficient resources for its language, meaning that present AI technologies serve a huge fraction of mankind. Besides the complexity of the Arabic language, due to its rich morphology and huge number of dialects, it takes a lot of work to develop templates that can portray the language the way it should appropriately. Therefore, creating a highly powerful dataset for Arabic is important to upscale the usefulness of the LLM models to a wider audience.

YOU MAY ALSO LIKE

How To Enter VR Mode On Steam

Peloton Has Made A Foldable (Treadmill)

Current prompt dataset generation approaches are mostly oriented towards English and include manual prompt creation or tools generating them based on existing datasets. For example, PromptSource and Super-NaturalInstructions have made millions of prompts available for English-language LLMs. However, these methods have yet to be adapted on any wide scale for other languages, and hence, the resources for training LLMs in non-English languages are considerably lacking. This, coupled with the limited availability of prompt datasets in languages like Arabic, may have hampered the ability of LLMs to excel in these languages, underlining that more focused efforts in dataset creation are necessary.

Researchers from aiXplain Inc. have introduced two innovative methods for creating large-scale Arabic prompt datasets to address this issue. The first method involves translating existing English prompt datasets into Arabic using an automatic translation system, followed by a rigorous quality assessment process. This method relies on state-of-the-art machine translation technologies and quality estimation tools to ensure that the translated prompts maintain high accuracy. By applying these techniques, researchers retained approximately 20% of the translated prompts, resulting in a dataset of around 20 million high-quality Arabic prompts. The second method focuses on creating new prompts directly from existing Arabic NLP datasets. This method uses a prompt sourcing tool to generate prompts for 78 publicly available Arabic datasets, covering tasks such as answering questions, summarization, and detecting hate speech. Over 67.4 million prompts were created through this process, significantly expanding the resources available for training Arabic LLMs.

The translation-based approach follows an end-to-end pipeline in data processing, starting from the tokenization of the English prompts into sentences further translated into Arabic by a neural machine translation model. Then, it performs quality estimation on such translations using a referenceless machine translation quality estimation model, where each sentence will be attributed some quality score. These prompts will be retained only if the set threshold for quality is met; therefore, the final dataset will be highly accurate. Manual verification is conducted on a random sample of prompts to increase the dataset’s quality further. Another approach is to generate prompts directly; PromptSource creates multiple templates for every task in the Arabic datasets. The approach allows the creation of diverse, contextually relevant prompts desirable for training effective language models.

The researchers then used these newly created prompts to fine-tune an open 7 billion parameter LLM, namely the Qwen2 7B model. The fine-tuned model was tested against several benchmarks and significantly improved handling Arabic prompts, outperforming a state-of-the-art 70 billion parameter instruction-tuned model, Llama3 70B. Specifically, the Qwen2 7B model fine-tuned on just 800,000 prompts achieved a ROUGE-L score of 0.184, while the model fine-tuned on 8 million prompts achieved a score of 0.224. These results highlight the effectiveness of the newly developed prompt datasets and demonstrate that fine-tuning with larger datasets leads to better model performance.

In a nutshell, this research speaks about a grave issue: no datasets of Arabic prompts are available to train large language models. The research has opened up the resources for training Arabic LLMs by introducing two new ways to create such datasets. Fine-tuning the Qwen2 7B model using these newly generated prompts produces a model at the top of all other existing models in terms of performance and places a gold standard for Arabic LLMs. It points to the need to develop robust, scalable methods for creating datasets in languages other than English.


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter and join our Telegram Channel and LinkedIn Group. If you like our work, you will love our newsletter..

Don’t Forget to join our 48k+ ML SubReddit

Find Upcoming AI Webinars here



Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.


Credit: Source link

ShareTweetSendSharePin

Related Posts

How To Enter VR Mode On Steam
AI & Technology

How To Enter VR Mode On Steam

September 22, 2026
Peloton Has Made A Foldable (Treadmill)
AI & Technology

Peloton Has Made A Foldable (Treadmill)

September 22, 2026
OpenAI Faces Lawsuit From British Columbia Over Tumbler Ridge Shooting
AI & Technology

OpenAI Faces Lawsuit From British Columbia Over Tumbler Ridge Shooting

September 22, 2026
NVIDIA Introduces SoL-Pi: Auto-Research Loops That Cut Coding Agent Token Traffic by Up to 49%
AI & Technology

NVIDIA Introduces SoL-Pi: Auto-Research Loops That Cut Coding Agent Token Traffic by Up to 49%

September 22, 2026
Next Post
‘We were sure the Russian army would protect us’: fury after Ukrainian incursion into Kursk – The Guardian

‘We were sure the Russian army would protect us’: fury after Ukrainian incursion into Kursk - The Guardian

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Novel climate change lawsuit seeking damages from oil companies heads to Supreme Court

Novel climate change lawsuit seeking damages from oil companies heads to Supreme Court

September 20, 2026
Trump names Adam Telle as acting Army secretary

Trump names Adam Telle as acting Army secretary

September 17, 2026
Raymond Horsch posted ads on Craigslist, preyed on prostitutes: Victim’s cousin – newsnationnow.com

Raymond Horsch posted ads on Craigslist, preyed on prostitutes: Victim’s cousin – newsnationnow.com

September 18, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!