• bitcoinBitcoin(BTC)$77,235.000.06%
  • ethereumEthereum(ETH)$2,524.260.49%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$727.180.04%
  • rippleXRP(XRP)$1.370.77%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$101.78-0.63%
  • tronTRON(TRX)$0.3398660.46%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.00-3.08%
  • zcashZcash(ZEC)$1,124.41-3.29%
  • HyperliquidHyperliquid(HYPE)$79.800.48%
  • dogecoinDogecoin(DOGE)$0.0847940.59%
  • RainRain(RAIN)$0.0157522.27%
  • moneroMonero(XMR)$539.243.69%
  • USDSUSDS(USDS)$1.00-0.01%
  • whitebitWhiteBIT Coin(WBT)$80.260.17%
  • chainlinkChainlink(LINK)$11.50-0.26%
  • leo-tokenLEO Token(LEO)$9.14-0.15%
  • cardanoCardano(ADA)$0.2072240.53%
  • stellarStellar(XLM)$0.1800321.06%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • daiDai(DAI)$1.00-0.02%
  • bitcoin-cashBitcoin Cash(BCH)$226.29-0.64%
  • USD1USD1(USD1)$1.000.00%
  • litecoinLitecoin(LTC)$53.600.57%
  • uniswapUniswap(UNI)$6.356.24%
  • CantonCanton(CC)$0.0980550.82%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.381.61%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • avalanche-2Avalanche(AVAX)$7.40-0.79%
  • hedera-hashgraphHedera(HBAR)$0.0746370.34%
  • shiba-inuShiba Inu(SHIB)$0.0000052.36%
  • nearNEAR Protocol(NEAR)$2.370.25%
  • suiSui(SUI)$0.72-0.02%
  • crypto-com-chainCronos(CRO)$0.0603357.01%
  • paypal-usdPayPal USD(PYUSD)$1.000.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • MemeCoreMemeCore(M)$1.18-0.49%
  • tether-goldTether Gold(XAUT)$4,350.010.00%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • okbOKB(OKB)$114.451.13%
  • BittensorBittensor(TAO)$232.83-1.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.07%
  • aaveAave(AAVE)$125.450.65%
  • pax-goldPAX Gold(PAXG)$4,354.62-0.01%
  • AsterAster(ASTER)$0.690.70%
  • mantleMantle(MNT)$0.56-3.58%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0572113.34%
  • polkadotPolkadot(DOT)$1.02-1.99%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Line Open-Sources ‘japanese-large-lm’: A Japanese Language Model With 3.6 Billion Parameters

August 20, 2023
in AI & Technology
Reading Time: 4 mins read
A A
Line Open-Sources ‘japanese-large-lm’: A Japanese Language Model With 3.6 Billion Parameters
ShareShareShareShareShare

Since November 2020, LINE has embarked on a transformative journey of research and development to create and harness the power of an advanced large-scale language model tailored specifically for the Japanese language. As a significant milestone in this journey, LINE’s Massive LM development unit has announced the release of their Japanese language models, “Japanese-large-lm,” as open-source software (OSS). This release is poised to significantly impact both the research community and businesses seeking to leverage cutting-edge language models.

These language models come in two variants—the 3.6 billion (3.6B) parameter model and the 1.7 billion (1.7B) parameter model, aptly referred to as the 3.6B model and 1.7B model. By unveiling these models and sharing their comprehensive insights into language model construction, LINE aims to provide a glimpse into the intricacies of their approach and contribute to the advancement of the field.

The 1.7B and 3.6B models are accessible via the HuggingFace Hub(1.7B model, 3.6B model), offering seamless integration into various projects through the popular transformers library. Licensing these models under Apache License 2.0 ensures that a wide spectrum of users, including researchers and commercial entities, can leverage their capabilities for diverse applications.

A cornerstone in developing any high-performing language model lies in utilizing an extensive and high-quality training dataset. LINE tapped into its proprietary Japanese web corpus—a repository enriched with diverse textual data to achieve this. However, the challenge web-derived content poses is its inherent noise, including source code and non-Japanese sentences. LINE’s response was to employ meticulous filtering processes powered by the HojiChar OSS library. These processes were instrumental in distilling a large-scale, high-quality dataset, forming the bedrock of the models’ robustness.

Efficiency in model training was a key consideration, and LINE rose to the occasion by implementing innovative techniques like 3D Parallelism and Activation Checkpointing. These advancements facilitated the efficient assimilation of voluminous data, effectively pushing the boundaries of computational capability. Astonishingly, the 1.7B model was developed using just 4000 GPU hours on an A100 80GB GPU—a testament to the efficacy of their learning approach.

Notably, the development trajectory of this Japanese language model diverged from that of HyperCLOVA. Built along a distinct development line, meticulously overseen by LINE’s dedicated Massive LM development unit, this model is a testament to LINE’s commitment to crafting exceptional pre-trained models for the Japanese language. Their overarching goal remains steadfast—integrating insights and lessons from their extensive experience with large-scale language models.

LINE delved into perplexity scores (PPL) and accuracy rates for question-answering and reading comprehension tasks to assess the models’ efficacy. PPL provides insight into the model’s predictive capabilities, while accuracy rates offer tangible performance measures. The results were promising, with LINE’s models showcasing competitive performance across various tasks, rivaling established models in the field.

Underpinning their success was a series of invaluable tips for effective large-scale language model training. These encompass considerations for fine-tuning, the hyperparameter Adam’s beta2, optimal learning rates, and applying a judicious learning rate scheduler. By delving into these technical intricacies, LINE has developed potent models and shared insights that benefit the wider community.

In conclusion, LINE’s release of the 1.7B and 3.6B Japanese language models marks a significant stride in natural language processing. Their commitment to releasing tuned models in the future underscores their dedication to enhancing the capabilities of language models. As LINE continues to make advancements, the global community eagerly anticipates the enduring impact of their ongoing contributions.


Check out the Reference Article. All Credit For This Research Goes To the Researchers on This Project. Also, don’t forget to join our 29k+ ML SubReddit, 40k+ Facebook Community, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.

If you like our work, please follow us on Twitter


YOU MAY ALSO LIKE

Blizzard Is Reviving StarCraft As An Open-World Shooter, But It’ll Be A Long Wait

Diablo V Is Coming Out In Spring 2029

Niharika is a Technical consulting intern at Marktechpost. She is a third year undergraduate, currently pursuing her B.Tech from Indian Institute of Technology(IIT), Kharagpur. She is a highly enthusiastic individual with a keen interest in Machine learning, Data science and AI and an avid reader of the latest developments in these fields.


🔥 Use SQL to predict the future (Sponsored)


Credit: Source link

ShareTweetSendSharePin

Related Posts

Blizzard Is Reviving StarCraft As An Open-World Shooter, But It’ll Be A Long Wait
AI & Technology

Blizzard Is Reviving StarCraft As An Open-World Shooter, But It’ll Be A Long Wait

September 12, 2026
Diablo V Is Coming Out In Spring 2029
AI & Technology

Diablo V Is Coming Out In Spring 2029

September 12, 2026
Fly Language Model (FLM) Wires the Full Fruit Fly Connectome Into a Frozen 1.2B LLM, and Its Own Controls Show the Wiring Does Not Help
AI & Technology

Fly Language Model (FLM) Wires the Full Fruit Fly Connectome Into a Frozen 1.2B LLM, and Its Own Controls Show the Wiring Does Not Help

September 12, 2026
Altman Says OpenAI Will Match Anthropic’s Embedded Evaluator Pledge – Unite.AI
AI & Technology

Altman Says OpenAI Will Match Anthropic’s Embedded Evaluator Pledge – Unite.AI

September 12, 2026
Next Post
TSLY: Best Move Is To Avoid This ETF

TSLY: Best Move Is To Avoid This ETF

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
What is ‘popcorn brain’ and how to help it

What is ‘popcorn brain’ and how to help it

September 6, 2026
Your 401k Contribution Limit Resets Every January

Your 401k Contribution Limit Resets Every January

September 8, 2026
Suspect in Missouri jumps off bridge while fleeing deputies

Suspect in Missouri jumps off bridge while fleeing deputies

September 7, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!