• bitcoinBitcoin(BTC)$79,139.000.61%
  • ethereumEthereum(ETH)$2,492.400.38%
  • tetherTether(USDT)$1.000.03%
  • binancecoinBNB(BNB)$750.44-0.71%
  • rippleXRP(XRP)$1.421.81%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$103.990.55%
  • tronTRON(TRX)$0.3388910.27%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.040.00%
  • zcashZcash(ZEC)$1,247.718.97%
  • HyperliquidHyperliquid(HYPE)$86.052.52%
  • dogecoinDogecoin(DOGE)$0.0907330.33%
  • RainRain(RAIN)$0.015977-5.86%
  • USDSUSDS(USDS)$1.000.02%
  • whitebitWhiteBIT Coin(WBT)$81.804.05%
  • moneroMonero(XMR)$497.05-2.93%
  • chainlinkChainlink(LINK)$12.06-4.78%
  • leo-tokenLEO Token(LEO)$9.18-0.05%
  • cardanoCardano(ADA)$0.218864-0.38%
  • stellarStellar(XLM)$0.188497-1.16%
  • bitcoin-cashBitcoin Cash(BCH)$258.130.38%
  • daiDai(DAI)$1.000.01%
  • Ethena USDeEthena USDe(USDE)$1.000.01%
  • USD1USD1(USD1)$1.00-0.01%
  • CantonCanton(CC)$0.1070691.99%
  • litecoinLitecoin(LTC)$54.01-2.77%
  • uniswapUniswap(UNI)$6.66-5.12%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.400.09%
  • hedera-hashgraphHedera(HBAR)$0.078530-2.33%
  • avalanche-2Avalanche(AVAX)$7.93-2.27%
  • suiSui(SUI)$0.81-1.34%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • nearNEAR Protocol(NEAR)$2.549.44%
  • shiba-inuShiba Inu(SHIB)$0.000005-1.02%
  • crypto-com-chainCronos(CRO)$0.0603562.95%
  • paypal-usdPayPal USD(PYUSD)$1.000.02%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • MemeCoreMemeCore(M)$1.190.64%
  • tether-goldTether Gold(XAUT)$4,395.090.09%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BittensorBittensor(TAO)$260.741.98%
  • Ripple USDRipple USD(RLUSD)$1.000.02%
  • okbOKB(OKB)$114.32-1.21%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.10%
  • mantleMantle(MNT)$0.642.29%
  • AsterAster(ASTER)$0.76-1.06%
  • polkadotPolkadot(DOT)$1.187.93%
  • aaveAave(AAVE)$129.13-1.44%
  • pax-goldPAX Gold(PAXG)$4,399.800.09%
  • Pump.funPump.fun(PUMP)$0.0044852.61%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

A New AI Research from China Introduces GLM-130B: A Bilingual (English and Chinese) Pre-Trained Language Model with 130B Parameters

November 3, 2023
in AI & Technology
Reading Time: 4 mins read
A A
A New AI Research from China Introduces GLM-130B: A Bilingual (English and Chinese) Pre-Trained Language Model with 130B Parameters
ShareShareShareShareShare

In recent times, the zero-shot and few-shot capabilities of Large Language Models (LLMs) have increased significantly, with those with over 100B parameters giving state-of-the-art performance on various benchmarks. Such an advancement also presents a critical challenge with respect to LLMs, i.e., transparency. Very limited knowledge about these large-scale models and their training process is available to the public, and releasing this information would facilitate the training of high-quality LLMs of this scale.

A group of researchers from Tsinghua University and Zhipu.AI have released GLM-130B, which is an open-source bilingual (English and Chinese) pre-trained language model with 130B parameters. The researchers in this paper have demonstrated the training process of the model, including the ways the process could be optimized, in an attempt to open-source a model at par with GPT-3, having parameters in the scale of 100B. Additionally, the researchers have shared both the successful and failed aspects of the training process.

GLM-130B uses a bidirectional General Language Model (GLM) as its base. The architecture uses autoregressive blank infilling as its training objective, which allows for a better understanding of contexts as compared to GPT-style models. GLM-130B is able to outperform both GPT-3 and PaLM 540B on zero-shot LAMBADA by achieving a zero-shot accuracy of 80.2%.

The authors of this paper experimented with different Layer Normalization (LN) techniques in order to stabilize the training process of GLM-130B. Existing practices such as Pre-LN, Post-LN, and Sandwich-LN were ineffective, but Post-LN initialized with DeepNorm showed promising results. The pre-training data of the model consists of more than 2TB of English and Chinese text corpora extracted from online forums, encyclopedias, etc., to form a well-balanced dataset.

As mentioned earlier, GLM-130B achieves a record accuracy on the LAMBADA dataset. On the Pile test set, which consists of a series of benchmarks for language modelling, the GLM model’s performance was at par with GPT-3 and Jurassic-1 models. The model also performs well on the MMLU benchmark, with its few-shot performance as good as GPT-3. 

Additionally, on the BIG-bench benchmark, GLM-130B was able to outperform both GPT-3 and PaLM in zero-shot settings. Even though the model gave significant performances, the researchers noticed that its performance growth with respect to few-shot samples is not as great as GPT-3’s. They hypothesize that it is due to multiple reasons, such as the model’s bidirectional nature, the limitation of a dataset that is at par with PaLM in terms of quality and diversity, etc.

The researchers also tested the zero-shot performance of the model on Chinese benchmarks. They concluded that GLM-130B not only outperformed ERNIE Titan 3.0 across more than ten tasks but also performed at least 260% better than the same on two abstractive MRC datasets. This may be due to the fact that the pre-training objective of GLM included autoregressive blank infilling that is similar to abstractive MRC.

In conclusion, the GLM-130B is a powerful, open-source, bilingual pre-trained language model that performs at the level of GPT-3 and PaLM across different benchmarks and even outperforms them in some of the tasks. Apart from its performance, what sets this model apart is the transparency of its development. The researchers have made the training process of the model public, along with their experiences of both success and failure. This approach reflects their commitment to fostering open and inclusive research within the field of LLMs.


Check out the Paper and Github. All Credit For This Research Goes To the Researchers on This Project. Also, don’t forget to join our 32k+ ML SubReddit, 40k+ Facebook Community, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.

If you like our work, you will love our newsletter..

We are also on Telegram and WhatsApp.


YOU MAY ALSO LIKE

Will We See The Foldable iPhone Ultra At The ‘Surprise And Shine’ Keynote Today?

Gradium Launches Voice Design: Write a Prompt, Get a Brand New Synthetic Voice in Seconds

I am a Civil Engineering Graduate (2022) from Jamia Millia Islamia, New Delhi, and I have a keen interest in Data Science, especially Neural Networks and their application in various areas.


🔥 Meet Retouch4me: A Family of Artificial Intelligence-Powered Plug-Ins for Photography Retouching

Credit: Source link

ShareTweetSendSharePin

Related Posts

Will We See The Foldable iPhone Ultra At The ‘Surprise And Shine’ Keynote Today?
AI & Technology

Will We See The Foldable iPhone Ultra At The ‘Surprise And Shine’ Keynote Today?

September 9, 2026
Gradium Launches Voice Design: Write a Prompt, Get a Brand New Synthetic Voice in Seconds
AI & Technology

Gradium Launches Voice Design: Write a Prompt, Get a Brand New Synthetic Voice in Seconds

September 9, 2026
Meta Introduces Muse, a Personal AI Agent That Runs on Its Own Dedicated Secure Cloud Computer
AI & Technology

Meta Introduces Muse, a Personal AI Agent That Runs on Its Own Dedicated Secure Cloud Computer

September 9, 2026
OpenAI Says Internal AI System Resolved the Navier–Stokes Problem – Unite.AI
AI & Technology

OpenAI Says Internal AI System Resolved the Navier–Stokes Problem – Unite.AI

September 9, 2026
Next Post
Uvalde Mass Shooting Footage Published By News Outlets

Uvalde Mass Shooting Footage Published By News Outlets

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Severe storms slam the Midwest, causing devastation

Severe storms slam the Midwest, causing devastation

September 4, 2026
Negotiating Your Monthly Bills Takes One Phone Call and Often Works

Negotiating Your Monthly Bills Takes One Phone Call and Often Works

September 2, 2026
‘Welcome to the AGI era’: OpenAI launches GPT-6 Astra

‘Welcome to the AGI era’: OpenAI launches GPT-6 Astra

September 3, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!