• bitcoinBitcoin(BTC)$78,937.000.71%
  • ethereumEthereum(ETH)$2,491.590.68%
  • tetherTether(USDT)$1.000.02%
  • binancecoinBNB(BNB)$749.72-0.52%
  • rippleXRP(XRP)$1.421.91%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$103.700.73%
  • tronTRON(TRX)$0.3388960.12%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.040.00%
  • zcashZcash(ZEC)$1,270.589.61%
  • HyperliquidHyperliquid(HYPE)$86.053.40%
  • dogecoinDogecoin(DOGE)$0.0905211.00%
  • RainRain(RAIN)$0.015973-5.31%
  • USDSUSDS(USDS)$1.000.01%
  • whitebitWhiteBIT Coin(WBT)$81.654.19%
  • moneroMonero(XMR)$493.47-3.43%
  • chainlinkChainlink(LINK)$12.04-3.46%
  • leo-tokenLEO Token(LEO)$9.180.02%
  • cardanoCardano(ADA)$0.2186680.52%
  • stellarStellar(XLM)$0.187758-0.60%
  • bitcoin-cashBitcoin Cash(BCH)$257.570.66%
  • daiDai(DAI)$1.000.00%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • USD1USD1(USD1)$1.00-0.02%
  • CantonCanton(CC)$0.1059581.22%
  • litecoinLitecoin(LTC)$53.80-2.94%
  • uniswapUniswap(UNI)$6.68-4.01%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.400.50%
  • hedera-hashgraphHedera(HBAR)$0.078451-2.15%
  • avalanche-2Avalanche(AVAX)$7.94-1.29%
  • Global DollarGlobal Dollar(USDG)$1.00-0.02%
  • suiSui(SUI)$0.81-1.17%
  • nearNEAR Protocol(NEAR)$2.5310.33%
  • shiba-inuShiba Inu(SHIB)$0.000005-0.14%
  • crypto-com-chainCronos(CRO)$0.0603691.99%
  • paypal-usdPayPal USD(PYUSD)$1.000.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,405.060.17%
  • MemeCoreMemeCore(M)$1.180.28%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BittensorBittensor(TAO)$263.013.29%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • okbOKB(OKB)$114.40-0.64%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.33%
  • mantleMantle(MNT)$0.642.17%
  • AsterAster(ASTER)$0.75-0.97%
  • polkadotPolkadot(DOT)$1.178.04%
  • aaveAave(AAVE)$129.21-0.64%
  • pax-goldPAX Gold(PAXG)$4,409.370.17%
  • OndoOndo(ONDO)$0.375057-0.60%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

New transformer architecture can make language models faster and resource-efficient

December 1, 2023
in AI & Technology
Reading Time: 5 mins read
A A
New transformer architecture can make language models faster and resource-efficient
ShareShareShareShareShare

Are you ready to bring more awareness to your brand? Consider becoming a sponsor for The AI Impact Tour. Learn more about the opportunities here.


Large language models like ChatGPT and Llama-2 are notorious for their extensive memory and computational demands, making them costly to run. Trimming even a small fraction of their size can lead to significant cost reductions. 

YOU MAY ALSO LIKE

Will We See The Foldable iPhone Ultra At The ‘Surprise And Shine’ Keynote Today?

Gradium Launches Voice Design: Write a Prompt, Get a Brand New Synthetic Voice in Seconds

To address this issue, researchers at ETH Zurich have unveiled a revised version of the transformer, the deep learning architecture underlying language models. The new design reduces the size of the transformer considerably while preserving accuracy and increasing inference speed, making it a promising architecture for more efficient language models.

Transformer blocks

Language models operate on a foundation of transformer blocks, uniform units adept at parsing sequential data, such as text passages.

Classic transformer block (source: arxiv.org)

The transformer block specializes in processing sequential data, such as a passage of text. Within each block, there are two key sub-blocks: the “attention mechanism” and the multi-layer perceptron (MLP). The attention mechanism acts like a highlighter, selectively focusing on different parts of the input data (like words in a sentence) to capture their context and importance relative to each other. This helps the model determine how the words in a sentence relate, even if they are far apart. 

VB Event

The AI Impact Tour

Connect with the enterprise AI community at VentureBeat’s AI Impact Tour coming to a city near you!

 

Learn More

After the attention mechanism has done its work, the MLP, a mini neural network, further refines and processes the highlighted information, helping to distill the data into a more sophisticated representation that captures complex relationships.

Beyond these core components, transformer blocks are equipped with additional features such as “residual connections” and “normalization layers.” These components accelerate learning and mitigate issues common in deep neural networks.

As transformer blocks stack to constitute a language model, their capacity to discern complex relationships in training data grows, enabling the sophisticated tasks performed by contemporary language models. Despite the transformative impact of these models, the fundamental design of the transformer block has remained largely unchanged since its creation. 

Making the transformer more efficient

“Given the exorbitant cost of training and deploying large transformer models nowadays, any efficiency gains in the training and inference pipelines for the transformer architecture represent significant potential savings,” write the ETH Zurich researchers. “Simplifying the transformer block by removing non-essential components both reduces the parameter count and increases throughput in our models.”

The team’s experiments demonstrate that paring down the transformer block does not compromise training speed or performance on downstream tasks. Standard transformer models feature multiple attention heads, each with its own set of key (K), query (Q), and value (V) parameters, which together map the interplay among input tokens. The researchers discovered that they could eliminate the V parameters and the subsequent projection layer that synthesizes the values for the MLP block, without losing efficacy.

Moreover, they removed the skip connections, which traditionally help avert the “vanishing gradients” issue in deep learning models. Vanishing gradients make training deep networks difficult, as the gradient becomes too small to effect significant learning in the earlier layers.

New transformer block, with V and projection parameters and skip connections removed (source: arxiv.org)

They also redesigned the transformer block to process attention heads and the MLP concurrently rather than sequentially. This parallel processing marks a departure from the conventional architecture.

To compensate for the reduction in parameters, the researchers adjusted other non-learnable parameters, refined the training methodology, and implemented architectural tweaks. These changes collectively maintain the model’s learning capabilities, despite the leaner structure.

Testing the new transformer block

The ETH Zurich team evaluated their compact transformer block across language models of varying depths. Their findings were significant: they managed to shrink the conventional transformer’s size by approximately 16% without sacrificing accuracy, and they achieved faster inference times. To put that in perspective, applying this new architecture to a large model like GPT-3, with its 175 billion parameters, could result in a memory saving of about 50 GB.

“Our simplified models are able to not only train faster but also to utilize the extra capacity that more depth provides,” the researchers write. While their technique has proven effective on smaller scales, its application to larger models remains untested. The potential for further enhancements, such as tailoring AI processors to this streamlined architecture, could amplify its impact.

“We believe our work can lead to simpler architectures being used in practice, thereby helping to bridge the gap between theory and practice in deep learning, and reducing the cost of large transformer models,” the researchers write.

VentureBeat’s mission is to be a digital town square for technical decision-makers to gain knowledge about transformative enterprise technology and transact. Discover our Briefings.

Credit: Source link

ShareTweetSendSharePin

Related Posts

Will We See The Foldable iPhone Ultra At The ‘Surprise And Shine’ Keynote Today?
AI & Technology

Will We See The Foldable iPhone Ultra At The ‘Surprise And Shine’ Keynote Today?

September 9, 2026
Gradium Launches Voice Design: Write a Prompt, Get a Brand New Synthetic Voice in Seconds
AI & Technology

Gradium Launches Voice Design: Write a Prompt, Get a Brand New Synthetic Voice in Seconds

September 9, 2026
Meta Introduces Muse, a Personal AI Agent That Runs on Its Own Dedicated Secure Cloud Computer
AI & Technology

Meta Introduces Muse, a Personal AI Agent That Runs on Its Own Dedicated Secure Cloud Computer

September 9, 2026
OpenAI Says Internal AI System Resolved the Navier–Stokes Problem – Unite.AI
AI & Technology

OpenAI Says Internal AI System Resolved the Navier–Stokes Problem – Unite.AI

September 9, 2026
Next Post
NASA’s Psyche mission launches to explore metal-filled asteroid

NASA’s Psyche mission launches to explore metal-filled asteroid

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Video of discredited bite mark evidence helps free man from death row

Video of discredited bite mark evidence helps free man from death row

September 4, 2026
Pros And Cons Of Using A Chromebook As Your Everyday PC

Pros And Cons Of Using A Chromebook As Your Everyday PC

September 8, 2026
How To Reset The Camera Settings On Your iPhone

How To Reset The Camera Settings On Your iPhone

September 8, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!