• bitcoinBitcoin(BTC)$77,686.00-2.18%
  • ethereumEthereum(ETH)$2,456.10-1.82%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$715.25-4.65%
  • rippleXRP(XRP)$1.37-3.89%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$101.06-3.21%
  • tronTRON(TRX)$0.3402910.34%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.030.00%
  • zcashZcash(ZEC)$1,213.06-4.20%
  • HyperliquidHyperliquid(HYPE)$82.40-5.05%
  • dogecoinDogecoin(DOGE)$0.085168-6.45%
  • RainRain(RAIN)$0.016086-1.44%
  • USDSUSDS(USDS)$1.000.00%
  • moneroMonero(XMR)$506.401.74%
  • whitebitWhiteBIT Coin(WBT)$80.27-2.13%
  • chainlinkChainlink(LINK)$11.78-3.19%
  • leo-tokenLEO Token(LEO)$9.230.33%
  • cardanoCardano(ADA)$0.212468-3.84%
  • stellarStellar(XLM)$0.179315-5.06%
  • bitcoin-cashBitcoin Cash(BCH)$243.68-6.03%
  • daiDai(DAI)$1.00-0.01%
  • Ethena USDeEthena USDe(USDE)$1.00-0.02%
  • USD1USD1(USD1)$1.000.00%
  • litecoinLitecoin(LTC)$52.17-3.75%
  • CantonCanton(CC)$0.101093-4.98%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.37-2.40%
  • uniswapUniswap(UNI)$6.00-10.80%
  • avalanche-2Avalanche(AVAX)$7.73-2.93%
  • hedera-hashgraphHedera(HBAR)$0.075982-3.41%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • nearNEAR Protocol(NEAR)$2.40-8.33%
  • suiSui(SUI)$0.76-6.49%
  • shiba-inuShiba Inu(SHIB)$0.000005-5.44%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.056453-5.88%
  • MemeCoreMemeCore(M)$1.201.22%
  • tether-goldTether Gold(XAUT)$4,370.92-0.60%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.00-0.02%
  • BittensorBittensor(TAO)$250.76-6.17%
  • okbOKB(OKB)$111.70-2.49%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.00%
  • mantleMantle(MNT)$0.59-7.83%
  • AsterAster(ASTER)$0.71-5.69%
  • aaveAave(AAVE)$122.78-4.91%
  • pax-goldPAX Gold(PAXG)$4,371.73-0.65%
  • polkadotPolkadot(DOT)$1.10-6.53%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0561920.71%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Google AI Introduces AltUp (Alternating Updates): An Artificial Intelligence Method that Takes Advantage of Increasing Scale in Transformer Networks without Increasing the Computation Cost

November 13, 2023
in AI & Technology
Reading Time: 5 mins read
A A
Google AI Introduces AltUp (Alternating Updates): An Artificial Intelligence Method that Takes Advantage of Increasing Scale in Transformer Networks without Increasing the Computation Cost
ShareShareShareShareShare

In deep learning, Transformer neural networks have garnered significant attention for their effectiveness in various domains, especially in natural language processing and emerging applications like computer vision, robotics, and autonomous driving. However, while enhancing performance, the ever-increasing scale of these models brings about a substantial rise in compute cost and inference latency. The fundamental challenge lies in leveraging the advantages of larger models without incurring impractical computational burdens.

The current landscape of deep learning models, particularly Transformers, showcases remarkable progress across diverse domains. Nevertheless, the scalability of these models often needs to be improved due to the escalating computational requirements. Prior efforts, exemplified by sparse mixture-of-experts models like Switch Transformer, Expert Choice, and V-MoE, have predominantly focused on efficiently scaling up network parameters, mitigating the increased compute per input. However, a research gap exists concerning the scaling up of the token representation dimension itself. Enter AltUp is a novel method introduced to address this gap.

AltUp stands out by providing a method to augment token representation without amplifying the computational overhead. This method ingeniously partitions a widened representation vector into equal-sized blocks, processing only one block at each layer. The crux of AltUp’s efficacy lies in its prediction-correction mechanism, enabling the inference of outputs for the non-processed blocks. By maintaining the model dimension and sidestepping the quadratic increase in computation associated with straightforward expansion, AltUp emerges as a promising solution to the computational challenges posed by larger Transformer networks.

AltUp’s mechanics delve into the intricacies of token embeddings and how they can be widened without triggering a surge in computational complexity. The method involves:

  • Invoking a 1x width transformer layer for one of the blocks.
  • Termed the “activated” block.
  • Concurrently employing a lightweight predictor.

This predictor computes a weighted combination of all input blocks, and the predicted values, along with the computed value of the activated block, undergo correction through a lightweight corrector. This correction mechanism facilitates the update of inactivated blocks based on the activated ones. Importantly, both prediction and correction steps involve minimal vector additions and multiplications, significantly faster than a conventional transformer layer.

The evaluation of AltUp on T5 models across benchmark language tasks demonstrates its consistent ability to outperform dense models at the same accuracy. Notably, a T5 Large model augmented with AltUp achieves notable speedups of 27%, 39%, 87%, and 29% on GLUE, SuperGLUE, SQuAD, and Trivia-QA benchmarks, respectively. AltUp’s relative performance improvements become more pronounced when applied to larger models, underscoring its scalability and enhanced efficacy as model size increases.

In conclusion, AltUp emerges as a noteworthy solution to the long-standing challenge of efficiently scaling up Transformer neural networks. Its ability to augment token representation without a proportional increase in computational cost holds significant promise for various applications. The innovative approach of AltUp, characterized by its partitioning and prediction-correction mechanism, offers a pragmatic way to harness the benefits of larger models without succumbing to impractical computational demands.

The researchers’ extension of AltUp, known as Recycled-AltUp, further showcases the adaptability of the proposed method. Recycled-AltUp, by replicating embeddings instead of widening the initial token embeddings, demonstrates strict improvements in pre-training performance without introducing perceptible slowdown. This dual-pronged approach, coupled with AltUp’s seamless integration with other techniques like MoE, exemplifies its versatility and opens avenues for future research in exploring the dynamics of training and model performance.

AltUp signifies a breakthrough in the quest for efficient scaling of Transformer networks, presenting a compelling solution to the trade-off between model size and computational efficiency. As outlined in this paper, the research team’s contributions mark a significant step towards making large-scale Transformer models more accessible and practical for a myriad of applications.


Check out the Paper and Google Article. All credit for this research goes to the researchers of this project. Also, don’t forget to join our 32k+ ML SubReddit, 41k+ Facebook Community, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.

If you like our work, you will love our newsletter..

We are also on Telegram and WhatsApp.


YOU MAY ALSO LIKE

IBM and NASA Open-Source Lunar Foundation Model With SomBench Dataset – Unite.AI

NASA And IBM Made An AI Model For Exploring The Moon

Madhur Garg is a consulting intern at MarktechPost. He is currently pursuing his B.Tech in Civil and Environmental Engineering from the Indian Institute of Technology (IIT), Patna. He shares a strong passion for Machine Learning and enjoys exploring the latest advancements in technologies and their practical applications. With a keen interest in artificial intelligence and its diverse applications, Madhur is determined to contribute to the field of Data Science and leverage its potential impact in various industries.


🔥 Meet Retouch4me: A Family of Artificial Intelligence-Powered Plug-Ins for Photography Retouching

Credit: Source link

ShareTweetSendSharePin

Related Posts

IBM and NASA Open-Source Lunar Foundation Model With SomBench Dataset – Unite.AI
AI & Technology

IBM and NASA Open-Source Lunar Foundation Model With SomBench Dataset – Unite.AI

September 10, 2026
NASA And IBM Made An AI Model For Exploring The Moon
AI & Technology

NASA And IBM Made An AI Model For Exploring The Moon

September 10, 2026
Fujitsu Signs New Palantir AIP Agreement, Becomes Global FDE Partner – Unite.AI
AI & Technology

Fujitsu Signs New Palantir AIP Agreement, Becomes Global FDE Partner – Unite.AI

September 10, 2026
AppleCare One Now Has A  Tier Per Month For Families
AI & Technology

AppleCare One Now Has A $50 Tier Per Month For Families

September 10, 2026
Next Post
U.S. ‘Will Make Every Effort’ To Rescue Captured Americans In Ukraine, Former Amb. Says

U.S. ‘Will Make Every Effort’ To Rescue Captured Americans In Ukraine, Former Amb. Says

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
AI News: The Most Insane Week So Far This Year!

AI News: The Most Insane Week So Far This Year!

September 4, 2026
Australian mother welcomes identical quadruplet girls

Australian mother welcomes identical quadruplet girls

September 7, 2026
Get Mortgage Pre-Approval Before You Start House Hunting

Get Mortgage Pre-Approval Before You Start House Hunting

September 4, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!