• bitcoinBitcoin(BTC)$77,213.001.26%
  • ethereumEthereum(ETH)$2,473.231.83%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$748.943.44%
  • rippleXRP(XRP)$1.321.57%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$104.675.14%
  • tronTRON(TRX)$0.3358970.09%
  • zcashZcash(ZEC)$1,511.2810.94%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.03-0.15%
  • HyperliquidHyperliquid(HYPE)$86.559.26%
  • dogecoinDogecoin(DOGE)$0.0841074.05%
  • moneroMonero(XMR)$514.572.84%
  • USDSUSDS(USDS)$1.000.02%
  • whitebitWhiteBIT Coin(WBT)$79.531.46%
  • RainRain(RAIN)$0.012674-1.52%
  • chainlinkChainlink(LINK)$11.745.86%
  • leo-tokenLEO Token(LEO)$8.89-0.57%
  • cardanoCardano(ADA)$0.2138899.29%
  • stellarStellar(XLM)$0.1877523.10%
  • uniswapUniswap(UNI)$8.6327.63%
  • bitcoin-cashBitcoin Cash(BCH)$244.8511.22%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • nearNEAR Protocol(NEAR)$3.4930.87%
  • daiDai(DAI)$1.00-0.01%
  • USD1USD1(USD1)$1.000.00%
  • CantonCanton(CC)$0.10704710.17%
  • litecoinLitecoin(LTC)$54.514.94%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.353.10%
  • avalanche-2Avalanche(AVAX)$7.864.41%
  • hedera-hashgraphHedera(HBAR)$0.0770884.20%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • suiSui(SUI)$0.788.61%
  • shiba-inuShiba Inu(SHIB)$0.0000057.13%
  • crypto-com-chainCronos(CRO)$0.0587240.57%
  • MemeCoreMemeCore(M)$1.2713.65%
  • paypal-usdPayPal USD(PYUSD)$1.000.01%
  • tether-goldTether Gold(XAUT)$4,353.861.45%
  • BittensorBittensor(TAO)$239.006.85%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • okbOKB(OKB)$113.862.39%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.29%
  • aaveAave(AAVE)$133.1210.03%
  • AsterAster(ASTER)$0.753.82%
  • polkadotPolkadot(DOT)$1.1311.33%
  • mantleMantle(MNT)$0.583.65%
  • Pump.funPump.fun(PUMP)$0.0040877.08%
  • pax-goldPAX Gold(PAXG)$4,352.051.37%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Consistency Large Language Models (CLLMs): A New Family of LLMs Specialized for the Jacobi Decoding Method for Latency Reduction

May 17, 2024
in AI & Technology
Reading Time: 5 mins read
A A
Consistency Large Language Models (CLLMs): A New Family of LLMs Specialized for the Jacobi Decoding Method for Latency Reduction
ShareShareShareShareShare

Large language models (LLMs), including GPT-4, LLaMA, and PaLM are pushing the boundaries of artificial intelligence. The inference latency of LLMs plays an important role because of LLMs integration in various applications, ensuring a positive user experience and high service quality. However, the LLM service operates within an AR paradigm, generating one token at a time because the attention mechanism relies on previous token states to generate the next token. To produce a longer response, a forward pass is executed using LLMs equivalent to the number of tokens generated, leading to high latency.

The efficient LLM Inference method is divided into two streams, a method that needs additional training and another that does not need it. Researchers explored this method due to the high AR inference cost for LLMs, mostly focused on increasing the AR decoding process. Another existing method is LLM Distillation, where the Knowledge distillation (KD) technique is used to create small models and replace the functionality of larger ones. However, traditional KD methods are not effective for LLMs. So, KD is used for autoregressive LLMs to minimize the reverse KL divergence between student and teacher models through student-driven decoding.   

Researchers from Shanghai Jiao University and the University of California proposed Consistency Large Language Models (CLLMs), a new family of LLMs specialized for the Jacobi decoding method for latency reduction. CLLM was compared with traditional methods such as speculative decoding and Medusa for adjusting auxiliary model components and didn’t use extra memory for this task, unlike others. When CLLMs are trained on ∼ 1M tokens for LLaMA-7B, it becomes 3.4 times faster on the Spider dataset showing that the cost of fine-tuning is moderate for this method. Two main factors for this speed-up are fast forwarding and stationary tokens. 

In fast forwarding, correct predictions are done in a single forward pass for multiple consecutive tokens whereas, stationary tokens show correct prediction with no change through subsequent iterations despite being preceded by inaccurate tokens. In target LLMs and CLLMs, when fast-forwarded and stationary counts are compared across all four datasets (in Table 3), there is an improvement of 2.0x to 6.8x in both token counts. Also, for both the token counts, such improvement in domain-specific datasets is better than in open-domain datasets profiled on MT-bench. This helps distinctive collocations and easy syntactical structures like blank space, newline tokens, and repetitive special characters in specialized domains like coding.

Researchers carried out experiments to evaluate the performance and inference speedup of CLLMs across multiple tasks such as comparing the stat-of-the-art (SOTA) baselines on the three domain-specific tasks and the open-domain profiled on MT-bench. CLLMs show outstanding performance on various benchmarks, e.g. they can achieve 2.4× to 3.4× speedup using Jacobi decoding with nearly no loss in accuracy on domain-specific benchmarks like GSM8K, CodeSearchNet Python, and Spider. CLLMs can achieve  2.4× speedup on ShareGPT with SOTA performance, with a 6.4 score on the open-domain benchmark MT-bench.

In conclusion, researchers introduced CLLMs, a new family of LLMs that excel in efficient parallel decoding and are designed in a way that they can improve the efficiency of Jacobi decoding. Additional architecture designs or managing two different models in a single system are complex and complexity is reduced with the help of CLLMs because this method is directly adapted from a target pre-trained LLM. Besides, fast-forwarded and stationary token counts are compared across four datasets, showing an enhancement of 2.0x to 6.8x In target LLMs and CLLMs. 


Check out the Paper and Project. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. Join our Telegram Channel, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 42k+ ML SubReddit


YOU MAY ALSO LIKE

eGPUs Do Work, But They Come With Some Notable Limitations

Google’s Revamped CC Is An AI Agent For Families And Groups

Sajjad Ansari is a final year undergraduate from IIT Kharagpur. As a Tech enthusiast, he delves into the practical applications of AI with a focus on understanding the impact of AI technologies and their real-world implications. He aims to articulate complex AI concepts in a clear and accessible manner.


🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

eGPUs Do Work, But They Come With Some Notable Limitations
AI & Technology

eGPUs Do Work, But They Come With Some Notable Limitations

September 17, 2026
Google’s Revamped CC Is An AI Agent For Families And Groups
AI & Technology

Google’s Revamped CC Is An AI Agent For Families And Groups

September 17, 2026
Anthropic Says Claude Leads 26% of Its AI Research and Development – Unite.AI
AI & Technology

Anthropic Says Claude Leads 26% of Its AI Research and Development – Unite.AI

September 17, 2026
FAA Says Laser Strikes On Aircraft Fell For The Third Consecutive Year
AI & Technology

FAA Says Laser Strikes On Aircraft Fell For The Third Consecutive Year

September 17, 2026
Next Post
Sumitomo Mitsui: Multiple Positives (NYSE:SMFG)

Sumitomo Mitsui: Multiple Positives (NYSE:SMFG)

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
LIVE NOW: FOMC RATE DECISION & PRESS CONFERENCE 2026

LIVE NOW: FOMC RATE DECISION & PRESS CONFERENCE 2026

September 17, 2026
Iran War Updates: U.N. panel cites possible U.S. war crimes as Trump again says Iran wants a deal – CBS News

Iran War Updates: U.N. panel cites possible U.S. war crimes as Trump again says Iran wants a deal – CBS News

September 17, 2026
Clancy jury unable to come to unanimous decision, judge signals mistrial

Clancy jury unable to come to unanimous decision, judge signals mistrial

September 17, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!