• bitcoinBitcoin(BTC)$79,644.00-1.72%
  • ethereumEthereum(ETH)$2,455.44-2.48%
  • tetherTether(USDT)$1.000.02%
  • binancecoinBNB(BNB)$746.824.01%
  • rippleXRP(XRP)$1.40-3.23%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$102.32-1.55%
  • tronTRON(TRX)$0.3327481.45%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.040.60%
  • HyperliquidHyperliquid(HYPE)$84.56-1.82%
  • zcashZcash(ZEC)$1,006.973.84%
  • dogecoinDogecoin(DOGE)$0.085724-1.68%
  • RainRain(RAIN)$0.016423-2.58%
  • moneroMonero(XMR)$528.06-3.06%
  • USDSUSDS(USDS)$1.000.00%
  • chainlinkChainlink(LINK)$11.76-2.00%
  • whitebitWhiteBIT Coin(WBT)$73.19-1.04%
  • leo-tokenLEO Token(LEO)$9.27-0.37%
  • cardanoCardano(ADA)$0.213540-3.46%
  • stellarStellar(XLM)$0.182203-0.92%
  • bitcoin-cashBitcoin Cash(BCH)$251.31-2.58%
  • daiDai(DAI)$1.000.00%
  • Ethena USDeEthena USDe(USDE)$1.000.02%
  • CantonCanton(CC)$0.108173-3.43%
  • USD1USD1(USD1)$1.000.01%
  • litecoinLitecoin(LTC)$53.193.92%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.423.87%
  • uniswapUniswap(UNI)$6.22-2.30%
  • hedera-hashgraphHedera(HBAR)$0.0796571.78%
  • Global DollarGlobal Dollar(USDG)$1.00-0.02%
  • avalanche-2Avalanche(AVAX)$7.48-0.34%
  • suiSui(SUI)$0.792.10%
  • shiba-inuShiba Inu(SHIB)$0.0000051.21%
  • nearNEAR Protocol(NEAR)$2.2212.83%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.056182-1.68%
  • tether-goldTether Gold(XAUT)$4,426.38-1.04%
  • Circle USYCCircle USYC(USYC)$1.140.04%
  • MemeCoreMemeCore(M)$1.128.50%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • okbOKB(OKB)$110.992.64%
  • BittensorBittensor(TAO)$236.352.90%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.31%
  • aaveAave(AAVE)$129.95-3.39%
  • AsterAster(ASTER)$0.741.22%
  • pax-goldPAX Gold(PAXG)$4,433.74-1.06%
  • mantleMantle(MNT)$0.571.63%
  • OndoOndo(ONDO)$0.3687171.22%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.056534-2.89%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Salesforce Introduces XGen-7B: A New 7B LLM Trained on up to 8K Sequence Length for 1.5T Tokens

July 2, 2023
in AI & Technology
Reading Time: 5 mins read
A A
Salesforce Introduces XGen-7B: A New 7B LLM Trained on up to 8K Sequence Length for 1.5T Tokens
ShareShareShareShareShare

With recent technological breakthroughs in artificial intelligence, Large Language Models, or LLMs in short, have become increasingly prevalent. Over the past few years, researchers have made rapid advancements in solving several complex language-related tasks by training these models on vast amounts of data in order to comprehend intricate language patterns, generate coherent responses, etc. One area of research that has particularly gained the interest of researchers and developers is the application of LLMs when it comes to handling long-form content to include broader contexts. Some examples of these tasks range from relatively simple tasks like text summarization and code generation to more complex problem statements like protein structure prediction and information retrieval. Long textual sequences consist of information in diverse forms, such as paragraphs, tables, images, etc.; thus, LLMs must be trained to process and understand such elements. Moreover, by effectively considering long-distance structural dependencies, LLMs can identify the connections between different parts of the text and extract the most relevant information. Thus, exposure to a broader range of knowledge allows LLMs to provide more accurate and contextually relevant answers to user queries. 

Yet, despite the numerous potential use cases, most available open-source LLMs, ranging from Meta’s LLaMA to MosaicML’s MPT LLM models, have been trained on sequences with a maximum of 2K tokens. This limitation presents a significant challenge when it comes to modeling longer sequences. Additionally, previous research on model scaling has shown that smaller models trained on a greater number of tokens outperform larger models when given a fixed computational budget. Thus, inspired by the problem at hand and current advances, Salesforce Research made groundbreaking achievements by introducing XGen-7B, a series of 7B LLMs trained on 8K sequence length for 1.5 trillion tokens. The series of models include XGen-7B-4K-Base (with support for 4K sequence length), XGen-7B-8K-Base (with support for 8K sequence length), and XGen-7B-8k-Inst which has been fine-tuned on public-domain instructional data (released only for research purposes). The striking characteristic of these LLMs is that on standard NLP benchmarks, XGen achieves comparable or better results when compared to other state-of-the-art LLMs of similar size like MPT, Falcon, LLaMA, etc.

The XGen-7b models employed in this study were trained using Salesforce’s proprietary library JaxFormer, which enables efficient training of LLMs utilizing data and model parallelism specifically optimized for TPU-v4 hardware. The training process followed the guidelines of LLaMA, augmented with two additional investigations. The first exploration focused on understanding “loss spikes,” where the loss suddenly and temporarily increases during training without a clear underlying cause. Although the root cause of these spikes remains unknown, the researchers identified factors such as “sequential over parallel circuits,” “swish-GLU over GeLU,” and “RMS-Norm over Layer-norm” as potential contributors to training instability. The second aspect addressed was sequence length. Since training with longer sequences incurs significantly higher computational costs due to the quadratic complexity of self-attention, a staged training approach was adopted. The training initially encompassed 800B tokens with a sequence length of 2k tokens, followed by 400B tokens with 4k length, and finally, 300B tokens with 8k length. 

🔥 Join The Fastest Growing ML Subreddit

To assess the capabilities of the XGen-7b 8k model in comprehending longer contexts, the researchers conducted evaluations using three primary tasks: long-form dialogue generation, text summarization, and question-answering. The researchers used the instruction-tuned model for their evaluations pertaining to the difficulty of the tasks at hand. Regarding long-form dialogue generation, the researchers utilized three tasks for assessment: AMI meeting summarization, ForeverDreaming, and TVMegaSite screenplay summarization. Across all metrics, the XGen-7B-inst model achieved the highest scores compared to several other instruction-tuned models, demonstrating its superior performance.

For long-form question-answering, the researchers generated questions using ChatGPT based on Wikipedia documents covering diverse topics like Physics, Engineering, History, and Entertainment, along with their corresponding summaries. The LLM-generated answers, which were 256 tokens long, were evaluated using GPT-4 based on their structure, organization, and relevance to the question and source document. In this scenario, the XGen-7B-8k-Inst model outperformed the baseline models, which are limited to 2k tokens, showcasing its superior performance. In terms of text summarization, the researchers employed two datasets from different domains, specifically meeting conversations and government reports, to evaluate the XGen-7b model. The results revealed that the XGen-7b model significantly outperformed other baseline models in these tasks, indicating its superior performance in text summarization as well. 

The evaluations demonstrated that the XGen-7b model excelled in understanding longer contexts across various tasks, including long-form dialogue generation, question-answering, and text summarization. Its performance surpassed that of other instruction-tuned and baseline models, showcasing its effectiveness in comprehending and generating coherent responses in extensive text contexts. Nevertheless, despite its efficacy, the researchers acknowledge a limitation of the XGen model, as it is not exempt from biases and has the potential to generate toxic responses, a characteristic it shares with many other AI models. Salesforce Research has also open-sourced its code to allow the community to explore its work.

Check Out the SF Blog and Github Link. Don’t forget to join our 25k+ ML SubReddit, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more. If you have any questions regarding the above article or if we missed anything, feel free to email us at [email protected]


Featured Tools:

🚀 Check Out 100’s AI Tools in AI Tools Club


YOU MAY ALSO LIKE

Google Launches Agentic Video Understanding for Gemini Flash Models, Cutting Video Tokens by Up to 88%

Seattle Times and Newsday Sue OpenAI and Microsoft Over News Content – Unite.AI

Khushboo Gupta is a consulting intern at MarktechPost. She is currently pursuing her B.Tech from the Indian Institute of Technology(IIT), Goa. She is passionate about the fields of Machine Learning, Natural Language Processing and Web Development. She enjoys learning more about the technical field by participating in several challenges.


🔥 StoryBird.ai just dropped some amazing features. Generate an illustrated story from a prompt. Check it out here. (Sponsored)

Credit: Source link

ShareTweetSendSharePin

Related Posts

Google Launches Agentic Video Understanding for Gemini Flash Models, Cutting Video Tokens by Up to 88%
AI & Technology

Google Launches Agentic Video Understanding for Gemini Flash Models, Cutting Video Tokens by Up to 88%

September 5, 2026
Seattle Times and Newsday Sue OpenAI and Microsoft Over News Content – Unite.AI
AI & Technology

Seattle Times and Newsday Sue OpenAI and Microsoft Over News Content – Unite.AI

September 5, 2026
How To See What’s Taking Up Space On Your Windows PC
AI & Technology

How To See What’s Taking Up Space On Your Windows PC

September 4, 2026
The Tetris Company Wants Nothing To Do With The White House’s New Copycat Game
AI & Technology

The Tetris Company Wants Nothing To Do With The White House’s New Copycat Game

September 4, 2026
Next Post
Newport News school reopens after six-year-old shot teacher

Newport News school reopens after six-year-old shot teacher

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Dr. Fauci’s opening statement at Covid hearing

Dr. Fauci’s opening statement at Covid hearing

September 3, 2026
Trump says Graham ‘never saw a war that he didn’t like’

Trump says Graham ‘never saw a war that he didn’t like’

September 3, 2026
This is everything that’s wrong with consumerism

This is everything that’s wrong with consumerism

September 1, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!