• bitcoinBitcoin(BTC)$65,401.00-0.60%
  • ethereumEthereum(ETH)$1,878.97-2.50%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$568.60-0.30%
  • usd-coinUSDC(USDC)$1.000.00%
  • rippleXRP(XRP)$1.11-2.40%
  • solanaSolana(SOL)$75.77-2.50%
  • tronTRON(TRX)$0.329339-0.10%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.043.10%
  • whitebitWhiteBIT Coin(WBT)$56.67-1.30%
  • HyperliquidHyperliquid(HYPE)$58.21-1.30%
  • dogecoinDogecoin(DOGE)$0.069330-4.50%
  • USDSUSDS(USDS)$1.000.00%
  • RainRain(RAIN)$0.014006-2.60%
  • leo-tokenLEO Token(LEO)$9.58-1.80%
  • zcashZcash(ZEC)$509.60-1.30%
  • moneroMonero(XMR)$354.10-0.10%
  • chainlinkChainlink(LINK)$8.44-2.00%
  • stellarStellar(XLM)$0.183066-1.90%
  • cardanoCardano(ADA)$0.167170-4.10%
  • CantonCanton(CC)$0.1221440.10%
  • daiDai(DAI)$1.000.00%
  • bitcoin-cashBitcoin Cash(BCH)$211.51-3.20%
  • USD1USD1(USD1)$1.000.00%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.46-1.90%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • litecoinLitecoin(LTC)$46.77-1.10%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • hedera-hashgraphHedera(HBAR)$0.070912-2.70%
  • suiSui(SUI)$0.74-2.30%
  • Circle USYCCircle USYC(USYC)$1.130.00%
  • crypto-com-chainCronos(CRO)$0.057357-0.30%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • avalanche-2Avalanche(AVAX)$6.26-4.70%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • nearNEAR Protocol(NEAR)$1.901.60%
  • tether-goldTether Gold(XAUT)$4,026.61-2.20%
  • shiba-inuShiba Inu(SHIB)$0.000004-2.60%
  • uniswapUniswap(UNI)$3.78-1.20%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.20%
  • OndoOndo(ONDO)$0.403488-0.70%
  • BittensorBittensor(TAO)$192.65-1.30%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0564825.30%
  • pax-goldPAX Gold(PAXG)$4,024.33-2.30%
  • okbOKB(OKB)$82.160.10%
  • AsterAster(ASTER)$0.620.30%
  • HTX DAOHTX DAO(HTX)$0.000002-0.60%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • usddUSDD(USDD)$1.000.00%
  • MemeCoreMemeCore(M)$1.151.10%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Meta AI Researchers Propose MEGABYTE: A Multiscale Decoder Architecture that Enables End-to-End Differentiable Modeling of Sequences of Over One Million Bytes

May 19, 2023
in AI & Technology
Reading Time: 4 mins read
A A
Meta AI Researchers Propose MEGABYTE: A Multiscale Decoder Architecture that Enables End-to-End Differentiable Modeling of Sequences of Over One Million Bytes
ShareShareShareShareShare

Million-byte sequences are common as music, picture, and video files frequently have several megabyte sizes. However, because of the quadratic cost of self-attention and, more significantly, the expense of large feedforward networks per position, large transformer decoders (LLMs) normally only require a few thousand tokens of context. This significantly reduces the range of tasks for which LLMs may be used. Researchers from META present MEGABYTE, a novel method for simulating lengthy byte sequences. Byte sequences are divided into fixed-sized patches roughly equivalent to tokens. 

Then, their model has three components: 

(1) A local module, a tiny autoregressive model that forecasts bytes within a patch. 

🚀 JOIN the fastest ML Subreddit Community

(2) A patch embedder merely encodes a patch by losslessly concatenating embeddings of each byte. 

(3) A global module, a big autoregressive transformer that inputs and outputs patch representations.

Importantly, most byte predictions are straightforward for many tasks (such as completing a word given the initial few letters), negating the need for massive networks per byte and allowing for considerably smaller models for intra-patch modeling. For extended sequence modeling, the MEGABYTE architecture offers three key advantages over Transformers: Self-attention that is sub-quadratic The vast majority of research on long sequence models has been devoted to reducing the quadratic cost of self-attention. Lengthy sequences are divided into two shorter sequences using MEGABYTE, and the self-attention cost is decreased to O(N(4/3)) by using optimum patch sizes, which are still tractable for lengthy sequences. Layers with per-patch feedforward. MEGABYTE allows for far bigger and more expressive models at the same cost by using huge feedforward layers per patch rather than per position. More than 98% of FLOPS are used in GPT3-size models to compute position-wise feedforward layers. 

Decoding Parallelism three Transformers must serially process all calculations during generation since each timestep’s input results from the initial output. MEGABYTE makes Greater parallelism during generation possible thanks to the parallel production of representations for patches. With patch size P, MEGABYTE may utilize a layer with mP parameters once for the same price as a baseline transformer, using the same feedforward layer with m parameters P times. For instance, when trained on the same compute, a MEGABYTE model with 1.5B parameters may create sequences 40% quicker than a conventional 350M Transformer while increasing perplexity. 

Together, these enhancements enable us to expand to lengthy sequences, increase generation speed during deployment, and train much bigger and better-performing models for the same computational budget. Sequences of bytes are translated into bigger discrete tokens in existing autoregressive models, which generally involve some tokenization. This is where MEGABYTE stands in stark contrast. Tokenization makes pre-processing, multi-modal modeling, and transfer to different domains more difficult while obscuring the model’s beneficial structure. Additionally, it implies that most cutting-edge models are still in progress. The most popular methods of tokenization lose information without language-specific heuristics. 

Therefore, switching from tokenization to performant and effective byte models would have several benefits. They carry out in-depth tests for both strong baselines and MEGABYTE. To concentrate their comparisons entirely on the model architecture rather than training resources, which are known to be advantageous to all models, they employ a single compute and data budget across all models. They discover that MEGABYTE enables byte-level models to reach state-of-the-art perplexities for density estimation on ImageNet, perform competitively with subword models on extended context language modeling, and allow audio modeling from raw audio data. These findings show that tokenization-free autoregressive sequence modeling is scaleable.


Check out the Paper. Don’t forget to join our 21k+ ML SubReddit, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more. If you have any questions regarding the above article or if we missed anything, feel free to email us at [email protected]

🚀 Check Out 100’s AI Tools in AI Tools Club


YOU MAY ALSO LIKE

Microsoft launches new in-house AI models it says cut costs up to 89% versus OpenAI

Scientists Develop Handheld Device For Measuring When Your Body Is Burning Fat

Aneesh Tickoo is a consulting intern at MarktechPost. He is currently pursuing his undergraduate degree in Data Science and Artificial Intelligence from the Indian Institute of Technology(IIT), Bhilai. He spends most of his time working on projects aimed at harnessing the power of machine learning. His research interest is image processing and is passionate about building solutions around it. He loves to connect with people and collaborate on interesting projects.


➡️ Meet Bright Data: The World’s #1 Web Data Platform

Credit: Source link

ShareTweetSendSharePin

Related Posts

Microsoft launches new in-house AI models it says cut costs up to 89% versus OpenAI
AI & Technology

Microsoft launches new in-house AI models it says cut costs up to 89% versus OpenAI

July 23, 2026
Scientists Develop Handheld Device For Measuring When Your Body Is Burning Fat
AI & Technology

Scientists Develop Handheld Device For Measuring When Your Body Is Burning Fat

July 23, 2026
Agentic coding goes hands-free as OpenAI brings GPT-Live’s full duplex voice control to Codex and ChatGPT on the desktop
AI & Technology

Agentic coding goes hands-free as OpenAI brings GPT-Live’s full duplex voice control to Codex and ChatGPT on the desktop

July 23, 2026
Meta’s Pro-AI Ad Campaign Is Conspicuously Light On AI
AI & Technology

Meta’s Pro-AI Ad Campaign Is Conspicuously Light On AI

July 23, 2026
Next Post
Bloomberg Technology 07/08/2022 Elon Musk Ends Twitter Deal

Bloomberg Technology 07/08/2022 Elon Musk Ends Twitter Deal

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Trump did not present any evidence that ballots were tampered with in speech

Trump did not present any evidence that ballots were tampered with in speech

July 23, 2026
Why Fixed Income ETFs Are Having A Moment — And How To Use Them

Why Fixed Income ETFs Are Having A Moment — And How To Use Them

July 21, 2026
This Spotify Setting Makes Your Music Feel DJ Mixed

This Spotify Setting Makes Your Music Feel DJ Mixed

July 21, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!