• bitcoinBitcoin(BTC)$84,628.000.65%
  • ethereumEthereum(ETH)$2,694.180.25%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$778.010.78%
  • rippleXRP(XRP)$1.53-0.32%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$122.701.16%
  • tronTRON(TRX)$0.334031-0.63%
  • zcashZcash(ZEC)$1,592.211.57%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.062.87%
  • HyperliquidHyperliquid(HYPE)$91.880.37%
  • dogecoinDogecoin(DOGE)$0.097478-0.16%
  • chainlinkChainlink(LINK)$14.15-0.50%
  • moneroMonero(XMR)$548.21-1.58%
  • whitebitWhiteBIT Coin(WBT)$84.370.55%
  • USDSUSDS(USDS)$1.000.00%
  • cardanoCardano(ADA)$0.256128-0.01%
  • RainRain(RAIN)$0.012577-3.89%
  • leo-tokenLEO Token(LEO)$9.050.99%
  • stellarStellar(XLM)$0.216818-0.69%
  • nearNEAR Protocol(NEAR)$5.309.18%
  • bitcoin-cashBitcoin Cash(BCH)$336.84-0.40%
  • uniswapUniswap(UNI)$9.741.31%
  • litecoinLitecoin(LTC)$71.22-1.30%
  • CantonCanton(CC)$0.1379632.81%
  • suiSui(SUI)$1.278.14%
  • Ethena USDeEthena USDe(USDE)$1.000.02%
  • avalanche-2Avalanche(AVAX)$11.001.10%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.699.27%
  • daiDai(DAI)$1.000.01%
  • USD1USD1(USD1)$1.00-0.01%
  • hedera-hashgraphHedera(HBAR)$0.0944900.79%
  • BittensorBittensor(TAO)$326.95-0.08%
  • shiba-inuShiba Inu(SHIB)$0.000006-0.51%
  • crypto-com-chainCronos(CRO)$0.0674962.72%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • BitwayBitway(BTW)$1.1911.65%
  • EthenaEthena(ENA)$0.2825384.06%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • MemeCoreMemeCore(M)$1.21-0.45%
  • OndoOndo(ONDO)$0.553.57%
  • quant-networkQuant(QNT)$185.5350.94%
  • tether-goldTether Gold(XAUT)$4,278.700.01%
  • okbOKB(OKB)$121.360.33%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • aaveAave(AAVE)$155.510.57%
  • Pump.funPump.fun(PUMP)$0.00501512.81%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.10%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Google DeepMind Researchers Unlock the Potential of Decoding-Based Regression for Tabular and Density Estimation Tasks

February 4, 2025
in AI & Technology
Reading Time: 6 mins read
A A
Google DeepMind Researchers Unlock the Potential of Decoding-Based Regression for Tabular and Density Estimation Tasks
ShareShareShareShareShare

Regression tasks, which involve predicting continuous numeric values, have traditionally relied on numeric heads such as Gaussian parameterizations or pointwise tensor projections. These traditional approaches have strong distributional assumption requirements, require a lot of labeled data, and tend to break down when modeling advanced numerical distributions. New research on large language models introduces a different approach—representing numerical values as sequences of discrete tokens and using auto-regressive decoding for prediction. This shift, however, comes with several serious challenges, including the need for an efficient tokenization mechanism, the potential for numeric precision loss, the need to maintain stable training, and the need to overcome the lack of inductive bias of sequential token forms for numerical values. Overcoming these challenges would lead to an even more powerful, data-efficient, and flexible regression framework, thus extending the application of deep learning models beyond traditional approaches.

Traditional regression models rely on numeric tensor projections or parametric distributional heads, such as Gaussian models. While these conventional approaches are widespread, they also have several drawbacks. Gaussian-based models have the drawback of assuming normally distributed outputs, restricting the ability to model more advanced, multimodal distributions. Pointwise regression heads struggle with highly non-linear or discontinuous relationships, which restricts their ability to generalize on various datasets. High-dimensional models, such as histogram-based Riemann distributions, are computationally and data-intensive and, therefore, inefficient. Furthermore, many traditional approaches require explicit normalization or scaling of output, introducing an additional layer of complexity and potential instability. While conventional work has tried to employ text-to-text regression using large language models, little systematic work has been done on “anything-to-text” regression, where numeric outputs are represented as sequences of tokens, thus introducing a new paradigm for numerical prediction.

YOU MAY ALSO LIKE

Why The iPhone Duo Could Be Beneficial For Samsung’s Galaxy Z Fold 8

How To Improve Your Router’s Security In 10 Minutes

Researchers from Google DeepMind propose an alternative regression formulation, reframing numeric prediction as an auto-regressive sequence generation problem.  Instead of generating scalar values directly, this method encodes numbers as token sequences and employs constrained decoding to generate valid numerical outputs. Encoding numeric values as discrete token sequences makes this method more flexible and expressive when modeling real-valued data. Unlike Gaussian-based approaches, this method does not entail strong distributional assumptions about data, thus making it more generalizable to real-world tasks with heterogeneous patterns. The model accommodates precise modeling of multimodal, complex distributions, thus improving its performance in density estimation as well as pointwise regression tasks. By leveraging the advantages of autoregressive decoders, it takes advantage of recent language modeling progress while still retaining competitive performance relative to standard numeric heads. This formulation presents a robust and flexible framework that can model a wide range of numeric relationships precisely, offering a practical substitute to standard regression methods that are usually regarded as inflexible.

The approach employs two tokenization methods for numeric representation: normalized tokenization and unnormalized tokenization. Normalized tokenization encodes numbers in a fixed range with base-B expansion to provide finer precision with increasing sequence length. Unnormalized tokenization extends the same idea to broader numeric ranges with a generalized floating-point representation such as IEEE-754 without the necessity of explicit normalization. A transformer auto-regressive model generates numeric outputs token by token subject to constraints to provide valid numeric sequences. The model is trained using cross-entropy loss over the token sequence to provide accurate numeric representation. Instead of predicting a scalar output directly, the system samples token sequences and employs statistical estimation techniques, such as mean or median computation, for final prediction. Evaluations are conducted on real-world tabular regression datasets of OpenML-CTR23 and AMLB benchmarks and compared with Gaussian mixture models, histogram-based regression, and standard pointwise regression heads. Hyperparameter tuning is conducted across various decoder settings, such as variations in the number of layers, hidden units, and token vocabularies, to provide optimized performance.

Experiments show that the model successfully captures intricate numeric relationships, achieving strong performance on a variety of regression tasks. It attains high Kendall-Tau correlation scores on tabular regression, often outperforming baseline models, especially in low-data settings where numeric stability is essential. The method is also better in density estimation, successfully capturing intricate distributions and outperforming Gaussian mixture models and Riemann-based approaches in negative log-likelihood tests. Model size tuning at the start improves performance, with overcapacity causing overfitting. Numeric stability is greatly improved by error correction methods like token repetition and majority voting, minimizing vulnerability to outliers. These results make this regression framework a robust and adaptive alternative to traditional methods, showing its capacity to successfully generalize across various datasets and modeling tasks.

This work introduces a novel approach to numeric prediction by leveraging tokenized representations and auto-regressive decoding. By substituting traditional numeric regression heads with token-based outputs, the framework improves flexibility in modeling real-valued data. It attains competitive performance on various regression tasks, especially in density estimation and tabular modeling, while providing theoretical guarantees for approximating arbitrary probability distributions. It outperforms traditional regression methods in important contexts, especially in modeling intricate distributions and sparse training data. Future work involves improving tokenization methods for better numeric precision and stability, extending the framework to multi-output regression and high-dimensional prediction tasks, and investigating its applications in reinforcement learning reward modeling and vision-based numeric estimation. These results make sequence-based numeric regression a promising alternative to traditional methods, expanding the scope of tasks that language models can successfully solve.


Check out the Paper and GitHub Page. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter and join our Telegram Channel and LinkedIn Group. Don’t Forget to join our 75k+ ML SubReddit.

🚨 Marktechpost is inviting AI Companies/Startups/Groups to partner for its upcoming AI Magazines on ‘Open Source AI in Production’ and ‘Agentic AI’.


Aswin AK is a consulting intern at MarkTechPost. He is pursuing his Dual Degree at the Indian Institute of Technology, Kharagpur. He is passionate about data science and machine learning, bringing a strong academic background and hands-on experience in solving real-life cross-domain challenges.

✅ [Recommended] Join Our Telegram Channel

Credit: Source link

ShareTweetSendSharePin

Related Posts

Why The iPhone Duo Could Be Beneficial For Samsung’s Galaxy Z Fold 8
AI & Technology

Why The iPhone Duo Could Be Beneficial For Samsung’s Galaxy Z Fold 8

September 27, 2026
How To Improve Your Router’s Security In 10 Minutes
AI & Technology

How To Improve Your Router’s Security In 10 Minutes

September 27, 2026
Humanoid Robots Are Getting Even Creepier (This One Can Cry On Command)
AI & Technology

Humanoid Robots Are Getting Even Creepier (This One Can Cry On Command)

September 27, 2026
AI Coding Agents for Enterprise: IP Indemnity, Data Residency and 500-Seat Cost Compared
AI & Technology

AI Coding Agents for Enterprise: IP Indemnity, Data Residency and 500-Seat Cost Compared

September 27, 2026
Next Post
Victims of the deadly midair collision will miss out on many of their daughter’s firsts

Victims of the deadly midair collision will miss out on many of their daughter’s firsts

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Secret Service responds to reports of Iranian threat against Trump’s son Barron

Secret Service responds to reports of Iranian threat against Trump’s son Barron

September 24, 2026
U.S.-Canada Tariff Talks Jolt Auto Industry; Drive to Decision Day at the Iowa State Fair | Aug. 21

U.S.-Canada Tariff Talks Jolt Auto Industry; Drive to Decision Day at the Iowa State Fair | Aug. 21

September 26, 2026
Judge lifts Trump's White House ban on CNN, MS NOW and Politico – Reuters

Judge lifts Trump's White House ban on CNN, MS NOW and Politico – Reuters

September 24, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!