• bitcoinBitcoin(BTC)$79,242.001.16%
  • ethereumEthereum(ETH)$2,509.701.21%
  • tetherTether(USDT)$1.000.02%
  • binancecoinBNB(BNB)$745.70-0.51%
  • rippleXRP(XRP)$1.431.34%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$103.960.86%
  • tronTRON(TRX)$0.338523-0.24%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.043.22%
  • zcashZcash(ZEC)$1,291.919.90%
  • HyperliquidHyperliquid(HYPE)$86.503.49%
  • dogecoinDogecoin(DOGE)$0.0903090.97%
  • RainRain(RAIN)$0.016447-1.31%
  • USDSUSDS(USDS)$1.000.03%
  • whitebitWhiteBIT Coin(WBT)$81.902.72%
  • moneroMonero(XMR)$511.463.15%
  • chainlinkChainlink(LINK)$12.12-3.21%
  • leo-tokenLEO Token(LEO)$9.19-0.29%
  • cardanoCardano(ADA)$0.219394-0.38%
  • stellarStellar(XLM)$0.188307-0.49%
  • bitcoin-cashBitcoin Cash(BCH)$259.000.99%
  • daiDai(DAI)$1.000.00%
  • Ethena USDeEthena USDe(USDE)$1.000.01%
  • USD1USD1(USD1)$1.00-0.01%
  • litecoinLitecoin(LTC)$54.33-0.58%
  • CantonCanton(CC)$0.1052790.64%
  • uniswapUniswap(UNI)$6.65-3.99%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.39-0.45%
  • hedera-hashgraphHedera(HBAR)$0.078807-1.54%
  • avalanche-2Avalanche(AVAX)$7.95-0.19%
  • nearNEAR Protocol(NEAR)$2.6212.92%
  • suiSui(SUI)$0.810.25%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • shiba-inuShiba Inu(SHIB)$0.0000050.12%
  • crypto-com-chainCronos(CRO)$0.059832-1.17%
  • paypal-usdPayPal USD(PYUSD)$1.000.02%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,420.730.67%
  • MemeCoreMemeCore(M)$1.180.78%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BittensorBittensor(TAO)$263.433.69%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • okbOKB(OKB)$114.01-0.34%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.03%
  • mantleMantle(MNT)$0.642.72%
  • AsterAster(ASTER)$0.75-1.09%
  • aaveAave(AAVE)$130.221.16%
  • Pump.funPump.fun(PUMP)$0.0046815.73%
  • polkadotPolkadot(DOT)$1.133.51%
  • pax-goldPAX Gold(PAXG)$4,424.910.67%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

A New Artificial Intelligence (AI) Research Approach Presents Prompt-Based In-Context Learning As An Algorithm Learning Problem From A Statistical Perspective

July 4, 2023
in AI & Technology
Reading Time: 6 mins read
A A
A New Artificial Intelligence (AI) Research Approach Presents Prompt-Based In-Context Learning As An Algorithm Learning Problem From A Statistical Perspective
ShareShareShareShareShare

In-context learning is a recent paradigm where a large language model (LLM) observes a test instance and a few training examples as its input and directly decodes the output without any update to its parameters. This implicit training contrasts with the usual training where the weights are changed based on the examples. 

Here comes the question of why In-context learning would be beneficial. You can suppose that you have two regression tasks that you want to model, but the only limitation is you can only use one model to fit both tasks. Here In-context learning comes in handy as it can learn the regression algorithms per task, which means the model will use separate fitted regressions for different sets of inputs. 

In the paper “Transformers as Algorithms: Generalization and Implicit Model Selection in In-context Learning,” they have formalized the problem of In-context learning as an algorithm learning problem. They have used a transformer as a learning algorithm that can be specialized by training to implement another target algorithm at inference time. In this paper, they have explored the statistical aspects of In-context learning through transformers and did numerical evaluations to verify the theoretical predictions.

[Sponsored] 🔥 Build your personal brand with Taplio  🚀 The 1st all-in-one AI-powered tool to grow on LinkedIn. Create better LinkedIn content 10x faster, schedule, analyze your stats & engage. Try it for free!

In this work, they have investigated two scenarios, in first the prompts are formed of a sequence of i.i.d (input, label) pairs, while in the other the sequence is a trajectory of a dynamic system (the next state depends on the previous state: xm+1 = f(xm) + noise).  

Now the question comes, how we train such a model?

In the training phase of ICL, T tasks are associated with a data distribution  {Dt}t=1T. They independently sample training sequences St from its corresponding distribution for each task. Then they pass a subsequence of St and a value x from sequence St to make a prediction on x. Here is like the meta-learning framework. After prediction, we minimize the loss. The intuition behind ICL training can be interpreted as searching for the optimal algorithm to fit the task at hand.

Next, to obtain generalization bounds on ICL, they borrowed some stability conditions from algorithm stability literature. In ICL, a training example in the prompt influences the future decisions of the algorithms from that point. So to deal with these input perturbations, they needed to impose some conditions on the input. You can read [paper] for more details. Figure 7 shows the results of experiments performed to assess the stability of the learning algorithm (Transformer here). 

RMTL  is the risk (~error) in multi-task learning. One of the insights from the derived bound is that the generalization error of ICL can be eliminated by increasing the sample size n or the number of sequences M per task. The same results can also extend to Stable dynamic systems.

Now let’s see the verification of these bounds using numerical evaluations. 

GPT-2 architecture containing 12 layers, 8 attention heads, and 256-dimensional embedding is used for all experiments. The experiments are performed on regression and linear dynamics. 

  1. Linear Regression: In both figures (2(a) and 2(b)), in-context learning results (Red) outperform the least squares results (Green) and are perfectly aligned with optimal ridge/weighted solution (Black dotted). This, in turn, provides evidence for transformers’ automated model selection ability by learning task priors. 
  2. Partially observed dynamic systems: In Figures (2(c) and 6), Results show that In-context learning outperforms Least square results of almost all orders H=1,2,3,4 (where H is the window size of that slides over the input state sequence to generate input to the model kind of similar to subsequence length)

In conclusion, they successfully showed that the experimental results align with the theoretical predictions. And for the future direction of works, several interesting questions would be worth exploring. 

(1) The proposed bounds are for MTL risk. How can the bounds on individual tasks be controlled?

(2) Can the same results from fully-observed dynamic systems be extended to more general dynamical systems like reinforcement learning?

 (3) From the observation, it was concluded that transfer risk depends only on MTL tasks and their complexity and is independent of the model complexity, so it would be interesting to characterize this inductive bias and what kind of algorithm is being learned by the transformer.


Check out the Paper. All Credit For This Research Goes To the Researchers on This Project. Also, don’t forget to join our Reddit Page, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.


YOU MAY ALSO LIKE

Why It’s Time to Abandon the ‘Set It and Forget It’ Model – Unite.AI

Lyft Is Now Offering Waymo Rides In Nashville

Vineet Kumar is a consulting intern at MarktechPost. He is currently pursuing his BS from the Indian Institute of Technology(IIT), Kanpur. He is a Machine Learning enthusiast. He is passionate about research and the latest advancements in Deep Learning, Computer Vision, and related fields.


🔥 StoryBird.ai just dropped some amazing features. Generate an illustrated story from a prompt. Check it out here. (Sponsored)

Credit: Source link

ShareTweetSendSharePin

Related Posts

Why It’s Time to Abandon the ‘Set It and Forget It’ Model – Unite.AI
AI & Technology

Why It’s Time to Abandon the ‘Set It and Forget It’ Model – Unite.AI

September 9, 2026
Lyft Is Now Offering Waymo Rides In Nashville
AI & Technology

Lyft Is Now Offering Waymo Rides In Nashville

September 9, 2026
Harvey Secures 0M in Fresh Funding, Valuation Climbs to .5B – Unite.AI
AI & Technology

Harvey Secures $550M in Fresh Funding, Valuation Climbs to $15.5B – Unite.AI

September 9, 2026
How To Take Full Advantage Of Gemini When Planning Your Next Trip
AI & Technology

How To Take Full Advantage Of Gemini When Planning Your Next Trip

September 9, 2026
Next Post
Top Three Oil and Gas Stocks to Bet on as Oil Prices Slide

Top Three Oil and Gas Stocks to Bet on as Oil Prices Slide

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
NBC Nightly News with Tom Llamas Full Episode – July 24

NBC Nightly News with Tom Llamas Full Episode – July 24

September 5, 2026
TODAY’s Sheinelle Jones talks about the ‘excruciating lows’ of losing her husband: Full interview

TODAY’s Sheinelle Jones talks about the ‘excruciating lows’ of losing her husband: Full interview

September 5, 2026
Lululemon billionaire founder Chip Wilson files for divorce after 20 years of marriage — with no prenup in place

Lululemon billionaire founder Chip Wilson files for divorce after 20 years of marriage — with no prenup in place

September 6, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!