• bitcoinBitcoin(BTC)$85,876.005.75%
  • ethereumEthereum(ETH)$2,761.795.35%
  • tetherTether(USDT)$1.000.04%
  • binancecoinBNB(BNB)$798.595.16%
  • rippleXRP(XRP)$1.496.59%
  • usd-coinUSDC(USDC)$1.000.04%
  • solanaSolana(SOL)$117.757.65%
  • tronTRON(TRX)$0.3448940.15%
  • zcashZcash(ZEC)$1,496.003.21%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.040.00%
  • HyperliquidHyperliquid(HYPE)$92.730.12%
  • dogecoinDogecoin(DOGE)$0.09650111.48%
  • moneroMonero(XMR)$567.824.36%
  • whitebitWhiteBIT Coin(WBT)$86.484.46%
  • RainRain(RAIN)$0.0140944.66%
  • chainlinkChainlink(LINK)$13.004.88%
  • USDSUSDS(USDS)$1.000.04%
  • cardanoCardano(ADA)$0.2443317.69%
  • leo-tokenLEO Token(LEO)$8.89-0.43%
  • stellarStellar(XLM)$0.2093037.13%
  • uniswapUniswap(UNI)$8.941.14%
  • bitcoin-cashBitcoin Cash(BCH)$264.395.73%
  • nearNEAR Protocol(NEAR)$4.06-2.09%
  • avalanche-2Avalanche(AVAX)$11.06-0.61%
  • Ethena USDeEthena USDe(USDE)$1.000.06%
  • litecoinLitecoin(LTC)$62.818.66%
  • CantonCanton(CC)$0.1174259.20%
  • daiDai(DAI)$1.000.02%
  • USD1USD1(USD1)$1.000.03%
  • suiSui(SUI)$1.0320.05%
  • hedera-hashgraphHedera(HBAR)$0.0929388.09%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.433.63%
  • shiba-inuShiba Inu(SHIB)$0.0000068.50%
  • MemeCoreMemeCore(M)$1.49-3.44%
  • BittensorBittensor(TAO)$289.3711.19%
  • Global DollarGlobal Dollar(USDG)$1.000.03%
  • crypto-com-chainCronos(CRO)$0.0634317.46%
  • paypal-usdPayPal USD(PYUSD)$1.000.05%
  • tether-goldTether Gold(XAUT)$4,351.13-0.51%
  • okbOKB(OKB)$122.654.47%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BitwayBitway(BTW)$0.9025.87%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.01%
  • aaveAave(AAVE)$142.994.44%
  • OndoOndo(ONDO)$0.4490825.97%
  • EthenaEthena(ENA)$0.212960-1.39%
  • mantleMantle(MNT)$0.646.27%
  • pepePepe(PEPE)$0.00000523.27%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

CMU Researchers Propose In-Context Abstraction Learning (ICAL): An AI Method that Builds a Memory of Multimodal Experience Insights from Sub-Optimal Demonstrations and Human Feedback

June 29, 2024
in AI & Technology
Reading Time: 5 mins read
A A
CMU Researchers Propose In-Context Abstraction Learning (ICAL): An AI Method that Builds a Memory of Multimodal Experience Insights from Sub-Optimal Demonstrations and Human Feedback
ShareShareShareShareShare

Humans are versatile; they can quickly apply what they’ve learned from little examples to larger contexts by combining new and old information. Not only can they foresee possible setbacks and determine what is important for success, but they swiftly learn to adjust to different situations by practicing and receiving feedback on what works. This process can refine and transfer knowledge across many jobs and situations.

Extraction of high-level insights from trajectories and experiences has been the subject of recent research utilizing visual-language models (VLMs) and large-language models (LLMs). The model’s introspection yields these insights, which are then used to improve performance by attaching them to prompts, using their remarkable ability to learn in context. The majority of current approaches rely on language in one of several ways: to communicate job rewards, to store human adjustments after failures, to have domain experts create or select examples without reflection, or to set regulations and incentives through language. The approaches in question mostly rely on text and don’t use visual cues or demonstrations. They also rely solely on introspection in the event of failure, which is just one of many ways that machines and humans can accumulate experiences and derive insights.

YOU MAY ALSO LIKE

Collaboration Must Sit At the Heart of Manufacturing’s Multi-Agentic AI Approach. Here’s How. – Unite.AI

How To Choose The Right USB To USB-C Adapter

A new study by Carnegie Mellon University and Google DeepMind demonstrates a novel approach to training VLMs. This approach, called In-Context Abstraction Learning (ICAL), guides VLMs to build multimodal abstractions in novel domains. In simpler terms, ICAL helps VLMs to understand and learn from their experiences in different situations, allowing them to adapt and perform better in new tasks. The approach emphasizes learning abstractions that encompass tasks’ dynamics and critical knowledge, in contrast to earlier efforts that store and recall successful action plans or trajectories. To be more precise, ICAL addresses four distinct kinds of cognitive abstractions: 

  1. Task and causal relationships, which reveal the underlying principles or actions required to accomplish a goal and the interconnectedness of its elements
  2. Changes in object states, which show the different shapes or states an object can take
  3. Temporal abstractions, which divide tasks into smaller objectives
  4. Task construals emphasize important visual aspects within a task. 

In response to good or bad demonstrations, ICAL tells a VLM to optimize the trajectories and generate relevant verbal and visual abstractions. Humans’ natural language input guides the execution of the trajectory in the environment, which further refines these abstractions. The model can enhance its execution and abstraction capabilities with each phase of abstraction generation, using previously derived abstractions. The acquired abstractions concisely summarize the rules, focal regions, action sequences, state transitions, and visual representations expressed in free-form natural language. 

Using the acquired example abstractions, the researchers conducted a thorough evaluation of their agent on three different benchmarks: VisualWebArena, TEACh, and Ego4D. These benchmarks are widely used in the field of AI and provide a standard for evaluating the performance of different models. VisualWebArena is used for multimodal autonomous web tasks, TEACh for dialogue-based training in the home, and Ego4D for video action anticipation. The effectiveness of ICAL-taught abstractions for in-context learning is demonstrated by their agent’s new state-of-the-art performance in TEACh, which outperforms VLM agents that rely on raw demos or extensive domain-expert hand-written examples. In particular, the proposed method improves the success of goal conditions by 12.6% compared to the prior SOTA, HELPER. After just ten cases, the findings show that this method delivers a speed boost of 14.7% on unseen jobs and grows with the size of the external memory. The goal-condition performance is enhanced by an additional 4.9% when the learned examples are combined with LoRA-based LLM fine-tuning [32]. With a success percentage of 22.7% in the VisualWebArena, the agent outperforms the state-of-the-art GPT4Vision + Set of Marks by a margin of 14.3%. Using the chain of thought, ICAL reduces the noun edit distance by 6.4 and the action edit distance by 1.7 in the Ego4D environment, outperforming few-shot GPT4V. It also competes closely with fully supervised approaches, even though it uses 639 times less in-domain training data. 

The potential of the ICAL method is vast, as it consistently outperforms in-context learning using action plans or trajectories without such abstractions, while significantly reducing the need for meticulously constructed examples. The team acknowledges several areas for further study and potential challenges for ICAL, such as its ability to handle noisy demos and its dependence on a static action API. However, these are seen as opportunities for growth and improvement rather than limitations, instilling a sense of optimism and hope for the future of ICAL.


Check out the Paper, Project, and GitHub. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. 

Join our Telegram Channel and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 45k+ ML SubReddit


🚀 Create, edit, and augment tabular data with the first compound AI system, Gretel Navigator, now generally available! [Advertisement]


Dhanshree Shenwai is a Computer Science Engineer and has a good experience in FinTech companies covering Financial, Cards & Payments and Banking domain with keen interest in applications of AI. She is enthusiastic about exploring new technologies and advancements in today’s evolving world making everyone’s life easy.

🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

Collaboration Must Sit At the Heart of Manufacturing’s Multi-Agentic AI Approach. Here’s How. – Unite.AI
AI & Technology

Collaboration Must Sit At the Heart of Manufacturing’s Multi-Agentic AI Approach. Here’s How. – Unite.AI

September 21, 2026
How To Choose The Right USB To USB-C Adapter
AI & Technology

How To Choose The Right USB To USB-C Adapter

September 21, 2026
A Laptop That Works Better With Your Android Phone
AI & Technology

A Laptop That Works Better With Your Android Phone

September 21, 2026
How AI Modernizes Lending Alongside Legacy Banking Systems Without a Teardown – Unite.AI
AI & Technology

How AI Modernizes Lending Alongside Legacy Banking Systems Without a Teardown – Unite.AI

September 21, 2026
Next Post
Fulton County DA: Leaked witness videos in Trump trial is meant to ‘intimidate witnesses’

Fulton County DA: Leaked witness videos in Trump trial is meant to 'intimidate witnesses'

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Sen. John Barrasso says Trump isn’t violating the Constitution by banning reporters from the White House – NBC News

Sen. John Barrasso says Trump isn’t violating the Constitution by banning reporters from the White House – NBC News

September 20, 2026
Vicinity Centres Stapled Securities (CNRAF) Discusses Capability Showcase With Focus on Development Strategy and Asset Portfolio – Slideshow (OTCMKTS:CNRAF) 2026-09-17

Vicinity Centres Stapled Securities (CNRAF) Discusses Capability Showcase With Focus on Development Strategy and Asset Portfolio – Slideshow (OTCMKTS:CNRAF) 2026-09-17

September 17, 2026
Nurix Therapeutics, Inc. (NRIX) Presents at H.C. Wainwright 28th Annual Global Investment Conference Transcript

Nurix Therapeutics, Inc. (NRIX) Presents at H.C. Wainwright 28th Annual Global Investment Conference Transcript

September 17, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!