• bitcoinBitcoin(BTC)$78,374.00-1.13%
  • ethereumEthereum(ETH)$2,477.84-1.19%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$721.99-4.25%
  • rippleXRP(XRP)$1.39-3.12%
  • usd-coinUSDC(USDC)$1.00-0.02%
  • solanaSolana(SOL)$101.94-2.50%
  • tronTRON(TRX)$0.3395870.26%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.03-0.94%
  • zcashZcash(ZEC)$1,242.950.25%
  • HyperliquidHyperliquid(HYPE)$84.10-2.30%
  • dogecoinDogecoin(DOGE)$0.085763-5.36%
  • RainRain(RAIN)$0.0164082.11%
  • USDSUSDS(USDS)$1.00-0.01%
  • moneroMonero(XMR)$514.661.62%
  • whitebitWhiteBIT Coin(WBT)$80.87-1.49%
  • chainlinkChainlink(LINK)$11.82-5.87%
  • leo-tokenLEO Token(LEO)$9.190.07%
  • cardanoCardano(ADA)$0.213362-3.41%
  • stellarStellar(XLM)$0.180441-5.16%
  • bitcoin-cashBitcoin Cash(BCH)$250.37-3.76%
  • daiDai(DAI)$1.000.00%
  • Ethena USDeEthena USDe(USDE)$1.00-0.03%
  • USD1USD1(USD1)$1.00-0.02%
  • CantonCanton(CC)$0.103687-5.15%
  • litecoinLitecoin(LTC)$52.76-2.97%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.38-1.07%
  • uniswapUniswap(UNI)$6.06-12.20%
  • avalanche-2Avalanche(AVAX)$7.83-2.55%
  • hedera-hashgraphHedera(HBAR)$0.077033-3.04%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • nearNEAR Protocol(NEAR)$2.517.09%
  • suiSui(SUI)$0.77-6.42%
  • shiba-inuShiba Inu(SHIB)$0.000005-4.04%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • crypto-com-chainCronos(CRO)$0.058075-3.31%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • MemeCoreMemeCore(M)$1.223.29%
  • tether-goldTether Gold(XAUT)$4,413.200.63%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BittensorBittensor(TAO)$256.66-0.90%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • okbOKB(OKB)$113.39-1.33%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.02%
  • mantleMantle(MNT)$0.60-5.48%
  • AsterAster(ASTER)$0.73-4.21%
  • aaveAave(AAVE)$125.04-3.82%
  • pax-goldPAX Gold(PAXG)$4,417.080.63%
  • polkadotPolkadot(DOT)$1.11-8.52%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0569051.72%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

New Machine Learning Research from MIT Proposes Compositional Foundation Models for Hierarchical Planning (HiP): Integrating Language, Vision, and Action for Long-Horizon Tasks Solutions

September 22, 2023
in AI & Technology
Reading Time: 5 mins read
A A
New Machine Learning Research from MIT Proposes Compositional Foundation Models for Hierarchical Planning (HiP): Integrating Language, Vision, and Action for Long-Horizon Tasks Solutions
ShareShareShareShareShare

Think about the challenge of preparing a cup of tea in a strange home. An efficient strategy for completing this task is to reason hierarchically at several levels, including an abstract level (for example, the high-level steps required to heat the tea), a concrete geometric level (for example, how they should physically move to and through the kitchen), and a control level (for example, how they should move their joints to lift a cup). An abstract plan to search cabinets for tea kettles must also be physically conceivable at the geometric level and executable given the actions they are capable of. This is why it is crucial that reasoning at each level is consistent with one another. In this study, they investigate the development of unique long-horizon task-solving bots capable of employing hierarchical reasoning. 

Large “foundation models” have taken the lead in tackling problems in mathematical reasoning, computer vision, and natural language processing. Creating a “foundation model” that can address unique and long-horizon decision-making problems is an issue that has attracted much attention in light of this paradigm. In several earlier studies, matched visual, linguistic, and action data were gathered, and a single neural network was trained to handle long-horizon tasks. However, it is expensive and challenging to scale up the coupled visual, linguistic, and action data collection. Another line of earlier research uses task-specific robot demonstrations to refine large language models (LLM) on visual and linguistic inputs. This is a concern since, in contrast to the wealth of material available on the Internet, examples of coupled vision and language robots are difficult to find and expensive to compile. 

Furthermore, because the model weights are not open-sourced, it is currently difficult to finetune high-performing language models like GPT3.5/4 and PaLM. The foundation model’s major feature is that it requires far less data to solve a new problem or adapt to a new environment than if it had to learn the job or domain from the start. In this work, they seek a scalable substitute for the time-consuming and expensive process of collecting paired data across three modalities to build a foundation model for long-term planning. Can they do this while still being reasonably effective at solving new planning tasks? 

Researchers from Improbable AI Lab, MIT-IBM Watson AI Lab and Massachusetts Institute Technology suggest Compositional Foundation Models for Hierarchical Planning (HiP), a foundation model made up of many expert models independently trained on language, vision, and action data. The amount of data needed to build the foundation models is significantly decreased since these models are introduced separately (Figure 1). HiP employs a big language model to discover a series of subtasks (i.e., planning) from an abstract language instruction specifying the intended task. HiP then develops a more intricate plan in the form of an observation-only trajectory using a large video diffusion model to gather geometric and physical information about the environment. Finally, HiP employs a sizable inverse model that has been previously trained and converts a series of egocentric pictures into actions. 

Figure 1: Compositional Foundation Models for Hierarchical Planning are shown above. HiP employs three models: a task model (represented by an LLM) to produce an abstract plan, a visual model (represented by a video model) to produce an image trajectory plan; and an ego-centric action model to deduce actions from the image trajectory.

Without needing to gather costly paired decision-making data across modalities, the compositional design choice enables various models to reason at different levels of the hierarchy and jointly make expert conclusions. Three separately trained models can generate conflicting results, which might fail in the whole planning process. For instance, choosing the output with the highest likelihood at each stage is a naive method for building models. A step in a plan, such as looking for a tea kettle in a cabinet, may have a high chance under one model but a zero likelihood under another, such as if the house does not contain a cabinet. Instead, it’s crucial to sample a strategy that jointly maximizes likelihood across all professional models. 

They provide an iterative refinement technique to assure consistency, utilizing feedback from the downstream models to develop consistent plans across their diverse models. The output distribution of the language model’s generative process incorporates intermediate feedback from a likelihood estimator conditioned on a representation of the current state at each stage. Similarly, intermediate input from the action model improves video creation at each stage of the development process. This iterative refinement process fosters consensus across the many models to create hierarchically consistent plans that are both responsive to the objective and executable given the existing state and agent. Their suggested iterative refinement method does not need extensive model finetuning, making training computationally efficient. 

Additionally, they don’t need to know the model’s weights, and their strategy applies to all models that provide input and output API access. In conclusion, they provide a foundation model for hierarchical planning that uses a composition of foundation models independently acquired on various Internet and egocentric robotics data modalities to create long-horizon plans. On three long-horizon tabletop manipulation situations, they show promising outcomes.


Check out the Paper. All Credit For This Research Goes To the Researchers on This Project. Also, don’t forget to join our 30k+ ML SubReddit, 40k+ Facebook Community, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.

If you like our work, you will love our newsletter..


YOU MAY ALSO LIKE

LandingAI Releases Agentic Document Extraction Gen2 with DPT-3 Pro and DPT-3 Verity

Apple Wallet Is Not The Same As Apple Pay: Here’s How They Differ

Aneesh Tickoo is a consulting intern at MarktechPost. He is currently pursuing his undergraduate degree in Data Science and Artificial Intelligence from the Indian Institute of Technology(IIT), Bhilai. He spends most of his time working on projects aimed at harnessing the power of machine learning. His research interest is image processing and is passionate about building solutions around it. He loves to connect with people and collaborate on interesting projects.


🚀 The end of project management by humans (Sponsored)

Credit: Source link

ShareTweetSendSharePin

Related Posts

LandingAI Releases Agentic Document Extraction Gen2 with DPT-3 Pro and DPT-3 Verity
AI & Technology

LandingAI Releases Agentic Document Extraction Gen2 with DPT-3 Pro and DPT-3 Verity

September 10, 2026
Apple Wallet Is Not The Same As Apple Pay: Here’s How They Differ
AI & Technology

Apple Wallet Is Not The Same As Apple Pay: Here’s How They Differ

September 9, 2026
Google Open-Sources Mantis: A Modular Skills Toolkit That Lets Coding Agents Find, Reproduce and Patch Vulnerabilities
AI & Technology

Google Open-Sources Mantis: A Modular Skills Toolkit That Lets Coding Agents Find, Reproduce and Patch Vulnerabilities

September 9, 2026
Muse, The Band, Lost Its Social Media Handles To Muse, Meta’s New AI Agent
AI & Technology

Muse, The Band, Lost Its Social Media Handles To Muse, Meta’s New AI Agent

September 9, 2026
Next Post
Full Show: Best of Bloomberg Technology (08/04)

Full Show: Best of Bloomberg Technology (08/04)

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Your Credit Card’s Purchase Protection Covers What Your Wallet Doesn’t

Your Credit Card’s Purchase Protection Covers What Your Wallet Doesn’t

September 3, 2026
Arab leaders push for de-escalation in Iran war

Arab leaders push for de-escalation in Iran war

September 5, 2026
Expedition to find Amelia Earhart’s plane set to begin

Expedition to find Amelia Earhart’s plane set to begin

September 5, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!