• bitcoinBitcoin(BTC)$77,388.000.17%
  • ethereumEthereum(ETH)$2,536.492.95%
  • tetherTether(USDT)$1.000.01%
  • binancecoinBNB(BNB)$726.891.55%
  • rippleXRP(XRP)$1.360.76%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$102.582.54%
  • tronTRON(TRX)$0.338436-0.23%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.030.00%
  • zcashZcash(ZEC)$1,179.424.44%
  • HyperliquidHyperliquid(HYPE)$80.830.47%
  • dogecoinDogecoin(DOGE)$0.0844900.31%
  • RainRain(RAIN)$0.015589-1.79%
  • USDSUSDS(USDS)$1.000.00%
  • moneroMonero(XMR)$519.141.24%
  • whitebitWhiteBIT Coin(WBT)$80.480.63%
  • chainlinkChainlink(LINK)$11.610.02%
  • leo-tokenLEO Token(LEO)$9.15-0.49%
  • cardanoCardano(ADA)$0.206489-1.44%
  • stellarStellar(XLM)$0.1790310.63%
  • Ethena USDeEthena USDe(USDE)$1.000.03%
  • bitcoin-cashBitcoin Cash(BCH)$229.451.00%
  • daiDai(DAI)$1.000.00%
  • USD1USD1(USD1)$1.000.04%
  • litecoinLitecoin(LTC)$53.692.49%
  • CantonCanton(CC)$0.098450-0.61%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.371.39%
  • uniswapUniswap(UNI)$6.080.18%
  • avalanche-2Avalanche(AVAX)$7.47-1.84%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • hedera-hashgraphHedera(HBAR)$0.074698-1.26%
  • nearNEAR Protocol(NEAR)$2.49-1.05%
  • shiba-inuShiba Inu(SHIB)$0.0000051.40%
  • suiSui(SUI)$0.73-1.64%
  • paypal-usdPayPal USD(PYUSD)$1.000.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.0565620.04%
  • MemeCoreMemeCore(M)$1.193.02%
  • tether-goldTether Gold(XAUT)$4,343.220.50%
  • Circle USYCCircle USYC(USYC)$1.140.03%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • okbOKB(OKB)$113.782.36%
  • BittensorBittensor(TAO)$236.30-1.40%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.24%
  • mantleMantle(MNT)$0.581.90%
  • aaveAave(AAVE)$124.941.56%
  • pax-goldPAX Gold(PAXG)$4,350.320.63%
  • AsterAster(ASTER)$0.68-3.09%
  • polkadotPolkadot(DOT)$1.05-5.00%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.054471-2.91%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

This AI Research from China Introduces LLaVA-Phi: A Vision Language Assistant Developed Using the Compact Language Model Phi-2

January 10, 2024
in AI & Technology
Reading Time: 4 mins read
A A
This AI Research from China Introduces LLaVA-Phi: A Vision Language Assistant Developed Using the Compact Language Model Phi-2
ShareShareShareShareShare

Large language models have shown notable achievements in executing instructions, multi-turn conversations, and image-based question-answering tasks. These models include Flamingo, GPT-4V, and Gemini. The fast development of open-source Large Language Models, such as LLaMA and Vicuna, has greatly accelerated the evolution of open-source vision language models. These advancements mainly center on improving visual understanding by utilizing language models with at least 7B parameters and integrating them with a vision encoder. Autonomous driving and robotics are two examples of time-sensitive or real-time interactive applications that could benefit from a faster inference speed and shorter test times.

Regarding mobile technology, Gemini has been a trailblazer for multimodal approaches. Gemini-Nano, a simplified version, contains 1.8/3.25 billion parameters and can be used on mobile devices. Yet, information such as the model’s design, training datasets, and training procedures is confidential and cannot be shared with anybody.

YOU MAY ALSO LIKE

New Images Show A Detailed View Of Meta’s Upcoming Mixed Reality Headset

Dzmitry Lazerka, Co-Founder of VictoriaMetrics – Interview Series – Unite.AI

A new study by Midea Group and East China Normal University provides LLaVA-Phi, a little language model-powered vision-language assistant. The most effective open-sourced tiny language model, Phi-2.2, and the robust open-sourced multimodal model, LLaVA-1.5, are combined in this study. The researchers use LLaVA’s high-quality visual instruction tuning data in a two-stage training pipeline. They tested LLaVA-Phi using eight different metrics.

Its performance is on par with, or even better than, other three times larger multimodal models, and it only has three billion parameters. 

The team used a wide variety of academic standards developed for multimodal models to thoroughly evaluate LLaVA-Phi. Examples of these tests include VQA-v2, VizWizQA, ScienceQA, and TextQA for general question-answering and more specialized assessments like POPE for object hallucination and MME, MMBench, and MMVet for a comprehensive evaluation of diverse multimodal abilities like visual understanding and visual commonsense reasoning. The proposed method outperformed other big multimodal models that were previously available by demonstrating that the model could answer questions based on visual cues. Amazingly, LLaVA-Phi achieved better results than models like IDEFICS, which rely on a 7B-parameter or greater LLMs. 

The top score the model achieved on ScienceQA stands out. The success of their multimodal model in answering math-based questions can be attributed to the Phi-2 language model, which has been trained on mathematical corpora and code production in particular. In the extensive multimodal benchmark of MMBench, LLaVA-Phi outperformed numerous prior art vision-language models based on 7B-LLM. 

Another parallel effort that constructs an effective vision-language model, MobileVLM, was also compared. LLaVA-Phi routinely beats all the approaches on all five measures.

The team highlights that since the model has not been fine-tuned to follow multilingual instructions, the LLaVA-Phi architecture cannot process instructions in various languages, including Chinese, because Phi-2 uses the codegenmono tokenizer. They intend to improve training procedures for small language models in the future and investigate the effect of visual encoder size, looking at methods like RLHF and direct preference optimization. These endeavors aim to further improve performance while decreasing model size.


Check out the Paper and Github. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. Join our 35k+ ML SubReddit, 41k+ Facebook Community, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..


Dhanshree Shenwai is a Computer Science Engineer and has a good experience in FinTech companies covering Financial, Cards & Payments and Banking domain with keen interest in applications of AI. She is enthusiastic about exploring new technologies and advancements in today’s evolving world making everyone’s life easy.


🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

New Images Show A Detailed View Of Meta’s Upcoming Mixed Reality Headset
AI & Technology

New Images Show A Detailed View Of Meta’s Upcoming Mixed Reality Headset

September 11, 2026
Dzmitry Lazerka, Co-Founder of VictoriaMetrics – Interview Series – Unite.AI
AI & Technology

Dzmitry Lazerka, Co-Founder of VictoriaMetrics – Interview Series – Unite.AI

September 11, 2026
Where Should Apple Go After The iPhone Duo? Bring On Smaller And Larger Foldables
AI & Technology

Where Should Apple Go After The iPhone Duo? Bring On Smaller And Larger Foldables

September 11, 2026
Why Falling AI Prices Aren’t Lowering Enterprise AI Bills – Unite.AI
AI & Technology

Why Falling AI Prices Aren’t Lowering Enterprise AI Bills – Unite.AI

September 11, 2026
Next Post
This Paper Explores How Deep Learning Enhances Osteoporosis Screening with Routine CT Scans

This Paper Explores How Deep Learning Enhances Osteoporosis Screening with Routine CT Scans

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
A Silicon Valley company with Eric Trump as an advisor is making robot soldiers

A Silicon Valley company with Eric Trump as an advisor is making robot soldiers

September 6, 2026
Google and NASA JPL Unveil AI Model Mapping Global Methane Plumes – Unite.AI

Google and NASA JPL Unveil AI Model Mapping Global Methane Plumes – Unite.AI

September 9, 2026
Seahawks vs. Patriots live updates: Seattle faces New England in rematch of Super Bowl LX to kick off season – CBS Sports

Seahawks vs. Patriots live updates: Seattle faces New England in rematch of Super Bowl LX to kick off season – CBS Sports

September 10, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!