• bitcoinBitcoin(BTC)$84,323.000.63%
  • ethereumEthereum(ETH)$2,686.680.58%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$775.440.70%
  • rippleXRP(XRP)$1.543.56%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$117.152.55%
  • tronTRON(TRX)$0.338412-0.82%
  • zcashZcash(ZEC)$1,584.904.55%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.03-0.75%
  • HyperliquidHyperliquid(HYPE)$93.551.69%
  • dogecoinDogecoin(DOGE)$0.0960182.80%
  • moneroMonero(XMR)$571.711.99%
  • chainlinkChainlink(LINK)$13.7211.53%
  • whitebitWhiteBIT Coin(WBT)$84.050.11%
  • USDSUSDS(USDS)$1.000.00%
  • cardanoCardano(ADA)$0.2516525.38%
  • RainRain(RAIN)$0.011896-1.73%
  • leo-tokenLEO Token(LEO)$8.81-1.45%
  • stellarStellar(XLM)$0.2187318.65%
  • bitcoin-cashBitcoin Cash(BCH)$334.850.70%
  • nearNEAR Protocol(NEAR)$4.629.25%
  • uniswapUniswap(UNI)$9.220.77%
  • litecoinLitecoin(LTC)$70.873.99%
  • Ethena USDeEthena USDe(USDE)$1.000.02%
  • CantonCanton(CC)$0.1183608.46%
  • daiDai(DAI)$1.000.02%
  • avalanche-2Avalanche(AVAX)$10.331.68%
  • USD1USD1(USD1)$1.000.00%
  • suiSui(SUI)$1.037.31%
  • hedera-hashgraphHedera(HBAR)$0.0927452.60%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.420.42%
  • BittensorBittensor(TAO)$303.346.07%
  • shiba-inuShiba Inu(SHIB)$0.0000062.48%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • crypto-com-chainCronos(CRO)$0.0647725.10%
  • BitwayBitway(BTW)$1.065.08%
  • OndoOndo(ONDO)$0.5631.01%
  • MemeCoreMemeCore(M)$1.21-2.96%
  • paypal-usdPayPal USD(PYUSD)$1.000.01%
  • tether-goldTether Gold(XAUT)$4,281.890.32%
  • okbOKB(OKB)$119.910.33%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.19%
  • EthenaEthena(ENA)$0.2227659.42%
  • mantleMantle(MNT)$0.68-2.51%
  • aaveAave(AAVE)$145.055.18%
  • MorphoMorpho(MORPHO)$2.896.97%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

InternLM-XComposer-2.5 (IXC-2.5): A Versatile Large-Vision Language Model that Supports Long-Contextual Input and Output

July 13, 2024
in AI & Technology
Reading Time: 5 mins read
A A
InternLM-XComposer-2.5 (IXC-2.5): A Versatile Large-Vision Language Model that Supports Long-Contextual Input and Output
ShareShareShareShareShare

Large Language Models (LLMs) have made significant strides in recent years, prompting researchers to explore the development of Large Vision Language Models (LVLMs). These models aim to integrate visual and textual information processing capabilities. However, current open-source LVLMs face challenges in matching the versatility of proprietary models like GPT-4, Gemini Pro, and Claude 3. The primary obstacles include limited diversity in training data and difficulties in handling long-context input and output. Researchers are striving to enhance open-source LVLMs’ ability to perform a wide range of vision-language comprehension and composition tasks, bridging the gap between open-source and closed-source leading paradigms in terms of versatility and performance across various benchmarks.

Researchers have made significant efforts to tackle the challenges in developing versatile LVLMs. These approaches include text-image conversation models, high-resolution image analysis techniques, and video understanding methods. For text-image conversations, most existing LVLMs focus on single-image multi-round interactions, with some extending to multi-image inputs. High-resolution image analysis has been tackled through two main strategies: high-resolution visual encoders and image patchification. Video understanding in LVLMs has employed techniques such as sparse sampling, temporal pooling, compressed video tokens, and memory banks.

YOU MAY ALSO LIKE

Fastino Releases GLiNER2.5-Decide: A 340M Open-Weight Decision Model That Runs on CPU

Black Forest Labs Releases FLUX 3 Action: A 7B Open-Weights World Action Model That Tops RoboLab-120

Also, researchers have explored webpage generation, moving from simple UI-to-code transformations to more complex tasks using large vision-language models trained on synthetic datasets. However, these approaches often lack diversity and real-world applicability. To align model outputs with human preferences, techniques like Reinforcement Learning from Human Feedback (RLHF) and Direct Preference Optimization (DPO) have been adapted for multimodal LVLMs, focusing on reducing hallucinations and improving response quality.

Researchers from Shanghai Artificial Intelligence Laboratory, The Chinese University of Hong Kong, SenseTime Group, and Tsinghua University have introduced InternLM-XComposer-2.5 (IXC-2.5), representing a significant advancement in LVLMs, offering versatility and long-context capabilities. This model excels in comprehension and composition tasks, including free-form text-image conversations, OCR, video understanding, article composition, and webpage crafting. IXC-2.5 supports a 24K interleaved image-text context window, extendable to 96K, enabling long-term human-AI interaction and content creation.

The model introduces three key comprehension upgrades: ultra-high resolution understanding, fine-grained video analysis, and multi-turn multi-image dialogue support. For composition tasks, IXC-2.5 incorporates additional LoRA parameters, enabling webpage creation and high-quality text-image article composition. The latter benefits from Chain-of-Thought and Direct Preference Optimization techniques to enhance content quality.

IXC-2.5 enhances its predecessors’ architecture with a ViT-L/14 Vision Encoder, InternLM2-7B Language Model, and Partial LoRA. It handles diverse inputs through a Unified Dynamic Image Partition strategy, processing images at 560×560 resolution with 400 tokens per sub-image. The model employs a scaled identity strategy for high-resolution images and treats videos as concatenated frames. Multi-image inputs are handled with interleaved formatting. IXC-2.5 also supports audio input/output using Whisper for transcription and MeloTTS for speech synthesis. This versatile architecture enables effective processing of various input types and complex tasks.

IXC-2.5 demonstrates exceptional performance across various benchmarks. In video understanding, it outperforms open-source models in 4 out of 5 benchmarks, matching closed-source APIs. For structural high-resolution tasks, IXC-2.5 competes with larger models, excelling in form and table understanding. It significantly improves multi-image multi-turn comprehension, outperforming previous models by 13.8% on the MMDU benchmark. In general visual QA tasks, IXC-2.5 matches or surpasses both open-source and closed-source models, notably outperforming GPT-4V and Gemini-Pro on some challenges. For screenshot-to-code translation, IXC-2.5 even surpasses GPT-4V in average performance, showcasing its versatility and effectiveness across diverse multimodal tasks.

IXC-2.5 represents a significant advancement in Large Vision-Language Models, offering long-contextual input and output capabilities. This model excels in ultra-high resolution image analysis, fine-grained video comprehension, multi-turn multi-image dialogues, webpage generation, and article composition. Despite utilizing a modest 7B Large Language Model backend, IXC-2.5 demonstrates competitive performance across various benchmarks. This achievement paves the way for future research into more contextual multi-modal environments, potentially extending to long-context video understanding and interaction history analysis. Such advancements promise to enhance AI’s capacity to assist humans in diverse real-world applications, marking a crucial step forward in multimodal AI technology.


Check out the Paper and GitHub. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. 

Join our Telegram Channel and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 46k+ ML SubReddit


Asjad is an intern consultant at Marktechpost. He is persuing B.Tech in mechanical engineering at the Indian Institute of Technology, Kharagpur. Asjad is a Machine learning and deep learning enthusiast who is always researching the applications of machine learning in healthcare.

🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

Fastino Releases GLiNER2.5-Decide: A 340M Open-Weight Decision Model That Runs on CPU
AI & Technology

Fastino Releases GLiNER2.5-Decide: A 340M Open-Weight Decision Model That Runs on CPU

September 25, 2026
Black Forest Labs Releases FLUX 3 Action: A 7B Open-Weights World Action Model That Tops RoboLab-120
AI & Technology

Black Forest Labs Releases FLUX 3 Action: A 7B Open-Weights World Action Model That Tops RoboLab-120

September 25, 2026
Warzone Is Adding A Button To Hide All The Goofy Skins
AI & Technology

Warzone Is Adding A Button To Hide All The Goofy Skins

September 24, 2026
How These AI Glasses Compare
AI & Technology

How These AI Glasses Compare

September 24, 2026
Next Post
Hunter Biden found guilty on all counts in federal gun case

Hunter Biden found guilty on all counts in federal gun case

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
U.S. strikes Iran for first time in weeks

U.S. strikes Iran for first time in weeks

September 20, 2026
Dolly Parton’s legacy as a philanthropist

Dolly Parton’s legacy as a philanthropist

September 23, 2026
Humanoid robot fails weightlifting test at Beijing games

Humanoid robot fails weightlifting test at Beijing games

September 25, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!