• bitcoinBitcoin(BTC)$79,124.001.59%
  • ethereumEthereum(ETH)$2,502.441.83%
  • tetherTether(USDT)$1.000.02%
  • binancecoinBNB(BNB)$745.07-0.13%
  • rippleXRP(XRP)$1.431.97%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$103.821.24%
  • tronTRON(TRX)$0.338467-0.07%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.040.00%
  • zcashZcash(ZEC)$1,270.729.81%
  • HyperliquidHyperliquid(HYPE)$86.524.80%
  • dogecoinDogecoin(DOGE)$0.0903261.44%
  • RainRain(RAIN)$0.016316-1.54%
  • USDSUSDS(USDS)$1.000.01%
  • whitebitWhiteBIT Coin(WBT)$81.813.35%
  • moneroMonero(XMR)$502.690.57%
  • chainlinkChainlink(LINK)$12.13-2.36%
  • leo-tokenLEO Token(LEO)$9.18-0.02%
  • cardanoCardano(ADA)$0.2183310.06%
  • stellarStellar(XLM)$0.1880100.15%
  • bitcoin-cashBitcoin Cash(BCH)$257.871.57%
  • daiDai(DAI)$1.000.01%
  • Ethena USDeEthena USDe(USDE)$1.000.02%
  • USD1USD1(USD1)$1.00-0.01%
  • litecoinLitecoin(LTC)$54.14-1.15%
  • CantonCanton(CC)$0.1050791.46%
  • uniswapUniswap(UNI)$6.62-4.01%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.390.44%
  • hedera-hashgraphHedera(HBAR)$0.078430-1.53%
  • avalanche-2Avalanche(AVAX)$7.950.47%
  • nearNEAR Protocol(NEAR)$2.5812.47%
  • suiSui(SUI)$0.811.00%
  • Global DollarGlobal Dollar(USDG)$1.000.01%
  • shiba-inuShiba Inu(SHIB)$0.0000050.41%
  • crypto-com-chainCronos(CRO)$0.0598560.15%
  • paypal-usdPayPal USD(PYUSD)$1.000.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,411.140.38%
  • MemeCoreMemeCore(M)$1.180.68%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BittensorBittensor(TAO)$261.694.16%
  • Ripple USDRipple USD(RLUSD)$1.000.02%
  • okbOKB(OKB)$114.050.03%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.38%
  • mantleMantle(MNT)$0.642.81%
  • AsterAster(ASTER)$0.75-0.39%
  • aaveAave(AAVE)$129.401.36%
  • polkadotPolkadot(DOT)$1.167.95%
  • Pump.funPump.fun(PUMP)$0.0046356.38%
  • pax-goldPAX Gold(PAXG)$4,414.410.35%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Microsoft Research Introduces Florence-2: A Novel Vision Foundation Model with a Unified Prompt-based Representation for a Variety of Computer Vision and Vision-Language Tasks

November 23, 2023
in AI & Technology
Reading Time: 5 mins read
A A
Microsoft Research Introduces Florence-2: A Novel Vision Foundation Model with a Unified Prompt-based Representation for a Variety of Computer Vision and Vision-Language Tasks
ShareShareShareShareShare

There has been a noticeable trend in Artificial General Intelligence (AGI) systems toward using pre-trained, adaptable representations, which provide task-agnostic advantages in various applications. Natural language processing (NLP) is a good example of this tendency since sophisticated models demonstrate flexibility with thorough knowledge covering several domains and tasks with straightforward instructions. The popularity of NLP encourages a complementary strategy in computer vision. Unique obstacles arise from the necessity for broad perceptual capacities in universal representation for various vision-related activities. Whereas natural language processing (NLP) focuses mostly on text, computer vision has to handle complex visual data such as characteristics, masked contours, and object placement. In computer vision, achieving universal representation necessitates skillful handling of various challenging tasks arranged in two dimensions, as shown in Figure 1. 

Figure 1

Spatial Hierarchy: The model has to recognize spatial information at different sizes, comprehending fine-grained pixel details and image-level ideas. To support the complex spatial hierarchy in vision, the model must be capable of managing a range of granularities.

Semantic Granularity: In computer vision, universal representation should cover a range of semantic granularities. The paradigm moves from abstract titles to more detailed explanations, providing flexible comprehension for various uses. 

This pursuit is characterized by distinctiveness and substantial challenges. A key hurdle is the need for more, hindering the development of a foundational model capable of capturing the intricate nuances of spatial hierarchy and semantic granularity. Existing datasets, such as ImageNet, COCO, and Flickr30k Entities, tailored for specialized applications, are extensively labeled by humans. To overcome this constraint, it is imperative to generate extensive annotations for each image on a larger scale. Another challenge is the absence of a that seamlessly integrates spatial hierarchy and semantic granularity in computer vision. With task-specific design, traditional models perform well in tasks like semantic segmentation, object identification, and picture captioning. However, creating a complete, cohesive model that can adjust to different vision tasks in a task-independent way is crucial, even taking on new duties with little to no task-specific fine-tuning.

Through unified pre-training and network design, the model pioneers the integration of spatial, temporal, and multi-modal features in computer vision. The first evolutionary iteration excels in transfer learning through task-specific fine-tuning using customized adapters and pre-training with noisy text-image pairings. However, its reliance on big task-specific datasets and adapters results in gaps when it comes to tackling the two major issues mentioned above. In this work, researchers from Azure provide a universal backbone that is attained using multitask learning with rich visual annotations. This leads to a prompt-based, unified representation for various vision tasks, which successfully tackles the issues of incomplete comprehensive data and lack of a uniform architecture.

Large-scale, high-quality annotated data is necessary for multitask learning. Rather than depending on time-consuming human annotation, their data engine creates an extensive visual dataset named \fld, which has 5.4B annotations for 126M photos. There are two effective processing modules in this engine. The first module departs from the conventional single and manual annotation strategy by using specialized models to annotate photos jointly and autonomously. Similar to the wisdom of crowds theory, many models collaborate to create a consensus, resulting in a more impartial and trustworthy picture interpretation. Using basic models that have been learned, the second module repeatedly refines and filters these automatic annotations.

Their model uses a sequence-to-sequence (seq2seq) architecture, integrating an image encoder and a multi-modality encoder-decoder by leveraging this large dataset. This architecture supports a range of vision tasks without requiring task-specific architectural adjustments, in line with the NLP community’s goal of flexible model creation with a uniform foundation. Every annotation in the dataset is consistently standardized into textual outputs. This enables the consistent optimization of a single multitask learning strategy using the same loss function as the goal. The result is a flexible vision foundation model, or model, that can handle a range of functions, including object recognition, captioning, and grounding, all under the control of a single model with standardized parameters. Textual prompts are utilized to activate tasks, consistent with the methodology employed by large language models (LLMs).

Their method achieves a universal representation and has wide-ranging use in many visual tasks. Key findings consist of:

  • The model is a flexible vision foundation model that provides new state-of-the-art zero-shot performance in tasks, including referencing expression comprehension on RefCOCO, visual grounding on Flick30k, and captioning on COCO.
  • Notwithstanding its small size, it competes with more specialized models after being fine-tuned using publicly available human-annotated data. Most notably, the improved model sets new benchmark state-of-the-art scores on RefCOCO.
  • The pre-trained backbone outperforms supervised and self-supervised models on downstream tasks, COCO object detection and instance segmentation, and ADE20K semantic segmentation. Their model, which uses the Mask-RCNN, DINO, and UperNet frameworks, delivers significant increases of 6.9, 5.5, and 5.9 points on COCO and ADE20K datasets, respectively and quadruples the training efficiency of pre-trained models on ImageNet.

Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to join our 33k+ ML SubReddit, 41k+ Facebook Community, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.

If you like our work, you will love our newsletter..


YOU MAY ALSO LIKE

Harvey Secures $550M in Fresh Funding, Valuation Climbs to $15.5B – Unite.AI

How To Take Full Advantage Of Gemini When Planning Your Next Trip

Aneesh Tickoo is a consulting intern at MarktechPost. He is currently pursuing his undergraduate degree in Data Science and Artificial Intelligence from the Indian Institute of Technology(IIT), Bhilai. He spends most of his time working on projects aimed at harnessing the power of machine learning. His research interest is image processing and is passionate about building solutions around it. He loves to connect with people and collaborate on interesting projects.


↗ Step by Step Tutorial on ‘How to Build LLM Apps that can See Hear Speak’

Credit: Source link

ShareTweetSendSharePin

Related Posts

Harvey Secures 0M in Fresh Funding, Valuation Climbs to .5B – Unite.AI
AI & Technology

Harvey Secures $550M in Fresh Funding, Valuation Climbs to $15.5B – Unite.AI

September 9, 2026
How To Take Full Advantage Of Gemini When Planning Your Next Trip
AI & Technology

How To Take Full Advantage Of Gemini When Planning Your Next Trip

September 9, 2026
Will We See The Foldable iPhone Ultra At The ‘Surprise And Shine’ Keynote Today?
AI & Technology

Will We See The Foldable iPhone Ultra At The ‘Surprise And Shine’ Keynote Today?

September 9, 2026
Gradium Launches Voice Design: Write a Prompt, Get a Brand New Synthetic Voice in Seconds
AI & Technology

Gradium Launches Voice Design: Write a Prompt, Get a Brand New Synthetic Voice in Seconds

September 9, 2026
Next Post
A timeline of the shootings in Lewiston, Maine

A timeline of the shootings in Lewiston, Maine

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
3 Big ChatGPT Updates You Need to Know

3 Big ChatGPT Updates You Need to Know

September 8, 2026
Wildberries says warehouses struck in drone attack

Wildberries says warehouses struck in drone attack

September 5, 2026
Urgent search for missing American in Caribbean

Urgent search for missing American in Caribbean

September 4, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!