• bitcoinBitcoin(BTC)$77,303.000.48%
  • ethereumEthereum(ETH)$2,512.842.25%
  • tetherTether(USDT)$1.000.02%
  • binancecoinBNB(BNB)$727.671.90%
  • rippleXRP(XRP)$1.361.26%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$102.042.69%
  • tronTRON(TRX)$0.338988-0.55%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.03-0.34%
  • zcashZcash(ZEC)$1,158.355.92%
  • HyperliquidHyperliquid(HYPE)$78.94-0.57%
  • dogecoinDogecoin(DOGE)$0.0842970.91%
  • RainRain(RAIN)$0.015347-2.67%
  • moneroMonero(XMR)$523.272.36%
  • USDSUSDS(USDS)$1.000.01%
  • whitebitWhiteBIT Coin(WBT)$80.210.68%
  • chainlinkChainlink(LINK)$11.540.29%
  • leo-tokenLEO Token(LEO)$9.15-0.40%
  • cardanoCardano(ADA)$0.2068520.08%
  • stellarStellar(XLM)$0.1796662.18%
  • bitcoin-cashBitcoin Cash(BCH)$229.651.43%
  • Ethena USDeEthena USDe(USDE)$1.000.03%
  • daiDai(DAI)$1.000.01%
  • USD1USD1(USD1)$1.000.05%
  • litecoinLitecoin(LTC)$53.331.22%
  • CantonCanton(CC)$0.098185-0.47%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.360.30%
  • uniswapUniswap(UNI)$6.03-0.32%
  • avalanche-2Avalanche(AVAX)$7.47-0.08%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • hedera-hashgraphHedera(HBAR)$0.074544-1.01%
  • shiba-inuShiba Inu(SHIB)$0.0000052.93%
  • nearNEAR Protocol(NEAR)$2.35-6.75%
  • suiSui(SUI)$0.73-0.90%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.056314-0.77%
  • MemeCoreMemeCore(M)$1.204.66%
  • tether-goldTether Gold(XAUT)$4,349.940.36%
  • Circle USYCCircle USYC(USYC)$1.140.03%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • okbOKB(OKB)$113.193.17%
  • BittensorBittensor(TAO)$236.17-0.17%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.00%
  • aaveAave(AAVE)$125.162.39%
  • mantleMantle(MNT)$0.581.27%
  • pax-goldPAX Gold(PAXG)$4,356.750.41%
  • AsterAster(ASTER)$0.68-2.73%
  • polkadotPolkadot(DOT)$1.05-6.47%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.054396-3.50%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Researchers from China Introduce DualToken-ViT: A Fusion of CNNs and Vision Transformers for Enhanced Image Processing Efficiency and Accuracy

October 2, 2023
in AI & Technology
Reading Time: 5 mins read
A A
Researchers from China Introduce DualToken-ViT: A Fusion of CNNs and Vision Transformers for Enhanced Image Processing Efficiency and Accuracy
ShareShareShareShareShare

In recent years, vision transformers (ViTs) have become a potent architecture for various vision applications, including object identification and picture classification. This is because, whereas the size of the convolutional kernel constrains convolutional neural networks (CNNs) and can only extract local information, self-attention can remove global information from the picture, delivering adequate and meaningful visual characteristics. There still needs to be an indication of performance saturation as the size of the dataset and the model for ViTs rise, which is a benefit over CNNs for both big models and huge datasets. Due to several inductive biases ViTs lack, CNNs are preferable over ViTs in lightweight models. 

Self-attention’s quadratic complexity contributes to the potentially high computational cost of ViTs. Consequently, it isn’t easy to build lightweight, effective ViTs. Propose a pyramid structure that separates the model into multiple stages, with the number of tokens reducing and the number of channels growing per stage to construct more effective and lightweight ViTs. Emphasis on streamlining and refining self-attention structure to mitigate its quadratic complexity, but at the expense of attention’s usefulness. A typical strategy is to downsample the key and value of self-attention, which reduces the amount of tokens engaged in the process. 

By conducting self-attention on the grouped tokens independently, certain locally grouped self-attention-based works lower the complexity of the overall attention component. Still, such techniques may harm the sharing of global knowledge. Some efforts additionally provide a few extra teachable parameters to enhance the backbone’s global information, including adding the branch of global tokens used at all stages. Local attention techniques like locally grouped self-attention-based and convolution-based structures can be enhanced using this method. However, all existing international token approaches only consider global information and disregard positional information, which is crucial for vision tasks. 

Figure 1: A visualization of the attention map for the key token (the most crucial component of the picture for the image classification challenge) and position-aware global tokens. The first picture in each row serves as the model’s input, while the second image depicts the correlation between each token in the position-aware global tokens, which each comprise seven tokens, where the red-boxed section is the first image’s key token.

In this study, researchers from East China Normal University and Alibaba Group put forth the DualToken-ViT, a compact and effective vision transformer model. Their suggested paradigm replaces self-attention with a more effective attentional framework. Convolution and self-attention are used together to extract local and global information. The outputs of the two processes are then fused to create an effective attention structure. Although window self-attention may also remove local information, they find that their lightweight model’s convolution is more effective. They step-wise downsample the feature map that creates key and value to retain more information throughout the downsampling process. This can lower the computational cost of self-attention in global information broadcasting. 

Additionally, they employ position-aware global tokens at every level to improve global data quality. Their position-aware global tokens can also maintain and pass on picture location information, providing their model an edge in vision tasks, in contrast to the standard global tokens. The efficacy of their position-aware global tokens is seen in Figure 1, where the key token in the image produces a greater correlation with the equivalent tokens in the position-aware global tokens. 

In a nutshell, their contributions are as follows: 

• They develop a compact and effective vision transformer model called DualToken-ViT, which fuses local and global tokens containing local and global information, respectively, to achieve an efficient attention structure by combining the benefits of convolution and self-attention. 

• They also suggest position-aware global tokens, which would expand the global information by including the image’s location data. 

• Their DualToken-ViT exhibits the greatest performance on image classification, object identification, and semantic segmentation among vision models of the same FLOPs magnitude.


Check out the Paper. All Credit For This Research Goes To the Researchers on This Project. Also, don’t forget to join our 31k+ ML SubReddit, 40k+ Facebook Community, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.

If you like our work, you will love our newsletter..


YOU MAY ALSO LIKE

Oracle’s Cloud Growth; Debate Around AI Risks

Balancing AI Risks With the Race to Stay Ahead of China

Aneesh Tickoo is a consulting intern at MarktechPost. He is currently pursuing his undergraduate degree in Data Science and Artificial Intelligence from the Indian Institute of Technology(IIT), Bhilai. He spends most of his time working on projects aimed at harnessing the power of machine learning. His research interest is image processing and is passionate about building solutions around it. He loves to connect with people and collaborate on interesting projects.


🚀 The end of project management by humans (Sponsored)

Credit: Source link

ShareTweetSendSharePin

Related Posts

Oracle’s Cloud Growth; Debate Around AI Risks
AI & Technology

Oracle’s Cloud Growth; Debate Around AI Risks

September 12, 2026
Balancing AI Risks With the Race to Stay Ahead of China
AI & Technology

Balancing AI Risks With the Race to Stay Ahead of China

September 12, 2026
Microsoft’s Data Center Plans Face Big Costs
AI & Technology

Microsoft’s Data Center Plans Face Big Costs

September 12, 2026
Baseten Adds DeepSeek-V4.1-Flash to Model APIs With 1M-Token Context – Unite.AI
AI & Technology

Baseten Adds DeepSeek-V4.1-Flash to Model APIs With 1M-Token Context – Unite.AI

September 11, 2026
Next Post
Tech Companies Said to Circulate Immigration Open Letter

Tech Companies Said to Circulate Immigration Open Letter

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Saudi Arabian oil pipeline system hit by projectiles triggering fires – CNN

Saudi Arabian oil pipeline system hit by projectiles triggering fires – CNN

September 11, 2026
Brent crude rises above 0 a barrel as Middle East conflict escalates – Reuters

Brent crude rises above $100 a barrel as Middle East conflict escalates – Reuters

September 9, 2026
I’m Losing My 0,000 Job During a Home Renovation

I’m Losing My $100,000 Job During a Home Renovation

September 9, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!