• bitcoinBitcoin(BTC)$80,966.00-0.45%
  • ethereumEthereum(ETH)$2,613.69-0.14%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$759.90-0.42%
  • rippleXRP(XRP)$1.40-1.29%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$109.79-3.17%
  • tronTRON(TRX)$0.3403120.57%
  • zcashZcash(ZEC)$1,462.75-6.11%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.02-1.50%
  • HyperliquidHyperliquid(HYPE)$91.22-2.50%
  • dogecoinDogecoin(DOGE)$0.087016-1.42%
  • moneroMonero(XMR)$541.48-6.40%
  • whitebitWhiteBIT Coin(WBT)$82.57-0.92%
  • RainRain(RAIN)$0.0136952.19%
  • USDSUSDS(USDS)$1.00-0.02%
  • chainlinkChainlink(LINK)$12.28-1.15%
  • cardanoCardano(ADA)$0.226713-1.70%
  • leo-tokenLEO Token(LEO)$8.910.06%
  • stellarStellar(XLM)$0.193755-0.53%
  • uniswapUniswap(UNI)$8.82-2.33%
  • bitcoin-cashBitcoin Cash(BCH)$250.41-2.69%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • nearNEAR Protocol(NEAR)$3.55-7.07%
  • daiDai(DAI)$1.000.00%
  • litecoinLitecoin(LTC)$57.91-1.42%
  • USD1USD1(USD1)$1.00-0.01%
  • avalanche-2Avalanche(AVAX)$9.6613.01%
  • CantonCanton(CC)$0.107624-3.64%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.391.14%
  • hedera-hashgraphHedera(HBAR)$0.0813252.35%
  • suiSui(SUI)$0.863.69%
  • MemeCoreMemeCore(M)$1.5317.02%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • shiba-inuShiba Inu(SHIB)$0.000005-0.01%
  • crypto-com-chainCronos(CRO)$0.059809-0.94%
  • BittensorBittensor(TAO)$261.713.24%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.02%
  • tether-goldTether Gold(XAUT)$4,370.58-0.06%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • okbOKB(OKB)$116.70-0.73%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.15-0.24%
  • aaveAave(AAVE)$141.04-1.37%
  • EthenaEthena(ENA)$0.21314621.92%
  • AsterAster(ASTER)$0.76-2.20%
  • mantleMantle(MNT)$0.62-0.60%
  • OndoOndo(ONDO)$0.4131150.20%
  • Pump.funPump.fun(PUMP)$0.004111-3.24%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Microsoft Research Introduces Gigapath: A Novel Vision Transformer For Digital Pathology

May 26, 2024
in AI & Technology
Reading Time: 5 mins read
A A
Microsoft Research Introduces Gigapath: A Novel Vision Transformer For Digital Pathology
ShareShareShareShareShare

Digital pathology converts traditional glass slides into digital images for viewing, analysis, and storage. Advances in imaging technology and software drive this transformation, which has significant implications for medical diagnostics, research, and education. There is a chance to speed up advancements in precision health by a factor of ten because of the present generative AI revolution and the parallel digital change in biomedicine. To generate evidence at the population level, digital pathology can be integrated with other multimodal, longitudinal patient data in multimodal generative AI.

A typical gigapixel slide maybe thousands of times wider and longer than ordinary natural images, which is both an exciting development and a sobering reminder that digital pathology has distinct computing hurdles. This massive magnitude is too much for traditional vision transformers to handle since the computation required for self-attention increases substantially as the input length does. Therefore, previous digital pathology work frequently fails to account for the complex interdependencies between picture tiles on each slide, leaving out crucial context at the slide level for vital applications like tumor microenvironment modeling.

Microsoft’s innovative vision transformer GigaPath has revolutionized whole-slide modeling by introducing dilated self-attention to manage computation. In a groundbreaking collaboration with Providence Health System and the University of Washington, they have developed Prov-GigaPath, an open-access whole-slide pathology foundation model. This model, pretrained on a staggering one billion 256 X 256 pathology image tiles from over 170,000 whole slides using real-world data from Providence, is a significant step forward in the field. All computations were conducted at the private tenant of Providence, with the approval of the Providence Institutional Review Board (IRB).

GigaPath’s two-stage curriculum learning involves pretraining at the tile level with DINOv2 and pre-training at the slide level using masked autoencoder and LongNet. The DINOv2 self-supervision method integrates masked reconstruction loss and contrastive loss to train student and teacher vision transformers. However, because self-attention is computationally challenging, it can only be used for small images like 256 × 256 tiles. The researchers modified LongNet’s dilated attention to digital pathology for slide-level modeling. They introduce a series of rising sizes for segmenting the tile sequence into pieces of the specified size to manage the lengthy sequence of image tiles for an entire presentation. They implement sparse attention for longer segments to counteract the quadratic expansion, where sparsity is directly proportional to segment length. The biggest part would encompass the whole slide but with sparsely subsampled self-attention. As a result, they keep computation tractable while systematically capturing long-range relationships.

By leveraging data from both the Providence and TCGA datasets, the team has established a digital pathology standard with nine tasks for cancer subtyping and seventeen tasks for pathomics. Through large-scale pretraining and whole-slide modeling, Prov-GigaPath has demonstrated exceptional performance, outperforming the second-best model on 18 of the 26 tasks by a significant margin. This impressive performance underscores the model’s versatility and potential for a wide range of applications in digital pathology. 

The objective of cancer subtyping is to categorize specific subtypes using pathology slides. Pathomics tasks aim to classify tumors as having particular therapeutically important genetic alterations using only the slide image. This has the potential to reveal significant associations between genetic pathways and tissue morphology that are too subtle to be detected by human eyes. 

In addition, the team aims to find universal signals for a gene mutation across all cancer kinds, and extremely different tumor morphologies in some studies called the pan-cancer scenario. Under these difficult conditions, Prov-GigaPath achieved state-of-the-art performance on 17 of 18 tasks, surpassing the second-best on 12 of 18 jobs. For instance, in the pan-cancer 5-gene study, Prov-GigaPath achieved a 6.5% improvement in AUROC and an 18.7% improvement in AUPRC compared to the top competing approaches. 

To further evaluate the Prov-generalizability GigaPath, researchers again ran a head-to-head comparison using TCGA data, dominating all other approaches. Evidence of the biological significance of the learned embeddings and the ability of Prov-Gigapath to extract genetically linked pan-cancer and subtype-specific morphological features at the whole-slide level pave the way for the utilization of real-world data in future research about the intricate biology of the tumor microenvironment.

Incorporating the pathology data further demonstrates GigaPath’s capability to perform vision-language tasks. The researchers utilize the report semantics to align the pathology slide representation and continue pretraining on these pairs. Then, they use them for downstream prediction tasks like zero-shot subtyping that don’t require supervised fine-tuning. They perform contrastive learning utilizing the slide-report pairings, with Prov-GigaPath as the whole-slide image encoder and PubMedBERT as the text encoder. This is far more difficult than conventional vision-language pretraining without granular alignment information between specific picture tiles and text samples. On benchmark vision-language tasks, including zero-shot cancer subtyping and gene mutation prediction, Prov-GigaPath significantly beats three top-tier pathology models, highlighting its promise for whole-slide vision-language modeling.

According to the researchers, this is the first digital pathology foundation model that has undergone large-scale pretraining on real-world data. Standard cancer classification, pathomics, and vision-language tasks are all areas where Prov-GigaPath achieves state-of-the-art performance. This offers new avenues for improving patient care and speeding up clinical discovery, and it also shows how important whole-slide modeling is on large-scale real-world data. They highlight that extensive work remains before one can fully realize the promise of a multimodal conversational assistant, particularly in integrating state-of-the-art multimodal frameworks like LLaVA-Med.


Check out the Paper and Blog. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter. Join our Telegram Channel, Discord Channel, and LinkedIn Group.

If you like our work, you will love our newsletter..

Don’t Forget to join our 42k+ ML SubReddit


YOU MAY ALSO LIKE

How Long Can You Expect Your Old Cassette Tapes To Last?

OpenClaw Releases 2026.9.5 With Atomic Updates, Plugin Hot Reload, Conversation Sharing, and Expanded GPT Live

Dhanshree Shenwai is a Computer Science Engineer and has a good experience in FinTech companies covering Financial, Cards & Payments and Banking domain with keen interest in applications of AI. She is enthusiastic about exploring new technologies and advancements in today’s evolving world making everyone’s life easy.


🐝 Join the Fastest Growing AI Research Newsletter Read by Researchers from Google + NVIDIA + Meta + Stanford + MIT + Microsoft and many others…


Credit: Source link

ShareTweetSendSharePin

Related Posts

How Long Can You Expect Your Old Cassette Tapes To Last?
AI & Technology

How Long Can You Expect Your Old Cassette Tapes To Last?

September 20, 2026
OpenClaw Releases 2026.9.5 With Atomic Updates, Plugin Hot Reload, Conversation Sharing, and Expanded GPT Live
AI & Technology

OpenClaw Releases 2026.9.5 With Atomic Updates, Plugin Hot Reload, Conversation Sharing, and Expanded GPT Live

September 19, 2026
Trump Proposes Renaming Artificial Intelligence, Announces AI Force – Unite.AI
AI & Technology

Trump Proposes Renaming Artificial Intelligence, Announces AI Force – Unite.AI

September 19, 2026
SpaceX Targets September 28 For Starship’s First Orbital Flight
AI & Technology

SpaceX Targets September 28 For Starship’s First Orbital Flight

September 19, 2026
Next Post
AT&T CEO: Focused on Always-On Connectivity

AT&T CEO: Focused on Always-On Connectivity

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Rollins: A Great Business At A Much Better Price (NYSE:ROL)

Rollins: A Great Business At A Much Better Price (NYSE:ROL)

September 13, 2026
‘Dancing With the Stars’ Season 35 Premiere Kicks Off With the Men as the First Celebrity Is Eliminated: See the Scores, Who Went Home – The Hollywood Reporter

‘Dancing With the Stars’ Season 35 Premiere Kicks Off With the Men as the First Celebrity Is Eliminated: See the Scores, Who Went Home – The Hollywood Reporter

September 16, 2026
Meet the Press NOW — September 9

Meet the Press NOW — September 9

September 14, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!