• bitcoinBitcoin(BTC)$77,689.001.09%
  • ethereumEthereum(ETH)$2,503.293.64%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$720.962.06%
  • rippleXRP(XRP)$1.360.55%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$101.402.06%
  • tronTRON(TRX)$0.336046-0.65%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.040.79%
  • zcashZcash(ZEC)$1,160.59-1.09%
  • HyperliquidHyperliquid(HYPE)$81.490.44%
  • dogecoinDogecoin(DOGE)$0.0848661.58%
  • RainRain(RAIN)$0.015813-0.20%
  • USDSUSDS(USDS)$1.00-0.01%
  • moneroMonero(XMR)$512.432.42%
  • whitebitWhiteBIT Coin(WBT)$80.551.60%
  • chainlinkChainlink(LINK)$11.670.10%
  • leo-tokenLEO Token(LEO)$9.15-0.35%
  • cardanoCardano(ADA)$0.207163-1.12%
  • stellarStellar(XLM)$0.1785870.92%
  • Ethena USDeEthena USDe(USDE)$1.000.00%
  • bitcoin-cashBitcoin Cash(BCH)$228.37-0.41%
  • daiDai(DAI)$1.000.00%
  • USD1USD1(USD1)$1.000.02%
  • litecoinLitecoin(LTC)$53.312.29%
  • CantonCanton(CC)$0.097573-3.34%
  • uniswapUniswap(UNI)$6.154.59%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.361.61%
  • nearNEAR Protocol(NEAR)$2.608.59%
  • avalanche-2Avalanche(AVAX)$7.54-0.55%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • hedera-hashgraphHedera(HBAR)$0.0753330.23%
  • shiba-inuShiba Inu(SHIB)$0.0000051.68%
  • suiSui(SUI)$0.73-2.00%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.03%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.0567510.99%
  • MemeCoreMemeCore(M)$1.191.50%
  • tether-goldTether Gold(XAUT)$4,381.250.65%
  • Circle USYCCircle USYC(USYC)$1.140.03%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • okbOKB(OKB)$113.122.62%
  • BittensorBittensor(TAO)$238.99-1.21%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.02%
  • mantleMantle(MNT)$0.591.55%
  • aaveAave(AAVE)$124.892.75%
  • pax-goldPAX Gold(PAXG)$4,385.360.75%
  • AsterAster(ASTER)$0.70-0.55%
  • polkadotPolkadot(DOT)$1.09-0.40%
  • OndoOndo(ONDO)$0.3553171.78%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Are Pre-Trained Foundation Models the Future of Molecular Machine Learning? Introducing Unprecedented Datasets and the Graphium Machine Learning Library

October 19, 2023
in AI & Technology
Reading Time: 5 mins read
A A
Are Pre-Trained Foundation Models the Future of Molecular Machine Learning? Introducing Unprecedented Datasets and the Graphium Machine Learning Library
ShareShareShareShareShare

The recent results of machine learning in drug discovery have been largely attributed to graph and geometric deep learning models. These techniques have proven effective in modeling atomistic interactions, molecular representation learning, 3D and 4D situations, activity and property prediction, force field creation, and molecular production. Like other deep learning techniques, they need a lot of training data to provide excellent modeling accuracy. However, most training datasets in the present literature on treatments have small sample sizes. Surprisingly, recent developments in self-supervised learning, foundation models for computer vision and natural language processing, and deep understanding have significantly increased data efficiency.

In reality, it is demonstrated that the learned inductive bias reduces the data needs for downstream tasks by spending upfront in pre-training huge models with plenty of data, a one-time expense. After these accomplishments, other research has examined the advantages of pre-training large molecular graph neural networks for low-data molecular modeling. Due to the lack of big, labeled molecular datasets, these investigations could only use self-supervised approaches like contrastive learning, autoencoders, or denoising tasks. Only a small portion of the improvement made by self-supervised models in NLP and CV has yet been produced by low-data modeling attempts by fine-tuning from these models. 

Since molecules’ and their conformers’ behavior depends on their environment and is primarily controlled by quantum physics, this is partially explained by the underspecification of molecules and their conformers as graphs. For instance, it is widely known that molecules with comparable structures can exhibit significantly varying levels of bioactivity, a phenomenon known as an activity cliff, which restricts graph modeling based only on structural data. According to their argument, developing efficient base models for molecular modeling necessitates supervised training using information derived from quantum mechanical descriptions and biological environment-dependent data. 

Researchers from Québec AI Institute ,Valence Labs ,Université de Montréal, ,McGill University ,Graphcore ,New Jersey Institute of Technology ,RWTH Aachen University and HEC Montré makes three contributions to molecular research. They start by presenting a brand-new family of multitask datasets that are orders of magnitude bigger than the state of the art. Second, they discuss Graphium, a graph machine learning package enabling effective training on enormous datasets. Third, various baseline models demonstrate the benefit of training on multiple tasks. They provide three comprehensive and rigorously maintained multi-label datasets, the largest currently, with approximately 100 million molecules and over 3000 activities with sparse definitions. These datasets combine labels that describe quantum and biological features that have been learned through simulation and wet lab testing, and they have been created for the supervised training of foundation models. The responsibilities covered by the labels span both the node-level and the graph-level. 

The variety of labels makes it easier to acquire transfer skills effectively. It makes it possible to build fundamental models by increasing the generalizability of such models for various downstream molecular modeling activities. They meticulously vetted and added new information to the existing data to produce these extensive databases. As a result, descriptions of each molecule in their collection include information about its quantum mechanical characteristics and biological functions. The QM characteristics’ energy, electrical, and geometric components are calculated using various cutting-edge techniques, including semi-empirical techniques like PM6 and approaches based on density functional theory, such as B3LYP. As shown in Figure 1, their databases on biological activity include molecular signatures from toxicological profiling, gene expression profiling, and dose-response bioassays. 

Figure 1: A visual overview of the suggested molecular dataset collections. The “mixes” are designed to be anticipated concurrently while doing several tasks. They comprise jobs at the graph level and node level, as well as quantum, chemical, and biological aspects, categorical and continuous data points.

The simultaneous modeling of quantum and biological effects promotes the capacity to characterize complicated environment-dependent features of molecules that would be impossible to obtain from what are often small experimental datasets. The Library of Graphium Has created a complete graph machine learning toolkit called Graphium to enable effective training on these enormous multitask datasets. This innovative library streamlines the creation and training of molecular graph foundation models by including feature ensembles and complicated feature interactions. Graphium addresses the limitations of previous frameworks primarily intended for sequential samples with little interaction between node, edge, and graph characteristics by considering features and representations as essential building components and adding cutting-edge GNN layers. 

Additionally, Graphium handles the crucial and otherwise hard engineering of training models on huge dataset ensembles in a simple and highly configurable manner by offering features like dataset combination, addressing missing data, and joint training. Baseline Findings For the dataset mixtures offered, they train various models in single-dataset and multi-dataset scenarios. These provide reliable baselines that may serve as a reference point for upcoming users of these datasets and also offer some insight into the advantages of training using this multi-dataset methodology. Results for these models specifically demonstrate that training low-resource tasks may be greatly enhanced by movement in conjunction with bigger datasets. 

In conclusion, this work offers the biggest 2D molecular datasets. These datasets were created expressly to train foundation models that can accurately understand molecules’ quantum characteristics and biological flexibility and, as a result, be tailored to various downstream applications. Additionally, they created the Graphium library to simplify the training of these models and provide different baseline results that demonstrate the potency of the datasets and library being used.


Check out the Paper. All Credit For This Research Goes To the Researchers on This Project. Also, don’t forget to join our 31k+ ML SubReddit, 40k+ Facebook Community, Discord Channel, and Email Newsletter, where we share the latest AI research news, cool AI projects, and more.

If you like our work, you will love our newsletter..

We are also on WhatsApp. Join our AI Channel on Whatsapp..


YOU MAY ALSO LIKE

Apple’s iPhone Handoff Feature Will Cost You $5 A Month On T-Mobile

Cohere Releases North Small Translate: A 218B MoE Translation Model That Scores 83.6 on WMT26 Across 50 Languages

Aneesh Tickoo is a consulting intern at MarktechPost. He is currently pursuing his undergraduate degree in Data Science and Artificial Intelligence from the Indian Institute of Technology(IIT), Bhilai. He spends most of his time working on projects aimed at harnessing the power of machine learning. His research interest is image processing and is passionate about building solutions around it. He loves to connect with people and collaborate on interesting projects.


▶️ Now Watch AI Research Updates On Our Youtube Channel [Watch Now]

Credit: Source link

ShareTweetSendSharePin

Related Posts

Apple’s iPhone Handoff Feature Will Cost You  A Month On T-Mobile
AI & Technology

Apple’s iPhone Handoff Feature Will Cost You $5 A Month On T-Mobile

September 11, 2026
Cohere Releases North Small Translate: A 218B MoE Translation Model That Scores 83.6 on WMT26 Across 50 Languages
AI & Technology

Cohere Releases North Small Translate: A 218B MoE Translation Model That Scores 83.6 on WMT26 Across 50 Languages

September 11, 2026
Sakana AI Launches Fugu Max and Fugu Ultra v2 for Cheaper, Stronger Multi-Agent Orchestration
AI & Technology

Sakana AI Launches Fugu Max and Fugu Ultra v2 for Cheaper, Stronger Multi-Agent Orchestration

September 11, 2026
How These XL Phones Compete
AI & Technology

How These XL Phones Compete

September 10, 2026
Next Post
Impact of the War on Israeli and Palestinian Tech Sectors

Impact of the War on Israeli and Palestinian Tech Sectors

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Carly Simon announces Parkinson’s diagnosis

Carly Simon announces Parkinson’s diagnosis

September 4, 2026
Seaplane crashes into rocks off Washington coast

Seaplane crashes into rocks off Washington coast

September 5, 2026
Canada’s retaliatory US tariffs set to take effect as trade dispute grows – The Guardian

Canada’s retaliatory US tariffs set to take effect as trade dispute grows – The Guardian

September 7, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!