• bitcoinBitcoin(BTC)$84,230.00-2.65%
  • ethereumEthereum(ETH)$2,663.81-3.15%
  • tetherTether(USDT)$1.00-0.01%
  • binancecoinBNB(BNB)$765.76-2.82%
  • rippleXRP(XRP)$1.50-4.74%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$114.17-3.20%
  • tronTRON(TRX)$0.338993-0.72%
  • zcashZcash(ZEC)$1,516.96-1.15%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.03-1.23%
  • HyperliquidHyperliquid(HYPE)$93.03-3.60%
  • dogecoinDogecoin(DOGE)$0.092124-7.29%
  • moneroMonero(XMR)$547.89-4.49%
  • whitebitWhiteBIT Coin(WBT)$84.47-2.67%
  • USDSUSDS(USDS)$1.00-0.02%
  • chainlinkChainlink(LINK)$12.26-5.63%
  • cardanoCardano(ADA)$0.238466-4.48%
  • RainRain(RAIN)$0.012411-5.98%
  • leo-tokenLEO Token(LEO)$8.97-0.01%
  • stellarStellar(XLM)$0.202885-4.95%
  • bitcoin-cashBitcoin Cash(BCH)$346.994.16%
  • uniswapUniswap(UNI)$9.19-0.02%
  • nearNEAR Protocol(NEAR)$4.32-1.17%
  • Ethena USDeEthena USDe(USDE)$1.00-0.02%
  • litecoinLitecoin(LTC)$60.42-2.52%
  • daiDai(DAI)$1.00-0.01%
  • avalanche-2Avalanche(AVAX)$10.24-6.63%
  • USD1USD1(USD1)$1.00-0.01%
  • CantonCanton(CC)$0.107945-4.53%
  • suiSui(SUI)$0.97-3.85%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.40-2.93%
  • hedera-hashgraphHedera(HBAR)$0.089911-6.24%
  • BittensorBittensor(TAO)$291.50-6.59%
  • shiba-inuShiba Inu(SHIB)$0.000006-6.84%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • crypto-com-chainCronos(CRO)$0.061412-7.48%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.01%
  • MemeCoreMemeCore(M)$1.21-7.60%
  • tether-goldTether Gold(XAUT)$4,292.99-1.48%
  • BitwayBitway(BTW)$0.9712.71%
  • okbOKB(OKB)$118.23-3.65%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.140.07%
  • mantleMantle(MNT)$0.65-0.81%
  • aaveAave(AAVE)$138.91-3.55%
  • EthenaEthena(ENA)$0.2103682.81%
  • OndoOndo(ONDO)$0.414558-4.03%
  • pax-goldPAX Gold(PAXG)$4,290.01-1.52%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Google AI Introduces CardBench: A Comprehensive Benchmark Featuring Over 20 Real-World Databases and Thousands of Queries to Revolutionize Learned Cardinality Estimation

September 2, 2024
in AI & Technology
Reading Time: 6 mins read
A A
Google AI Introduces CardBench: A Comprehensive Benchmark Featuring Over 20 Real-World Databases and Thousands of Queries to Revolutionize Learned Cardinality Estimation
ShareShareShareShareShare

Cardinality estimation (CE) is crucial in optimizing query performance in relational databases. It involves predicting the number of intermediate results a database query will return, directly influencing the choice of execution plans by query optimizers. Accurate cardinality estimates are essential for selecting efficient join orders, determining whether to use an index and choosing the best join method. These decisions significantly impact query execution times and overall database performance. Inaccurate estimates can lead to poor execution plans, resulting in significantly slower performance, sometimes by several orders of magnitude. This makes CE a fundamental aspect of database management, with extensive research dedicated to improving its accuracy and efficiency.

The challenge, however, lies in the limitations of current methods for cardinality estimation. Traditional CE techniques, widely used in modern database systems, rely on heuristics and simplified models, such as assuming data uniformity and column independence. While computationally efficient, these methods often need to accurately predict cardinalities, especially in complex queries involving multiple tables and filters. Learned CE models have emerged as a promising alternative, offering better accuracy by leveraging data-driven approaches. However, these models must overcome significant barriers to adoption in practical settings. High training overheads, the need for large datasets, and a systematic benchmark for evaluating these models’ performance across diverse databases have hindered their widespread use.

YOU MAY ALSO LIKE

Logitech’s Yeti 2 Brings The 17-Year-Old USB Mic Into The Modern Age

Never Use ChatGPT For These Five Tasks

Existing methods, including traditional heuristic-based approaches, have been supplemented by learned models that utilize instance-specific features from the data. These learned models can improve accuracy but often at the cost of extensive training requirements. For example, workload-driven approaches necessitate running tens of thousands of queries to collect true cardinalities for training, leading to significant computational overheads. More recent data-driven methods attempt to model the data distribution within and across tables without executing queries, reducing some overhead but still requiring re-training as data changes. Despite these advancements, the lack of a comprehensive benchmark has made it difficult to compare different models and assess their generalizability across various datasets.

Researchers from Google Inc. have introduced CardBench, a benchmark designed to address the need for a systematic evaluation framework for learned cardinality estimation models. CardBench is a comprehensive benchmark that includes thousands of queries across 20 distinct real-world databases, significantly more than any previous benchmarks. This allows for a more thorough evaluation of learned CE models under various conditions. The benchmark supports three key setups: instance-based models, which are trained on a single dataset; zero-shot models, which are pre-trained on multiple datasets and then tested on an unseen dataset; and fine-tuned models, which are pre-trained and then fine-tuned with a small amount of data from the target dataset.

CardBench’s design includes tools for calculating necessary data statistics, generating realistic SQL queries, and creating annotated query graphs for training CE models. The benchmark offers two sets of training data: one for single table queries with multiple filter predicates and another for binary join queries involving two tables. The benchmark includes 9125 single table queries and 8454 binary join queries for one of its smaller datasets, ensuring a robust and challenging environment for model evaluation. The training data labels, derived from Google BigQuery, required seven CPU years of query execution time, highlighting the significant computational investment in creating this benchmark. By providing these datasets and tools, CardBench lowers the barrier for researchers interested in developing and testing new CE models.

Performance evaluations using CardBench show promising results, particularly for fine-tuned models. While zero-shot models struggle with accuracy when applied to unseen datasets, especially in complex queries involving joins, fine-tuned models achieve accuracy comparable to instance-based methods with far less training data. For instance, fine-tuned graph neural network (GNN) models achieved a median q-error of 1.32 and a 95th percentile q-error of 120 in binary join queries, significantly outperforming zero-shot models. The results suggest fine-tuning pre-trained models can substantially improve their performance even with 500 queries. This makes them viable for practical applications where training data may be limited.

In conclusion, CardBench represents a significant advancement in learned cardinality estimation. Researchers can systematically evaluate and compare different CE models by providing a comprehensive and diverse benchmark, fostering further innovation in this critical area. The benchmark’s ability to support fine-tuned models, which require less data and training time, offers a practical solution for real-world applications where the cost of training new models can be prohibitive. 


Check out the Paper. All credit for this research goes to the researchers of this project. Also, don’t forget to follow us on Twitter and join our Telegram Channel and LinkedIn Group. If you like our work, you will love our newsletter..

Don’t Forget to join our 50k+ ML SubReddit

Here is a highly recommended webinar from our sponsor: ‘Building Performant AI Applications with NVIDIA NIMs and Haystack’


Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.

▶• ılıılıılıılıılı Upcoming Live Session: ‘Building Performant AI Applications with NVIDIA NIMs and Haystack’.


Credit: Source link

ShareTweetSendSharePin

Related Posts

Logitech’s Yeti 2 Brings The 17-Year-Old USB Mic Into The Modern Age
AI & Technology

Logitech’s Yeti 2 Brings The 17-Year-Old USB Mic Into The Modern Age

September 23, 2026
Never Use ChatGPT For These Five Tasks
AI & Technology

Never Use ChatGPT For These Five Tasks

September 23, 2026
Apple Links Landmarks On Its Maps App To Hidden Histories Podcast Episodes
AI & Technology

Apple Links Landmarks On Its Maps App To Hidden Histories Podcast Episodes

September 23, 2026
Nokia Open-Sources AnyJev: A Training-Free Layer That Turns Any Open LLM Into a Calibrated Decision Model
AI & Technology

Nokia Open-Sources AnyJev: A Training-Free Layer That Turns Any Open LLM Into a Calibrated Decision Model

September 23, 2026
Next Post
Morning News NOW Full Broadcast – July 17

Morning News NOW Full Broadcast – July 17

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Lindsay Clancy jury deadlocked again as judge orders jurors to keep deliberating

Lindsay Clancy jury deadlocked again as judge orders jurors to keep deliberating

September 19, 2026
Waymo Announces Singapore Expansion Targeting 2028 Ride-Hailing Launch – Unite.AI

Waymo Announces Singapore Expansion Targeting 2028 Ride-Hailing Launch – Unite.AI

September 18, 2026
The 50/30/20 Budgeting Rule Explained in Simple Terms

The 50/30/20 Budgeting Rule Explained in Simple Terms

September 22, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!