• bitcoinBitcoin(BTC)$78,870.00-1.70%
  • ethereumEthereum(ETH)$2,460.67-1.25%
  • tetherTether(USDT)$1.000.01%
  • binancecoinBNB(BNB)$694.81-2.67%
  • rippleXRP(XRP)$1.44-4.72%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$96.91-3.95%
  • tronTRON(TRX)$0.337762-1.94%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.01-2.79%
  • HyperliquidHyperliquid(HYPE)$81.841.21%
  • dogecoinDogecoin(DOGE)$0.086538-5.90%
  • zcashZcash(ZEC)$785.80-7.15%
  • RainRain(RAIN)$0.01774620.96%
  • USDSUSDS(USDS)$1.000.00%
  • whitebitWhiteBIT Coin(WBT)$72.84-1.61%
  • leo-tokenLEO Token(LEO)$9.31-0.48%
  • chainlinkChainlink(LINK)$11.40-2.96%
  • moneroMonero(XMR)$442.61-0.96%
  • cardanoCardano(ADA)$0.210652-6.43%
  • stellarStellar(XLM)$0.183919-7.03%
  • bitcoin-cashBitcoin Cash(BCH)$266.54-3.44%
  • CantonCanton(CC)$0.118677-4.84%
  • daiDai(DAI)$1.000.00%
  • Ethena USDeEthena USDe(USDE)$1.000.02%
  • USD1USD1(USD1)$1.00-0.02%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.43-3.23%
  • litecoinLitecoin(LTC)$50.21-4.14%
  • hedera-hashgraphHedera(HBAR)$0.078471-4.04%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • avalanche-2Avalanche(AVAX)$7.35-3.76%
  • shiba-inuShiba Inu(SHIB)$0.000005-6.69%
  • suiSui(SUI)$0.76-7.11%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • crypto-com-chainCronos(CRO)$0.059128-2.74%
  • tether-goldTether Gold(XAUT)$4,624.460.36%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • paypal-usdPayPal USD(PYUSD)$1.000.02%
  • uniswapUniswap(UNI)$4.27-2.88%
  • MemeCoreMemeCore(M)$1.172.30%
  • nearNEAR Protocol(NEAR)$1.87-5.45%
  • okbOKB(OKB)$114.10-3.24%
  • BittensorBittensor(TAO)$233.48-3.27%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.22%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • pax-goldPAX Gold(PAXG)$4,633.360.33%
  • aaveAave(AAVE)$127.29-1.67%
  • AsterAster(ASTER)$0.700.45%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.0585800.42%
  • Pump.funPump.fun(PUMP)$0.004766-2.63%
  • OndoOndo(ONDO)$0.366304-6.71%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Liquid AI Open-Sources Pipette: A Reproducible Benchmarking Suite That Measures On-Device Models, Quantization, Runtime and Hardware Together

August 25, 2026
in AI & Technology
Reading Time: 17 mins read
A A
Liquid AI Open-Sources Pipette: A Reproducible Benchmarking Suite That Measures On-Device Models, Quantization, Runtime and Hardware Together
ShareShareShareShareShare

Model cards report quality under server-class, full-precision conditions. Those numbers rarely predict how the same model behaves on a phone. This week, Liquid AI released Pipette. It is an open-source platform for benchmarking foundation models on edge devices, built in partnership with Artificial Analysis as an independent methodology validator. Pipette treats on-device behavior as a property of the deployed system, not the model in isolation. Its unit of measurement is a full configuration: model + quantization + runtime + device. The launch dataset covers five on-device performance metrics across more than 1,000 model × quantization × runtime × device × context configurations, spanning 30+ models, llama.cpp builds for macOS, iOS, Windows and Android, and context lengths from 256 to 8,192 tokens. Initial verified results come from a MacBook Pro with M5 Max, an iPhone 17 Pro and a Galaxy S26 Ultra. The practical claim is testable: two 350M models at the same quantization on the same phone retain 78.4% and 33.8% of decode throughput at 4,096 tokens.

Is it deployable?

Yes, Pipette ships as Apache 2.0 infrastructure (pipette-mgmt, pipette-clients, pipette-scores), a public results dataset, a hosted dashboard, and native iOS and Android benchmark apps. Nothing is waitlisted. Publication of community-submitted results is still in beta.

  • Which companies: Any team shipping a model onto hardware it does not own. Solo developers and seed-stage startups can use the dashboard and apps without infrastructure. Mid-market product teams can run the clients across an internal device fleet. Large OEMs, chip vendors and enterprises can operate the whole pipeline behind their own firewall.
  • Industries: Consumer electronics and smartphone OEMs, automotive, industrial and robotics, healthcare devices, financial services, defense — anywhere latency, privacy or connectivity forces inference onto the device.
  • Applications: Model and quantization selection before a sprint commits; SoC and hardware procurement validation; regression testing when a runtime, OS or driver updates; context-length capacity planning; independent verification of vendor performance claims.

What Liquid AI shipped

Liquid AI released Pipette in partnership with Artificial Analysis, an independent validator that reviewed and verified the methodology. The premise is narrow and useful: on-device behavior is a property of the deployed system, not of the model in isolation.

The launch dataset covers five on-device performance metrics across more than 1,000 model × quantization × runtime × device × context configurations. It spans 30+ models, multiple quantization formats, llama.cpp builds for macOS, iOS, Windows and Android, and context lengths from 256 to 8,192 tokens. Initial published results come from a MacBook Pro with M5 Max, an iPhone 17 Pro and a Galaxy S26 Ultra, with AMD Ryzen AI Max+ 395 and Radeon 8060S results listed as coming soon.

In Pipette, the unit of measurement is a deployment configuration: model + quantization + runtime + device. A benchmark then defines the metric and token shape, producing a latency, throughput or memory result. Quality is tracked separately on IFBench, GPQA Diamond and MATH-500. Those quality scores currently come from llama.cpp evaluation runs on NVIDIA H100 80GB reference systems, then get matched to on-device runs sharing the same model and quantization — a quality number shown next to phone throughput was not produced on the phone.

Why the deployment context changes the answer

Four published comparisons show how far a configuration can move a decision:

  • Context scaling can diverge at identical parameter counts. At Q4_K_M on Galaxy S26 Ultra, Granite-4.0-H-350M retains 78.4% of its decode throughput from 256 to 4,096 input tokens, while Granite-4.0-350M retains only 33.8%.
  • Sparse activation buys speed, not memory. At 2,048 input tokens on the same phone, LFM2.5-8B-A1B decodes 2.4x faster than Qwen3.5-4B and 2.6x faster than Ministral-3-3B-Instruct-2512. It activates 1.5B of 8.5B parameters per token, yet still peaks at 5.29 GiB because all expert weights occupy memory.
  • Speed and quality do not co-locate. On iPhone 17 Pro at Q4_K_M, MiniCPM5-1B completes a 2,048-in / 256-out workload in 3.47 seconds versus 4.12 seconds for LFM2.5-1.2B-Instruct, a 15.8% reduction in elapsed time. On the same artifacts, LFM scores 9.0 points higher on MATH-500.
  • Near-identical system profiles can hide task-level reversals. At Q4_K_M and 2,048 input tokens on M5 Max, Granite-4.1-8B and Ministral-3-8B-Instruct-2512 differ by 2.4% in decode throughput and 1.2% in peak RAM. Granite leads IFBench by 7.3 points; Ministral leads GPQA Diamond by 14.0 points.

How the measurements are produced

Performance runs follow a published methodology: fixed token shapes, greedy decoding, a discarded warm-up, five measured repetitions and readiness gating. Before each timed repetition, a platform-specific check verifies thermal and load conditions; failing runs are not published. Evaluations use a separate protocol with deterministic, model-blind scoring, and pipette-scores never sees generation provenance. Every submission records benchmark version, token shape, model artifact, quantization, runtime version and settings, and device hardware and OS.

Interactive explainer

Key Takeaways

  • Pipette benchmarks configurations, not models: model + quantization + runtime + device.
  • Apache 2.0 stack, 1,000+ configurations, 30+ models, three verified devices at launch.
  • Quality evals run on H100 references and are matched to on-device performance, not measured on-device.
  • Identical parameter counts can differ 78.4% vs 33.8% in context-scaling retention.

Check out the Technical Details and Leaderboard. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.


YOU MAY ALSO LIKE

Adam Mosseri Says It’s ‘News To Me’ That Instagram Employees Limited His Exposure To Teen Safety Data

Here Are Three Sick Gibson Guitar Controllers Made For Stage Tour’s Dec. 10 Release

Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.

Credit: Source link

ShareTweetSendSharePin

Related Posts

Adam Mosseri Says It’s ‘News To Me’ That Instagram Employees Limited His Exposure To Teen Safety Data
AI & Technology

Adam Mosseri Says It’s ‘News To Me’ That Instagram Employees Limited His Exposure To Teen Safety Data

August 25, 2026
Here Are Three Sick Gibson Guitar Controllers Made For Stage Tour’s Dec. 10 Release
AI & Technology

Here Are Three Sick Gibson Guitar Controllers Made For Stage Tour’s Dec. 10 Release

August 25, 2026
Lego Skylines Is A Cozy City Builder From The Studio That’s Repairing Cities: Skylines II
AI & Technology

Lego Skylines Is A Cozy City Builder From The Studio That’s Repairing Cities: Skylines II

August 25, 2026
AI Method Reveals What Genomic Models Learn From DNA and Exposes Hidden Experimental Bias – Unite.AI
AI & Technology

AI Method Reveals What Genomic Models Learn From DNA and Exposes Hidden Experimental Bias – Unite.AI

August 25, 2026
Next Post
California AG says no settlement talks scheduled with Paramount after confidential details leaked

California AG says no settlement talks scheduled with Paramount after confidential details leaked

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Alex Murdaugh appears in court before murder retrial

Alex Murdaugh appears in court before murder retrial

August 23, 2026
Honda says it may not build new plant in North America without trade deal extension

Honda says it may not build new plant in North America without trade deal extension

August 25, 2026
Humanoids Market Worth 0BN By 2035: Barclays’ Todorova

Humanoids Market Worth $200BN By 2035: Barclays’ Todorova

August 20, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!