• bitcoinBitcoin(BTC)$76,804.00-1.87%
  • ethereumEthereum(ETH)$2,448.80-0.76%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$712.36-1.44%
  • rippleXRP(XRP)$1.34-3.67%
  • usd-coinUSDC(USDC)$1.000.01%
  • solanaSolana(SOL)$99.08-2.31%
  • tronTRON(TRX)$0.3402530.38%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.040.98%
  • zcashZcash(ZEC)$1,086.86-12.34%
  • HyperliquidHyperliquid(HYPE)$78.83-5.55%
  • dogecoinDogecoin(DOGE)$0.083466-2.92%
  • RainRain(RAIN)$0.015751-1.38%
  • USDSUSDS(USDS)$1.000.00%
  • moneroMonero(XMR)$509.71-0.13%
  • whitebitWhiteBIT Coin(WBT)$79.50-1.61%
  • chainlinkChainlink(LINK)$11.51-2.21%
  • leo-tokenLEO Token(LEO)$9.190.03%
  • cardanoCardano(ADA)$0.205902-2.87%
  • stellarStellar(XLM)$0.175722-2.77%
  • daiDai(DAI)$1.00-0.01%
  • bitcoin-cashBitcoin Cash(BCH)$225.94-10.03%
  • Ethena USDeEthena USDe(USDE)$1.00-0.02%
  • USD1USD1(USD1)$1.00-0.01%
  • litecoinLitecoin(LTC)$52.35-1.56%
  • CantonCanton(CC)$0.098583-5.87%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.35-1.08%
  • uniswapUniswap(UNI)$6.00-2.02%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • hedera-hashgraphHedera(HBAR)$0.075122-1.82%
  • nearNEAR Protocol(NEAR)$2.47-0.21%
  • avalanche-2Avalanche(AVAX)$7.46-4.32%
  • suiSui(SUI)$0.73-4.84%
  • shiba-inuShiba Inu(SHIB)$0.000005-3.50%
  • paypal-usdPayPal USD(PYUSD)$1.000.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • crypto-com-chainCronos(CRO)$0.056286-3.77%
  • tether-goldTether Gold(XAUT)$4,320.62-1.71%
  • MemeCoreMemeCore(M)$1.15-5.47%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • okbOKB(OKB)$109.84-2.86%
  • BittensorBittensor(TAO)$235.63-6.84%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.23%
  • AsterAster(ASTER)$0.70-4.11%
  • polkadotPolkadot(DOT)$1.10-1.49%
  • mantleMantle(MNT)$0.57-5.21%
  • pax-goldPAX Gold(PAXG)$4,325.53-1.62%
  • aaveAave(AAVE)$121.02-3.21%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.056391-0.08%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Meet Redis LangCache: A Managed Semantic Cache That Cuts LLM API Costs by Up to 90% and Returns Cache Hits Up to 15x Faster

September 10, 2026
in AI & Technology
Reading Time: 11 mins read
A A
Meet Redis LangCache: A Managed Semantic Cache That Cuts LLM API Costs by Up to 90% and Returns Cache Hits Up to 15x Faster
ShareShareShareShareShare

Production LLM applications rarely receive a question nobody has asked before. Support assistants and RAG pipelines field the same intents thousands of times a day, each phrased differently, and most stacks treat every phrasing as a fresh, fully billed request. Redis LangCache is a fully managed semantic caching service that sits between the application and the model, matches incoming prompts against previously answered ones by meaning rather than exact text, and returns the stored response when a close enough match exists. Redis reports API cost savings of up to 90% and cache-hit responses up to 15x faster than re-querying the model.

Is it deployable? Yes. LangCache is available today as a public preview on Redis Cloud, accessed through a REST API with Python and JavaScript SDKs, and Redis notes that features and behavior may change during the preview.

YOU MAY ALSO LIKE

How These XL Phones Compete

Meta Is Testing Community Notes In Latin America. Fact Checkers Are Worried.

The Problem: Paraphrases Are Still Full LLM Calls

Consider three requests to a customer-support assistant:

  • “Can I get a refund after buying the monthly plan?”
  • “Is the monthly subscription refundable?”
  • “Can I cancel the plan and get my money back?”

The wording differs, but the question and answer are identical. Without a semantic cache, each version triggers a complete generation: input tokens processed, output tokens decoded, user waiting.

Prefix caching only removes part of that cost. When requests share a system prompt or context, the engine reuses the KV states computed for that prefix, but the request still reaches the LLM, new tokens still get processed, and the full answer still gets decoded. A prefix-cache hit is a cheaper generation call, not an avoided one.

How LangCache Works

LangCache moves the cache outside the model and stores the generated response itself. The architecture is a two-call loop:

  1. Before invoking the model, the app sends the prompt to POST /v1/caches/{cacheId}/entries/search.
  2. LangCache generates an embedding for the prompt and runs a vector search over stored entries.
  3. If a semantically similar entry clears the configured similarity threshold, the cached response is returned and no LLM call occurs.
  4. On a miss, the app calls its chosen LLM as usual, then stores the prompt and new response through POST /v1/caches/{cacheId}/entries for future matches.

Embedding generation is handled by the service, with default models or bring-your-own. Cache behavior is controlled through similarity thresholds, TTLs, and eviction policies, plus adaptive controls that tune precision and recall. Built on Redis’s vector database and exposed as a REST API, it works with any LLM provider and language. Hit rates and savings are monitored from the Redis Cloud console.

Credit: Source link

ShareTweetSendSharePin

Related Posts

How These XL Phones Compete
AI & Technology

How These XL Phones Compete

September 10, 2026
Meta Is Testing Community Notes In Latin America. Fact Checkers Are Worried.
AI & Technology

Meta Is Testing Community Notes In Latin America. Fact Checkers Are Worried.

September 10, 2026
IDScan Is Offering Free Credit Monitoring And ID Protection After Leaking Driver’s Licenses
AI & Technology

IDScan Is Offering Free Credit Monitoring And ID Protection After Leaking Driver’s Licenses

September 10, 2026
OpenAI Launches ChatGPT for Financial Services With Built-In Data – Unite.AI
AI & Technology

OpenAI Launches ChatGPT for Financial Services With Built-In Data – Unite.AI

September 10, 2026
Next Post
US Mint commemorates 9/11 25th anniversary with new half-dollar coin: ‘NEVER FORGET’

US Mint commemorates 9/11 25th anniversary with new half-dollar coin: 'NEVER FORGET'

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
IDScan Is Offering Free Credit Monitoring And ID Protection After Leaking Driver’s Licenses

IDScan Is Offering Free Credit Monitoring And ID Protection After Leaking Driver’s Licenses

September 10, 2026
A major earthquake strikes southern Japan

A major earthquake strikes southern Japan

September 4, 2026
OpenAI Commits B to Frontline Cyber Defense, Launches MS-ISAC Pilot – Unite.AI

OpenAI Commits $1B to Frontline Cyber Defense, Launches MS-ISAC Pilot – Unite.AI

September 4, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!