• bitcoinBitcoin(BTC)$84,015.00-0.88%
  • ethereumEthereum(ETH)$2,698.650.07%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$770.03-1.11%
  • rippleXRP(XRP)$1.51-1.65%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$119.96-2.59%
  • tronTRON(TRX)$0.3364030.59%
  • zcashZcash(ZEC)$1,512.27-5.91%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.060.00%
  • HyperliquidHyperliquid(HYPE)$88.75-3.54%
  • dogecoinDogecoin(DOGE)$0.094900-2.86%
  • chainlinkChainlink(LINK)$15.257.85%
  • moneroMonero(XMR)$542.20-0.94%
  • whitebitWhiteBIT Coin(WBT)$83.87-0.76%
  • USDSUSDS(USDS)$1.00-0.02%
  • cardanoCardano(ADA)$0.249438-2.97%
  • RainRain(RAIN)$0.012585-0.07%
  • leo-tokenLEO Token(LEO)$9.04-0.11%
  • stellarStellar(XLM)$0.2305936.03%
  • nearNEAR Protocol(NEAR)$4.92-9.35%
  • bitcoin-cashBitcoin Cash(BCH)$310.83-7.42%
  • hedera-hashgraphHedera(HBAR)$0.12740833.99%
  • uniswapUniswap(UNI)$8.91-8.96%
  • litecoinLitecoin(LTC)$69.53-2.42%
  • CantonCanton(CC)$0.131250-5.29%
  • avalanche-2Avalanche(AVAX)$10.52-4.79%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • suiSui(SUI)$1.17-8.63%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.65-0.38%
  • daiDai(DAI)$1.000.02%
  • USD1USD1(USD1)$1.000.00%
  • quant-networkQuant(QNT)$246.0829.12%
  • BittensorBittensor(TAO)$307.48-6.62%
  • crypto-com-chainCronos(CRO)$0.0696073.32%
  • shiba-inuShiba Inu(SHIB)$0.000006-3.89%
  • Global DollarGlobal Dollar(USDG)$1.000.02%
  • tether-goldTether Gold(XAUT)$4,144.35-3.17%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.02%
  • BitwayBitway(BTW)$0.99-17.54%
  • EthenaEthena(ENA)$0.263994-9.84%
  • MemeCoreMemeCore(M)$1.17-3.25%
  • OndoOndo(ONDO)$0.52-5.99%
  • okbOKB(OKB)$118.89-2.21%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • Pump.funPump.fun(PUMP)$0.0052405.28%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • aaveAave(AAVE)$149.23-4.26%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.15-0.12%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

How to Build an Advanced AI Agent with Summarized Short-Term and Vector-Based Long-Term Memory

September 2, 2025
in AI & Technology
Reading Time: 7 mins read
A A
How to Build an Advanced AI Agent with Summarized Short-Term and Vector-Based Long-Term Memory
ShareShareShareShareShare

In this tutorial, we walk you through building an advanced AI Agent that not only chats but also remembers. We start from scratch and demonstrate how to combine a lightweight LLM, FAISS vector search, and a summarization mechanism to create both short-term and long-term memory. By working together with embeddings and auto-distilled facts, we can craft an agent that adapts to our instructions, recalls important details in future conversations, and intelligently compresses context, ensuring the interaction remains smooth and efficient. Check out the FULL CODES here.

!pip -q install transformers accelerate bitsandbytes sentence-transformers faiss-cpu


import os, json, time, uuid, math, re
from datetime import datetime
import torch, faiss
from transformers import AutoModelForCausalLM, AutoTokenizer, pipeline, BitsAndBytesConfig
from sentence_transformers import SentenceTransformer
DEVICE = "cuda" if torch.cuda.is_available() else "cpu"

We begin by installing the essential libraries and importing all the required modules for our agent. We set up the environment to determine whether we are using a GPU or a CPU, allowing us to run the model efficiently. Check out the FULL CODES here.

YOU MAY ALSO LIKE

Darktrace CEO: AI Agents Are the New ‘Insider Threat’

Meta Is Starting An Enterprise Business To Justify Its Massive AI Spending

def load_llm(model_name="TinyLlama/TinyLlama-1.1B-Chat-v1.0"):
   try:
       if DEVICE=="cuda":
           bnb=BitsAndBytesConfig(load_in_4bit=True,bnb_4bit_compute_dtype=torch.bfloat16,bnb_4bit_quant_type="nf4")
           tok=AutoTokenizer.from_pretrained(model_name, use_fast=True)
           mdl=AutoModelForCausalLM.from_pretrained(model_name, quantization_config=bnb, device_map="auto")
       else:
           tok=AutoTokenizer.from_pretrained(model_name, use_fast=True)
           mdl=AutoModelForCausalLM.from_pretrained(model_name, torch_dtype=torch.float16 if torch.cuda.is_available() else torch.float32, low_cpu_mem_usage=True)
       return pipeline("text-generation", model=mdl, tokenizer=tok, device=0 if DEVICE=="cuda" else -1, do_sample=True)
   except Exception as e:
       raise RuntimeError(f"Failed to load LLM: {e}")

We define a function to load our language model. We set it up so that if a GPU is available, we use 4-bit quantization for efficiency; otherwise, we fall back to the CPU with optimized settings. This ensures we can generate text smoothly regardless of the hardware we are running on. Check out the FULL CODES here.

class VectorMemory:
   def __init__(self, path="/content/agent_memory.json", dim=384):
       self.path=path; self.dim=dim; self.items=[]
       self.embedder=SentenceTransformer("sentence-transformers/all-MiniLM-L6-v2", device=DEVICE)
       self.index=faiss.IndexFlatIP(dim)
       if os.path.exists(path):
           data=json.load(open(path))
           self.items=data.get("items",[])
           if self.items:
               X=torch.tensor([x["emb"] for x in self.items], dtype=torch.float32).numpy()
               self.index.add(X)
   def _emb(self, text):
       v=self.embedder.encode([text], normalize_embeddings=True)[0]
       return v.tolist()
   def add(self, text, meta=None):
       e=self._emb(text); self.index.add(torch.tensor([e]).numpy())
       rec={"id":str(uuid.uuid4()),"text":text,"meta":meta or {}, "emb":e}
       self.items.append(rec); self._save(); return rec["id"]
   def search(self, query, k=5, thresh=0.25):
       if len(self.items)==0: return []
       q=self.embedder.encode([query], normalize_embeddings=True)
       D,I=self.index.search(q, min(k, len(self.items)))
       out=[]
       for d,i in zip(D[0],I[0]):
           if i==-1: continue
           if d>=thresh: out.append((d,self.items[i]))
       return out
   def _save(self):
       slim=[{k:v for k,v in it.items()} for it in self.items]
       json.dump({"items":slim}, open(self.path,"w"), indent=2)

We create a VectorMemory class that gives our agent long-term memory. We store past interactions as embeddings using MiniLM and index them with FAISS, allowing us to search and recall relevant information later. Each memory is saved to disk, enabling the agent to retain its memory across sessions. Check out the FULL CODES here.

def now_iso(): return datetime.now().isoformat(timespec="seconds")
def clamp(txt, n=1600): return txt if len(txt)<=n else txt[:n]+" …"
def strip_json(s):
   m=re.search(r"\{.*\}", s, flags=re.S);
   return m.group(0) if m else None


SYS_GUIDE = (
"You are a helpful, concise assistant with memory. Use provided MEMORY when relevant. "
"Prefer facts from MEMORY over guesses. Answer directly; keep code blocks tight. If unsure, say so."
)


SUMMARIZE_PROMPT = lambda convo: f"Summarize the conversation below in 4-6 bullet points focusing on stable facts and tasks:\n\n{convo}\n\nSummary:"
DISTILL_PROMPT = lambda user: (
f"""Decide if the USER text contains durable info worth long-term memory (preferences, identity, projects, deadlines, facts).
Return compact JSON only: {{"save": true/false, "memory": "one-sentence memory"}}.
USER: {user}""")


class MemoryAgent:
   def __init__(self):
       self.llm=load_llm()
       self.mem=VectorMemory()
       self.turns=[]    
       self.summary=""   
       self.max_turns=10
   def _gen(self, prompt, max_new_tokens=256, temp=0.7):
       out=self.llm(prompt, max_new_tokens=max_new_tokens, temperature=temp, top_p=0.95, num_return_sequences=1, pad_token_id=self.llm.tokenizer.eos_token_id)[0]["generated_text"]
       return out[len(prompt):].strip() if out.startswith(prompt) else out.strip()
   def _chat_prompt(self, user, memory_context):
       convo="\n".join([f"{r.upper()}: {t}" for r,t in self.turns[-8:]])
       sys=f"System: {SYS_GUIDE}\nTime: {now_iso()}\n\n"
       mem = f"MEMORY (relevant excerpts):\n{memory_context}\n\n" if memory_context else ""
       summ=f"CONTEXT SUMMARY:\n{self.summary}\n\n" if self.summary else ""
       return sys+mem+summ+convo+f"\nUSER: {user}\nASSISTANT:"
   def _distill_and_store(self, user):
       try:
           raw=self._gen(DISTILL_PROMPT(user), max_new_tokens=120, temp=0.1)
           js=strip_json(raw)
           if js:
               obj=json.loads(js)
               if obj.get("save") and obj.get("memory"):
                   self.mem.add(obj["memory"], {"ts":now_iso(),"source":"distilled"})
                   return True, obj["memory"]
       except Exception: pass
       if re.search(r"\b(my name is|call me|I like|deadline|due|email|phone|working on|prefer|timezone|birthday|goal|exam)\b", user, flags=re.I):
           m=f"User said: {clamp(user,120)}"
           self.mem.add(m, {"ts":now_iso(),"source":"heuristic"})
           return True, m
       return False, ""
   def _maybe_summarize(self):
       if len(self.turns)>self.max_turns:
           convo="\n".join([f"{r}: {t}" for r,t in self.turns])
           s=self._gen(SUMMARIZE_PROMPT(clamp(convo, 3500)), max_new_tokens=180, temp=0.2)
           self.summary=s; self.turns=self.turns[-4:]
   def recall(self, query, k=5):
       hits=self.mem.search(query, k=k)
       return "\n".join([f"- ({d:.2f}) {h['text']} [meta={h['meta']}]" for d,h in hits])
   def ask(self, user):
       self.turns.append(("user", user))
       saved, memline = self._distill_and_store(user)
       mem_ctx=self.recall(user, k=6)
       prompt=self._chat_prompt(user, mem_ctx)
       reply=self._gen(prompt)
       self.turns.append(("assistant", reply))
       self._maybe_summarize()
       status=f"💾 memory_saved: {saved}; " + (f"note: {memline}" if saved else "note: -")
       print(f"\nUSER: {user}\nASSISTANT: {reply}\n{status}")
       return reply

We bring everything together into the MemoryAgent class. We design the agent to generate responses with context, distill important facts into long-term memory, and periodically summarize conversations to manage short-term context. With this setup, we create an assistant that remembers, recalls, and adapts to our interactions with it. Check out the FULL CODES here.

agent=MemoryAgent()


print("✅ Agent ready. Try these:\n")
agent.ask("Hi! My name is Nicolaus, I prefer being called Nik. I'm preparing for UPSC in 2027.")
agent.ask("Also, I work at  Visa in analytics and love concise answers.")
agent.ask("What's my exam year and how should you address me next time?")
agent.ask("Reminder: I like agentic RAG tutorials with single-file Colab code.")
agent.ask("Given my prefs, suggest a study focus for this week in one paragraph.")

We instantiate our MemoryAgent and immediately exercise it with a few messages to seed long-term memories and verify recall. We confirm it remembers our preferred name and exam year, adapts replies to our concise style, and uses past preferences (agentic RAG, single-file Colab) to tailor study guidance in the present.

In conclusion, we see how powerful it is when we give our AI Agent the ability to remember. We now have an agent that stores key details, recalls them when relevant, and summarizes conversations to stay efficient. This approach keeps our interactions contextual and evolving, making the agent feel more personal and intelligent with each exchange. With this foundation, we are ready to extend memory further, explore richer schemas, and experiment with more advanced memory-augmented agent designs.


Check out the FULL CODES here. Feel free to check out our GitHub Page for Tutorials, Codes and Notebooks. Also, feel free to follow us on Twitter and don’t forget to join our 100k+ ML SubReddit and Subscribe to our Newsletter.


Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is committed to harnessing the potential of Artificial Intelligence for social good. His most recent endeavor is the launch of an Artificial Intelligence Media Platform, Marktechpost, which stands out for its in-depth coverage of machine learning and deep learning news that is both technically sound and easily understandable by a wide audience. The platform boasts of over 2 million monthly views, illustrating its popularity among audiences.

Credit: Source link

ShareTweetSendSharePin

Related Posts

Darktrace CEO: AI Agents Are the New ‘Insider Threat’
AI & Technology

Darktrace CEO: AI Agents Are the New ‘Insider Threat’

September 28, 2026
Meta Is Starting An Enterprise Business To Justify Its Massive AI Spending
AI & Technology

Meta Is Starting An Enterprise Business To Justify Its Massive AI Spending

September 28, 2026
AI Takes Center Stage From Oracle to the White House
AI & Technology

AI Takes Center Stage From Oracle to the White House

September 28, 2026
Trump-Xi Optics Overshadow Substance, Gewirtz Says
AI & Technology

Trump-Xi Optics Overshadow Substance, Gewirtz Says

September 28, 2026
Next Post
Putin tells Kremlin officials that Alaska summit with Trump was ‘frank’ and ‘meaningful’

Putin tells Kremlin officials that Alaska summit with Trump was 'frank' and 'meaningful'

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
LIVE: Trumps announce IndyCar partnership for ‘Fostering the Future’ initiative | NBC News

LIVE: Trumps announce IndyCar partnership for ‘Fostering the Future’ initiative | NBC News

September 26, 2026
Anger grows across Spain after elderly woman evicted from her home of 70 years – NBC News

Anger grows across Spain after elderly woman evicted from her home of 70 years – NBC News

September 27, 2026
AWS Strands Agents Team Releases Strands Harness: An Open-Source Agent Harness With 28% Lower Token Cost at Comparable Accuracy

AWS Strands Agents Team Releases Strands Harness: An Open-Source Agent Harness With 28% Lower Token Cost at Comparable Accuracy

September 21, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!