• bitcoinBitcoin(BTC)$81,014.005.82%
  • ethereumEthereum(ETH)$2,632.377.41%
  • tetherTether(USDT)$1.000.04%
  • binancecoinBNB(BNB)$764.344.35%
  • rippleXRP(XRP)$1.418.46%
  • usd-coinUSDC(USDC)$1.000.02%
  • solanaSolana(SOL)$113.8012.58%
  • tronTRON(TRX)$0.3384120.98%
  • zcashZcash(ZEC)$1,478.77-1.18%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.032.40%
  • HyperliquidHyperliquid(HYPE)$92.2610.36%
  • dogecoinDogecoin(DOGE)$0.0877907.36%
  • moneroMonero(XMR)$559.288.57%
  • whitebitWhiteBIT Coin(WBT)$83.295.55%
  • USDSUSDS(USDS)$1.000.03%
  • RainRain(RAIN)$0.0135115.17%
  • chainlinkChainlink(LINK)$12.379.10%
  • cardanoCardano(ADA)$0.2222099.84%
  • leo-tokenLEO Token(LEO)$8.90-0.26%
  • stellarStellar(XLM)$0.1942454.25%
  • uniswapUniswap(UNI)$9.0017.22%
  • bitcoin-cashBitcoin Cash(BCH)$253.759.33%
  • Ethena USDeEthena USDe(USDE)$1.000.04%
  • nearNEAR Protocol(NEAR)$3.6319.41%
  • daiDai(DAI)$1.000.00%
  • litecoinLitecoin(LTC)$57.126.04%
  • CantonCanton(CC)$0.11029711.29%
  • USD1USD1(USD1)$1.000.05%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.372.04%
  • avalanche-2Avalanche(AVAX)$8.238.33%
  • hedera-hashgraphHedera(HBAR)$0.0791815.00%
  • suiSui(SUI)$0.8110.26%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • shiba-inuShiba Inu(SHIB)$0.0000056.52%
  • MemeCoreMemeCore(M)$1.319.52%
  • crypto-com-chainCronos(CRO)$0.0597313.67%
  • BittensorBittensor(TAO)$250.288.59%
  • paypal-usdPayPal USD(PYUSD)$1.000.09%
  • tether-goldTether Gold(XAUT)$4,379.190.90%
  • Circle USYCCircle USYC(USYC)$1.140.03%
  • okbOKB(OKB)$116.994.02%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.05%
  • aaveAave(AAVE)$138.958.21%
  • mantleMantle(MNT)$0.629.99%
  • Pump.funPump.fun(PUMP)$0.00439711.79%
  • AsterAster(ASTER)$0.751.55%
  • polkadotPolkadot(DOT)$1.135.71%
  • OndoOndo(ONDO)$0.3961905.78%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

How to Build an Advanced End-to-End Voice AI Agent Using Hugging Face Pipelines?

September 17, 2025
in AI & Technology
Reading Time: 6 mins read
A A
How to Build an Advanced End-to-End Voice AI Agent Using Hugging Face Pipelines?
ShareShareShareShareShare

In this tutorial, we build an advanced voice AI agent using Hugging Face’s freely available models, and we keep the entire pipeline simple enough to run smoothly on Google Colab. We combine Whisper for speech recognition, FLAN-T5 for natural language reasoning, and Bark for speech synthesis, all connected through transformers pipelines. By doing this, we avoid heavy dependencies, API keys, or complicated setups, and we focus on showing how we can turn voice input into meaningful conversation and get back natural-sounding voice responses in real time. Check out the FULL CODES here.

Copy CodeCopiedUse a different Browser
!pip -q install "transformers>=4.42.0" accelerate torchaudio sentencepiece gradio soundfile


import os, torch, tempfile, numpy as np
import gradio as gr
from transformers import pipeline, AutoTokenizer, AutoModelForSeq2SeqLM


DEVICE = 0 if torch.cuda.is_available() else -1


asr = pipeline(
   "automatic-speech-recognition",
   model="openai/whisper-small.en",
   device=DEVICE,
   chunk_length_s=30,
   return_timestamps=False
)


LLM_MODEL = "google/flan-t5-base"
tok = AutoTokenizer.from_pretrained(LLM_MODEL)
llm = AutoModelForSeq2SeqLM.from_pretrained(LLM_MODEL, device_map="auto")


tts = pipeline("text-to-speech", model="suno/bark-small")

We install the necessary libraries and load three Hugging Face pipelines: Whisper for speech-to-text, FLAN-T5 for generating responses, and Bark for text-to-speech. We set the device automatically so that we can use GPU if available. Check out the FULL CODES here.

YOU MAY ALSO LIKE

Ben Bernstein, Manager of Cybersecurity Advisors at Huntress – Interview Series – Unite.AI

The New Resident Evil Movie Captures The Survival Horror Magic Of The Games

Copy CodeCopiedUse a different Browser
SYSTEM_PROMPT = (
   "You are a helpful, concise voice assistant. "
   "Prefer direct, structured answers. "
   "If the user asks for steps or code, use short bullet points."
)


def format_dialog(history, user_text):
   turns = []
   for u, a in history:
       if u: turns.append(f"User: {u}")
       if a: turns.append(f"Assistant: {a}")
   turns.append(f"User: {user_text}")
   prompt = (
       "Instruction:\n"
       f"{SYSTEM_PROMPT}\n\n"
       "Dialog so far:\n" + "\n".join(turns) + "\n\n"
       "Assistant:"
   )
   return prompt

We define a system prompt that guides our agent to stay concise and structured, and we implement a format_dialog function that takes past conversation history along with the user input and builds a prompt string for the model to generate the assistant’s reply. Check out the FULL CODES here.

Copy CodeCopiedUse a different Browser
def transcribe(filepath):
   out = asr(filepath)
   text = out["text"].strip()
   return text


def generate_reply(history, user_text, max_new_tokens=256):
   prompt = format_dialog(history, user_text)
   inputs = tok(prompt, return_tensors="pt", truncation=True).to(llm.device)
   with torch.no_grad():
       ids = llm.generate(
           **inputs,
           max_new_tokens=max_new_tokens,
           temperature=0.7,
           do_sample=True,
           top_p=0.9,
           repetition_penalty=1.05,
       )
   reply = tok.decode(ids[0], skip_special_tokens=True).strip()
   return reply


def synthesize_speech(text):
   out = tts(text)
   audio = out["audio"]
   sr = out["sampling_rate"]
   audio = np.asarray(audio, dtype=np.float32)
   return (sr, audio)

We create three core functions for our voice agent: transcribe converts recorded audio into text using Whisper, generate_reply builds a context-aware response from FLAN-T5, and synthesize_speech turns that response back into spoken audio with Bark. Check out the FULL CODES here.

Copy CodeCopiedUse a different Browser
def clear_history():
   return [], []


def voice_to_voice(mic_file, history):
   history = history or []
   if not mic_file:
       return history, None, "Please record something!"
   try:
       user_text = transcribe(mic_file)
   except Exception as e:
       return history, None, f"ASR error: {e}"


   if not user_text:
       return history, None, "Didn't catch that. Try again?"


   try:
       reply = generate_reply(history, user_text)
   except Exception as e:
       return history, None, f"LLM error: {e}"


   try:
       sr, wav = synthesize_speech(reply)
   except Exception as e:
       return history + [(user_text, reply)], None, f"TTS error: {e}"


   return history + [(user_text, reply)], (sr, wav), f"User: {user_text}\nAssistant: {reply}"


def text_to_voice(user_text, history):
   history = history or []
   user_text = (user_text or "").strip()
   if not user_text:
       return history, None, "Type a message first."
   try:
       reply = generate_reply(history, user_text)
       sr, wav = synthesize_speech(reply)
   except Exception as e:
       return history, None, f"Error: {e}"
   return history + [(user_text, reply)], (sr, wav), f"User: {user_text}\nAssistant: {reply}"


def export_chat(history):
   lines = []
   for u, a in history or []:
       lines += [f"User: {u}", f"Assistant: {a}", ""]
   text = "\n".join(lines).strip() or "No conversation yet."
   with tempfile.NamedTemporaryFile(delete=False, suffix=".txt", mode="w") as f:
       f.write(text)
       path = f.name
   return path

We add interactive functions for our agent: clear_history resets the conversation, voice_to_voice handles speech input and returns a spoken reply, text_to_voice processes typed input and speaks back, and export_chat saves the entire dialog into a downloadable text file. Check out the FULL CODES here.

Copy CodeCopiedUse a different Browser
with gr.Blocks(title="Advanced Voice AI Agent (HF Pipelines)") as demo:
   gr.Markdown(
       "##  Advanced Voice AI Agent (Hugging Face Pipelines Only)\n"
       "- **ASR**: openai/whisper-small.en\n"
       "- **LLM**: google/flan-t5-base\n"
       "- **TTS**: suno/bark-small\n"
       "Speak or type; the agent replies with voice + text."
   )


   with gr.Row():
       with gr.Column(scale=1):
           mic = gr.Audio(sources=["microphone"], type="filepath", label="Record")
           say_btn = gr.Button("🎤 Speak")
           text_in = gr.Textbox(label="Or type instead", placeholder="Ask me anything…")
           text_btn = gr.Button("💬 Send")
           export_btn = gr.Button("⬇ Export Chat (.txt)")
           reset_btn = gr.Button("♻ Reset")
       with gr.Column(scale=1):
           audio_out = gr.Audio(label="Assistant Voice", autoplay=True)
           transcript = gr.Textbox(label="Transcript", lines=6)
           chat = gr.Chatbot(height=360)
   state = gr.State([])


   def update_chat(history):
       return [(u, a) for u, a in (history or [])]


   say_btn.click(voice_to_voice, [mic, state], [state, audio_out, transcript]).then(
       update_chat, inputs=state, outputs=chat
   )
   text_btn.click(text_to_voice, [text_in, state], [state, audio_out, transcript]).then(
       update_chat, inputs=state, outputs=chat
   )
   reset_btn.click(clear_history, None, [chat, state])
   export_btn.click(export_chat, state, gr.File(label="Download chat.txt"))


demo.launch(debug=False)

We build a clean Gradio UI that lets us speak or type and then hear the agent’s response. We wire buttons to our callbacks, maintain chat state, and stream results into a chatbot, transcript, and audio player, all launched in one Colab app.

In conclusion, we see how seamlessly Hugging Face pipelines enable us to create a voice-driven conversational agent that listens, thinks, and responds. We now have a working demo that captures audio, transcribes it, generates intelligent responses, and returns speech output, all inside Colab. With this foundation, we can experiment with larger models, add multilingual support, or even extend the system with custom logic. Still, the core idea remains the same: we can bring together ASR, LLM, and TTS into one smooth workflow for an interactive voice AI experience.


Check out the FULL CODES here. Feel free to check out our GitHub Page for Tutorials, Codes and Notebooks. Also, feel free to follow us on Twitter and don’t forget to join our 100k+ ML SubReddit and Subscribe to our Newsletter.

The post How to Build an Advanced End-to-End Voice AI Agent Using Hugging Face Pipelines? appeared first on MarkTechPost.

Credit: Source link

ShareTweetSendSharePin

Related Posts

Ben Bernstein, Manager of Cybersecurity Advisors at Huntress – Interview Series – Unite.AI
AI & Technology

Ben Bernstein, Manager of Cybersecurity Advisors at Huntress – Interview Series – Unite.AI

September 18, 2026
The New Resident Evil Movie Captures The Survival Horror Magic Of The Games
AI & Technology

The New Resident Evil Movie Captures The Survival Horror Magic Of The Games

September 18, 2026
AWS Reworks Bedrock AgentCore Runtime for Elastic Memory, Fast Cold Starts – Unite.AI
AI & Technology

AWS Reworks Bedrock AgentCore Runtime for Elastic Memory, Fast Cold Starts – Unite.AI

September 18, 2026
Still The Best (And It’s Not Close)
AI & Technology

Still The Best (And It’s Not Close)

September 18, 2026
Next Post
Watch Live: Ousted CDC Director Susan Monarez says RFK Jr. politicizing vaccine decisions "really concerns me" – CBS News

Watch Live: Ousted CDC Director Susan Monarez says RFK Jr. politicizing vaccine decisions "really concerns me" - CBS News

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Paramount ‘deadly serious’ about leaving Hollywood as merger fight continues

Paramount ‘deadly serious’ about leaving Hollywood as merger fight continues

September 17, 2026
Amazon raises minimum hourly pay by  for US workers

Amazon raises minimum hourly pay by $20 for US workers

September 16, 2026
Secret new documentary on convicted tech founder Elizabeth Holmes

Secret new documentary on convicted tech founder Elizabeth Holmes

September 16, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!