• bitcoinBitcoin(BTC)$77,974.001.44%
  • ethereumEthereum(ETH)$2,501.221.79%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$746.192.23%
  • rippleXRP(XRP)$1.320.90%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$105.514.53%
  • tronTRON(TRX)$0.3376270.84%
  • zcashZcash(ZEC)$1,458.806.09%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.03-0.15%
  • HyperliquidHyperliquid(HYPE)$89.8112.00%
  • dogecoinDogecoin(DOGE)$0.0850884.45%
  • moneroMonero(XMR)$536.696.83%
  • USDSUSDS(USDS)$1.00-0.01%
  • whitebitWhiteBIT Coin(WBT)$80.311.43%
  • RainRain(RAIN)$0.012855-2.48%
  • chainlinkChainlink(LINK)$11.794.85%
  • leo-tokenLEO Token(LEO)$8.91-0.03%
  • cardanoCardano(ADA)$0.2134256.36%
  • stellarStellar(XLM)$0.1853731.06%
  • uniswapUniswap(UNI)$8.7926.62%
  • bitcoin-cashBitcoin Cash(BCH)$248.0710.27%
  • Ethena USDeEthena USDe(USDE)$1.00-0.01%
  • daiDai(DAI)$1.00-0.01%
  • nearNEAR Protocol(NEAR)$3.4922.03%
  • USD1USD1(USD1)$1.000.01%
  • litecoinLitecoin(LTC)$54.973.74%
  • CantonCanton(CC)$0.1072357.14%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.362.09%
  • avalanche-2Avalanche(AVAX)$7.974.90%
  • hedera-hashgraphHedera(HBAR)$0.0768322.26%
  • Global DollarGlobal Dollar(USDG)$1.000.00%
  • suiSui(SUI)$0.798.74%
  • shiba-inuShiba Inu(SHIB)$0.0000055.72%
  • crypto-com-chainCronos(CRO)$0.0592562.12%
  • MemeCoreMemeCore(M)$1.2913.69%
  • paypal-usdPayPal USD(PYUSD)$1.000.03%
  • BittensorBittensor(TAO)$244.387.29%
  • tether-goldTether Gold(XAUT)$4,369.430.26%
  • Circle USYCCircle USYC(USYC)$1.140.01%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • okbOKB(OKB)$114.051.85%
  • Ripple USDRipple USD(RLUSD)$1.00-0.01%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.12%
  • aaveAave(AAVE)$135.689.81%
  • AsterAster(ASTER)$0.750.29%
  • Pump.funPump.fun(PUMP)$0.0041916.29%
  • mantleMantle(MNT)$0.594.30%
  • polkadotPolkadot(DOT)$1.139.01%
  • OndoOndo(ONDO)$0.39030410.68%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Alibaba Qwen Releases Qwen3.8-Omni-Flash: A 1M-Context Omni-Modal Model Built Around Agentic Audio-Video Understanding and Tool Use

September 18, 2026
in AI & Technology
Reading Time: 19 mins read
A A
Alibaba Qwen Releases Qwen3.8-Omni-Flash: A 1M-Context Omni-Modal Model Built Around Agentic Audio-Video Understanding and Tool Use
ShareShareShareShareShare

Alibaba’s Qwen team has released Qwen3.8-Omni-Flash. They called it its first omni-modal model built around agentic capabilities. It accepts text, images, audio, and video, and it returns text. Audio-video understanding, reasoning, and tool use sit inside one model. The stated workflow is simple: understand the content, plan the task, execute with tools, deliver the result.

Is it deployable? Yes, as a hosted API today. It is live on QwenCloud, Alibaba Cloud Model Studio, and Qwen Studio. No open weights were announced at launch, so self-hosting is not an option.

YOU MAY ALSO LIKE

SK Hynix Debuts Ventures CVC Brand at Inaugural Silicon Valley Event – Unite.AI

NASA’s Moon Orbiter Has Spotted An Impact Crater That Only Happens Once A Century

What is Qwen3.8-Omni-Flash

The Qwen3.8-Omni-Flash is built on the Qwen3.8-Flash-Next architecture. That base model shipped with open weights in August 2026.

The context window is 1M tokens. QwenCloud lists 991K max input and 131K max output. Max reasoning length is 262K tokens.

Output is text only. The Model Studio docs point developers to Qwen3.5-Omni when they need generated speech. Thinking is on by default, with reasoning_effort set to xhigh. Setting it to none disables thinking.

The API follows both the DashScope and OpenAI protocols. It works with Chat Completions and the Responses API. Function calling, web search, structured outputs, context caching, and batch calls are supported.

Agentic Perception for Long Video

Most video models read a long file from start to finish. That holds even when the answer sits in 3 minutes of footage.

Qwen research team describes a different path. The agent starts from the question. It decides what to watch and hear. It then gathers evidence over several coarse-to-fine rounds. Compute and tokens go to the segments that matter.

The research team reports the result on OmniVideoBench. Accuracy rises from 63.4 to 67.8. Token use drops from 145,736 to 79,117. That is about 45.7% fewer tokens.

Reported Benchmarks

All figures here come from Qwen. Independent results were not available at publication.

  • Across 29 evaluations, the average score improves more than 25% over Qwen3.5-Omni-Plus.
  • WildClawBench-MM improves by 36.5 points. AgenticVBench improves by 22.3 points.
  • UniClawBench reaches 69.6.
  • LongAudioSpan gains 8.3 points. OmniVideoBench gains 9.6 points.
  • OmniCap-IF CSR and ISR improve by 8.5 and 14.1 points.

The research team states that audio-visual performance is close to Gemini 3.8 Flash. It also claims overall audio performance above Gemini 3.8 Flash. The X post summarizes the agent gains as +19.5 points on average across WildClawBench-MM and UniClawBench.

🚀 Meet Qwen3.8-Omni-Flash, Qwen’s first omni-modal model built around agentic capabilities!

Native audio-video understanding, reasoning, and tool use come together in one model: understand the content, plan the task, execute with tools, and deliver the result.

Highlights: 🥳
-… pic.twitter.com/iJypeohw7y

— Qwen (@Alibaba_Qwen) September 18, 2026

Pricing and Input Limits

QwenCloud lists $0.15 per 1M input tokens and $0.47 per 1M output tokens. Implicit cache hits cost $0.016 per 1M tokens.

The research team reports large cost cuts against Qwen3.5-Omni-Plus. Audio input costs over 98% less per hour. Audio-visual input costs over 93% less per hour. The X post puts the video input reduction at about 89%.

Key limits from the Model Studio docs:

  • Video files up to 2 hours and 2 GB by URL.
  • Audio files up to 3 hours.
  • Audio input in 113 languages and dialects.
  • Stable results with video sampled at up to 15 fps.
  • Two-channel stereo and four-channel FOA spatial audio through use_multichannel.
  • Availability in 6 regions: Beijing, Singapore, Hong Kong, Tokyo, Frankfurt, and Virginia.

Calling it takes a few lines with the OpenAI SDK:

import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["DASHSCOPE_API_KEY"],
    base_url=os.environ["DASHSCOPE_BASE_URL"],
)
completion = client.chat.completions.create(
    model="qwen3.8-omni-flash",
    messages=[{"role": "user", "content": [
        {"type": "video_url", "video_url": {"url": os.environ["VIDEO_URL"]}},
        {"type": "text", "text": "List the key moments with timestamps."},
    ]}],
    modalities=["text"],
    stream=True,
)
for chunk in completion:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="")

The model returns text, so tools do the media work. Qwen team is open-sourcing 2 projects to support that.

Qwen-MM-Plugins is live under Apache-2.0. Its tagline is ‘Make any agent harness multimodal-native.’ Each capability installs as a Skill plus an optional MCP server. The guided installer supports Claude Code, CodeBuddy, Codex, Qoder, OpenClaw, Qwen Code, and Gemini CLI.

The Omni capabilities map to the launch demos:

  • omni-memory builds an audio-visual memory of a long video.
  • omni-video2note converts a tutorial video into an illustrated PDF.
  • omni-chatcut covers Music-to-MV, movie commentary, and speaker-preserving video translation.

A core plugin lets the main model read local images and video frames natively. The README notes one current gap. Most harnesses cannot feed audio to the main model natively yet. Audio is routed through the API for now.

Interactive Explainer


Credit: Source link

ShareTweetSendSharePin

Related Posts

SK Hynix Debuts Ventures CVC Brand at Inaugural Silicon Valley Event – Unite.AI
AI & Technology

SK Hynix Debuts Ventures CVC Brand at Inaugural Silicon Valley Event – Unite.AI

September 18, 2026
NASA’s Moon Orbiter Has Spotted An Impact Crater That Only Happens Once A Century
AI & Technology

NASA’s Moon Orbiter Has Spotted An Impact Crater That Only Happens Once A Century

September 18, 2026
Waymo Is Expanding To Singapore
AI & Technology

Waymo Is Expanding To Singapore

September 18, 2026
Meta Launches Muse Mac App With File, Messages, and Calendar Access – Unite.AI
AI & Technology

Meta Launches Muse Mac App With File, Messages, and Calendar Access – Unite.AI

September 18, 2026
Next Post
Judge responds to Clancy lawyer request to remove juror

Judge responds to Clancy lawyer request to remove juror

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
Apple’s Redesigned Health App Is Available Now In The iOS 27.2 Developer Beta

Apple’s Redesigned Health App Is Available Now In The iOS 27.2 Developer Beta

September 16, 2026
Why Do Routers Have So Many Antennas?

Why Do Routers Have So Many Antennas?

September 13, 2026
Recall of soups and dips with tainted jalapeños gets highest risk level

Recall of soups and dips with tainted jalapeños gets highest risk level

September 17, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!