• bitcoinBitcoin(BTC)$76,712.00-0.66%
  • ethereumEthereum(ETH)$2,477.39-1.78%
  • tetherTether(USDT)$1.00-0.02%
  • binancecoinBNB(BNB)$715.94-1.38%
  • rippleXRP(XRP)$1.34-1.79%
  • usd-coinUSDC(USDC)$1.00-0.01%
  • solanaSolana(SOL)$99.94-1.67%
  • tronTRON(TRX)$0.339560-0.12%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.000.00%
  • zcashZcash(ZEC)$1,076.18-4.23%
  • HyperliquidHyperliquid(HYPE)$77.53-2.67%
  • dogecoinDogecoin(DOGE)$0.082299-2.77%
  • RainRain(RAIN)$0.015169-3.71%
  • USDSUSDS(USDS)$1.00-0.02%
  • moneroMonero(XMR)$521.58-3.14%
  • whitebitWhiteBIT Coin(WBT)$79.58-0.86%
  • chainlinkChainlink(LINK)$11.18-2.69%
  • leo-tokenLEO Token(LEO)$9.04-1.16%
  • cardanoCardano(ADA)$0.202657-1.97%
  • stellarStellar(XLM)$0.176823-1.67%
  • Ethena USDeEthena USDe(USDE)$1.00-0.03%
  • daiDai(DAI)$1.000.02%
  • bitcoin-cashBitcoin Cash(BCH)$220.65-2.16%
  • USD1USD1(USD1)$1.00-0.03%
  • litecoinLitecoin(LTC)$53.610.16%
  • uniswapUniswap(UNI)$6.14-3.14%
  • CantonCanton(CC)$0.094708-2.68%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.34-2.82%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • hedera-hashgraphHedera(HBAR)$0.0748560.33%
  • avalanche-2Avalanche(AVAX)$7.30-1.14%
  • shiba-inuShiba Inu(SHIB)$0.000005-3.03%
  • nearNEAR Protocol(NEAR)$2.31-1.79%
  • suiSui(SUI)$0.70-2.99%
  • paypal-usdPayPal USD(PYUSD)$1.00-0.02%
  • crypto-com-chainCronos(CRO)$0.057119-4.33%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • tether-goldTether Gold(XAUT)$4,334.14-0.36%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • MemeCoreMemeCore(M)$1.14-3.05%
  • Ripple USDRipple USD(RLUSD)$1.000.00%
  • okbOKB(OKB)$112.04-1.73%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.14-0.02%
  • BittensorBittensor(TAO)$231.63-0.42%
  • BitwayBitway(BTW)$0.7231.45%
  • aaveAave(AAVE)$124.54-0.65%
  • pax-goldPAX Gold(PAXG)$4,338.19-0.37%
  • AsterAster(ASTER)$0.690.26%
  • mantleMantle(MNT)$0.56-1.48%
  • World Liberty FinancialWorld Liberty Financial(WLFI)$0.056648-1.39%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

Netflix AI Team Just Open-Sourced VOID: an AI Model That Erases Objects From Videos — Physics and All

April 4, 2026
in AI & Technology
Reading Time: 6 mins read
A A
Netflix AI Team Just Open-Sourced VOID: an AI Model That Erases Objects From Videos — Physics and All
ShareShareShareShareShare

Video editing has always had a dirty secret: removing an object from footage is easy; making the scene look like it was never there is brutally hard. Take out a person holding a guitar, and you’re left with a floating instrument that defies gravity. Hollywood VFX teams spend weeks fixing exactly this kind of problem. A team of researchers from Netflix and INSAIT, Sofia University ‘St. Kliment Ohridski,’ released VOID (Video Object and Interaction Deletion) model that can do it automatically.

VOID removes objects from videos along with all interactions they induce on the scene — not just secondary effects like shadows and reflections, but physical interactions like objects falling when a person is removed.

YOU MAY ALSO LIKE

How To Fix iMessage “Not Delivered” Error On iPhones

How To Adjust The Liquid Glass Effect On Your iPhone With iOS 27

What Problem Is VOID Actually Solving?

Standard video inpainting models — the kind used in most editing workflows today — are trained to fill in the pixel region where an object was. They’re essentially very sophisticated background painters. What they don’t do is reason about causality: if I remove an actor who is holding a prop, what should happen to that prop?

Existing video object removal methods excel at inpainting content ‘behind’ the object and correcting appearance-level artifacts such as shadows and reflections. However, when the removed object has more significant interactions, such as collisions with other objects, current models fail to correct them and produce implausible results.

VOID is built on top of CogVideoX and fine-tuned for video inpainting with interaction-aware mask conditioning. The key innovation is in how the model understands the scene — not just ‘what pixels should I fill?’ but ‘what is physically plausible after this object disappears?’

The canonical example from the research paper: if a person holding a guitar is removed, VOID also removes the person’s effect on the guitar — causing it to fall naturally. That’s not trivial. The model has to understand that the guitar was being supported by the person, and that removing the person means gravity takes over.

And unlike prior work, VOID was evaluated head-to-head against real competitors. Experiments on both synthetic and real data show that the approach better preserves consistent scene dynamics after object removal compared to prior video object removal methods including ProPainter, DiffuEraser, Runway, MiniMax-Remover, ROSE, and Gen-Omnimatte.

https://arxiv.org/pdf/2604.02296

The Architecture: CogVideoX Under the Hood

VOID is built on CogVideoX-Fun-V1.5-5b-InP — a model from Alibaba PAI — and fine-tuned for video inpainting with interaction-aware quadmask conditioning. CogVideoX is a 3D Transformer-based video generation model. Think of it like a video version of Stable Diffusion — a diffusion model that operates over temporal sequences of frames rather than single images. The specific base model (CogVideoX-Fun-V1.5-5b-InP) is released by Alibaba PAI on Hugging Face, which is the checkpoint engineers will need to download separately before running VOID.

The fine-tuned architecture specs: a CogVideoX 3D Transformer with 5B parameters, taking video, quadmask, and a text prompt describing the scene after removal as input, operating at a default resolution of 384×672, processing a maximum of 197 frames, using the DDIM scheduler, and running in BF16 with FP8 quantization for memory efficiency.

The quadmask is arguably the most interesting technical contribution here. Rather than a binary mask (remove this pixel / keep this pixel), the quadmask is a 4-value mask that encodes the primary object to remove, overlap regions, affected regions (falling objects, displaced items), and background to keep.

In practice, each pixel in the mask gets one of four values: 0 (primary object being removed), 63 (overlap between primary and affected regions), 127 (interaction-affected region — things that will move or change as a result of the removal), and 255 (background, keep as-is). This gives the model a structured semantic map of what’s happening in the scene, not just where the object is.

Two-Pass Inference Pipeline

VOID uses two transformer checkpoints, trained sequentially. You can run inference with Pass 1 alone or chain both passes for higher temporal consistency.

Pass 1 (void_pass1.safetensors) is the base inpainting model and is sufficient for most videos. Pass 2 serves a specific purpose: correcting a known failure mode. If the model detects object morphing — a known failure mode of smaller video diffusion models — an optional second pass re-runs inference using flow-warped noise derived from the first pass, stabilizing object shape along the newly synthesized trajectories.

It’s worth understanding the distinction: Pass 2 isn’t just for longer clips — it’s specifically a shape-stability fix. When the diffusion model produces objects that gradually warp or deform across frames (a well-documented artifact in video diffusion), Pass 2 uses optical flow to warp the latents from Pass 1 and feeds them as initialization into a second diffusion run, anchoring the shape of synthesized objects frame-to-frame.

How the Training Data Was Generated

This is where things get genuinely interesting. Training a model to understand physical interactions requires paired videos — the same scene, with and without the object, where the physics plays out correctly in both. Real-world paired data at this scale doesn’t exist. So the team built it synthetically.

Training used paired counterfactual videos generated from two sources: HUMOTO — human-object interactions rendered in Blender with physics simulation — and Kubric — object-only interactions using Google Scanned Objects.

HUMOTO uses motion-capture data of human-object interactions. The key mechanic is a Blender re-simulation: the scene is set up with a human and objects, rendered once with the human present, then the human is removed from the simulation and physics is re-run forward from that point. The result is a physically correct counterfactual — objects that were being held or supported now fall, exactly as they should. Kubric, developed by Google Research, applies the same idea to object-object collisions. Together, they produce a dataset of paired videos where the physics is provably correct, not approximated by a human annotator.

Key Takeaways

  • VOID goes beyond pixel-filling. Unlike existing video inpainting tools that only correct visual artifacts like shadows and reflections, VOID understands physical causality — if you remove a person holding an object, the object falls naturally in the output video.
  • The quadmask is the core innovation. Instead of a simple binary remove/keep mask, VOID uses a 4-value quadmask (values 0, 63, 127, 255) that encodes not just what to remove, but which surrounding regions of the scene will be physically affected — giving the diffusion model structured scene understanding to work with.
  • Two-pass inference solves a real failure mode. Pass 1 handles most videos; Pass 2 exists specifically to fix object morphing artifacts — a known weakness of video diffusion models — by using optical flow-warped latents from Pass 1 as initialization for a second diffusion run.
  • Synthetic paired data made training possible. Since real-world paired counterfactual video data doesn’t exist at scale, the research team built it using Blender physics re-simulation (HUMOTO) and Google’s Kubric framework, generating ground-truth before/after video pairs where the physics is provably correct.

Check out the Paper, Model Weight and Repo.  Also, feel free to follow us on Twitter and don’t forget to join our 120k+ ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.

The post Netflix AI Team Just Open-Sourced VOID: an AI Model That Erases Objects From Videos — Physics and All appeared first on MarkTechPost.

Credit: Source link

ShareTweetSendSharePin

Related Posts

How To Fix iMessage “Not Delivered” Error On iPhones
AI & Technology

How To Fix iMessage “Not Delivered” Error On iPhones

September 13, 2026
How To Adjust The Liquid Glass Effect On Your iPhone With iOS 27
AI & Technology

How To Adjust The Liquid Glass Effect On Your iPhone With iOS 27

September 13, 2026
Hierarchical NeRF with JAX3D for Volumetric Rendering, Novel-View Synthesis, and 3D Reconstruction
AI & Technology

Hierarchical NeRF with JAX3D for Volumetric Rendering, Novel-View Synthesis, and 3D Reconstruction

September 13, 2026
Car Manufacturers Are Ditching CarPlay In 2026: Here’s Why
AI & Technology

Car Manufacturers Are Ditching CarPlay In 2026: Here’s Why

September 13, 2026
Next Post
Sitting On .69B In Cash: Why UiPath Stock Is Too Cheap To Ignore (NYSE:PATH)

Sitting On $1.69B In Cash: Why UiPath Stock Is Too Cheap To Ignore (NYSE:PATH)

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
6,600 noncitizens mistakenly registered to vote in New Jersey

6,600 noncitizens mistakenly registered to vote in New Jersey

September 7, 2026
BrainChip Holdings Ltd (BCHPY) Q2 2026 Earnings Call Transcript

BrainChip Holdings Ltd (BCHPY) Q2 2026 Earnings Call Transcript

September 11, 2026
Reducto Releases r-1: A Single Pass Document Parsing Model That Cuts Errors 20% at 1 Cent Per Page

Reducto Releases r-1: A Single Pass Document Parsing Model That Cuts Errors 20% at 1 Cent Per Page

September 8, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!