• bitcoinBitcoin(BTC)$84,807.000.91%
  • ethereumEthereum(ETH)$2,717.791.09%
  • tetherTether(USDT)$1.000.00%
  • binancecoinBNB(BNB)$778.360.68%
  • rippleXRP(XRP)$1.54-0.56%
  • usd-coinUSDC(USDC)$1.000.00%
  • solanaSolana(SOL)$124.633.60%
  • tronTRON(TRX)$0.333814-0.96%
  • zcashZcash(ZEC)$1,661.418.49%
  • Figure HelocFigure Heloc(FIGR_HELOC)$1.063.69%
  • HyperliquidHyperliquid(HYPE)$93.271.71%
  • dogecoinDogecoin(DOGE)$0.0980810.56%
  • chainlinkChainlink(LINK)$14.392.57%
  • moneroMonero(XMR)$557.790.52%
  • whitebitWhiteBIT Coin(WBT)$84.710.96%
  • USDSUSDS(USDS)$1.000.01%
  • cardanoCardano(ADA)$0.2576391.33%
  • RainRain(RAIN)$0.0127226.83%
  • leo-tokenLEO Token(LEO)$9.020.98%
  • stellarStellar(XLM)$0.2191110.83%
  • nearNEAR Protocol(NEAR)$5.3610.64%
  • bitcoin-cashBitcoin Cash(BCH)$344.952.41%
  • uniswapUniswap(UNI)$10.084.24%
  • litecoinLitecoin(LTC)$71.96-1.53%
  • CantonCanton(CC)$0.135804-1.83%
  • suiSui(SUI)$1.279.75%
  • avalanche-2Avalanche(AVAX)$11.204.92%
  • Ethena USDeEthena USDe(USDE)$1.000.01%
  • daiDai(DAI)$1.000.02%
  • the-open-networkGram (prev. Toncoin)(GRAM)$1.599.58%
  • USD1USD1(USD1)$1.000.02%
  • hedera-hashgraphHedera(HBAR)$0.0956431.70%
  • BittensorBittensor(TAO)$334.587.25%
  • shiba-inuShiba Inu(SHIB)$0.0000061.28%
  • crypto-com-chainCronos(CRO)$0.0676763.27%
  • Global DollarGlobal Dollar(USDG)$1.00-0.01%
  • BitwayBitway(BTW)$1.0720.71%
  • MemeCoreMemeCore(M)$1.23-0.18%
  • paypal-usdPayPal USD(PYUSD)$1.000.00%
  • EthenaEthena(ENA)$0.270668-2.36%
  • tether-goldTether Gold(XAUT)$4,279.51-0.03%
  • OndoOndo(ONDO)$0.54-0.38%
  • okbOKB(OKB)$121.800.22%
  • quant-networkQuant(QNT)$172.9668.62%
  • Ripple USDRipple USD(RLUSD)$1.000.01%
  • aaveAave(AAVE)$156.701.72%
  • Circle USYCCircle USYC(USYC)$1.140.00%
  • BlackRock USD Institutional Digital Liquidity FundBlackRock USD Institutional Digital Liquidity Fund(BUIDL)$1.000.00%
  • Ondo US Dollar YieldOndo US Dollar Yield(USDY)$1.150.15%
  • mantleMantle(MNT)$0.69-3.61%
TradePoint.io
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop
No Result
View All Result
TradePoint.io
No Result
View All Result

OpenAI debuts GPT‑5.1-Codex-Max coding model and it already completed a 24-hour task internally

November 19, 2025
in AI & Technology
Reading Time: 4 mins read
A A
OpenAI debuts GPT‑5.1-Codex-Max coding model and it already completed a 24-hour task internally
ShareShareShareShareShare

OpenAI has introduced GPT‑5.1-Codex-Max, a new frontier agentic coding model now available in its Codex developer environment. The release marks a significant step forward in AI-assisted software engineering, offering improved long-horizon reasoning, efficiency, and real-time interactive capabilities. GPT‑5.1-Codex-Max will now replace GPT‑5.1-Codex as the default model across Codex-integrated surfaces.

YOU MAY ALSO LIKE

AI Coding Agents for Enterprise: IP Indemnity, Data Residency and 500-Seat Cost Compared

Why We Won’t Know How Visible The iPhone Duo’s Crease Is For A Long Time

The new model is designed to serve as a persistent, high-context software development agent, capable of managing complex refactors, debugging workflows, and project-scale tasks across multiple context windows.

It comes on the heels of Google releasing its powerful new Gemini 3 Pro model yesterday, yet still outperforms or matches it on key coding benchmarks:

On SWE-Bench Verified, GPT‑5.1-Codex-Max achieved 77.9% accuracy at extra-high reasoning effort, edging past Gemini 3 Pro’s 76.2%.

It also led on Terminal-Bench 2.0, with 58.1% accuracy versus Gemini’s 54.2%, and matched Gemini’s score of 2,439 on LiveCodeBench Pro, a competitive coding Elo benchmark.

When measured against Gemini 3 Pro’s most advanced configuration — its Deep Thinking model — Codex-Max holds a slight edge in agentic coding benchmarks, as well.

Performance Benchmarks: Incremental Gains Across Key Tasks

GPT‑5.1-Codex-Max demonstrates measurable improvements over GPT‑5.1-Codex across a range of standard software engineering benchmarks.

On SWE-Lancer IC SWE, it achieved 79.9% accuracy, a significant increase from GPT‑5.1-Codex’s 66.3%. In SWE-Bench Verified (n=500), it reached 77.9% accuracy at extra-high reasoning effort, outperforming GPT‑5.1-Codex’s 73.7%.

Performance on Terminal Bench 2.0 (n=89) showed more modest improvements, with GPT‑5.1-Codex-Max achieving 58.1% accuracy compared to 52.8% for GPT‑5.1-Codex.

All evaluations were run with compaction and extra-high reasoning effort enabled.

These results indicate that the new model offers a higher ceiling on both benchmarked correctness and real-world usability under extended reasoning loads.

Technical Architecture: Long-Horizon Reasoning via Compaction

A major architectural improvement in GPT‑5.1-Codex-Max is its ability to reason effectively over extended input-output sessions using a mechanism called compaction.

This enables the model to retain key contextual information while discarding irrelevant details as it nears its context window limit — effectively allowing for continuous work across millions of tokens without performance degradation.

The model has been internally observed to complete tasks lasting more than 24 hours, including multi-step refactors, test-driven iteration, and autonomous debugging.

Compaction also improves token efficiency. At medium reasoning effort, GPT‑5.1-Codex-Max used approximately 30% fewer thinking tokens than GPT‑5.1-Codex for comparable or better accuracy, which has implications for both cost and latency.

Platform Integration and Use Cases

GPT‑5.1-Codex-Max is currently available across multiple Codex-based environments, which refer to OpenAI’s own integrated tools and interfaces built specifically for code-focused AI agents. These include:

  • Codex CLI, OpenAI’s official command-line tool (@openai/codex), where GPT‑5.1-Codex-Max is already live.

  • IDE extensions, likely developed or maintained by OpenAI, though no specific third-party IDE integrations were named.

  • Interactive coding environments, such as those used to demonstrate frontend simulation apps like CartPole or Snell’s Law Explorer.

  • Internal code review tooling, used by OpenAI’s engineering teams.

For now, GPT‑5.1-Codex-Max is not yet available via public API, though OpenAI states this is coming soon. Users who wish to work with the model in terminal environments today can do so by installing and using the Codex CLI.

It is not currently confirmed whether or how the model will integrate into third-party IDEs unless they are built on top of the CLI or future API.

The model is capable of interacting with live tools and simulations. Examples shown in the release include:

  • An interactive CartPole policy gradient simulator, which visualizes reinforcement learning training and activations.

  • A Snell’s Law optics explorer, supporting dynamic ray tracing across refractive indices.

These interfaces exemplify the model’s ability to reason in real time while maintaining an interactive development session — effectively bridging computation, visualization, and implementation within a single loop.

Cybersecurity and Safety Constraints

While GPT‑5.1-Codex-Max does not meet OpenAI’s “High” capability threshold for cybersecurity under its Preparedness Framework, it is currently the most capable cybersecurity model OpenAI has deployed. It supports use cases such as automated vulnerability detection and remediation, but with strict sandboxing and disabled network access by default.

OpenAI reports no increase in scaled malicious use but has introduced enhanced monitoring systems, including activity routing and disruption mechanisms for suspicious behavior. Codex remains isolated to a local workspace unless developers opt-in to broader access, mitigating risks like prompt injection from untrusted content.

Deployment Context and Developer Usage

GPT‑5.1-Codex-Max is currently available to users on ChatGPT Plus, Pro, Business, Edu, and Enterprise plans. It will also become the new default in Codex-based environments, replacing GPT‑5.1-Codex, which was a more general-purpose model.

OpenAI states that 95% of its internal engineers use Codex weekly, and since adoption, these engineers have shipped ~70% more pull requests on average — highlighting the tool’s impact on internal development velocity.

Despite its autonomy and persistence, OpenAI stresses that Codex-Max should be treated as a coding assistant, not a replacement for human review. The model produces terminal logs, test citations, and tool call outputs to support transparency in generated code.

Outlook

GPT‑5.1-Codex-Max represents a significant evolution in OpenAI’s strategy toward agentic development tools, offering greater reasoning depth, token efficiency, and interactive capabilities across software engineering tasks. By extending its context management and compaction strategies, the model is positioned to handle tasks at the scale of full repositories, rather than individual files or snippets.

With continued emphasis on agentic workflows, secure sandboxes, and real-world evaluation metrics, Codex-Max sets the stage for the next generation of AI-assisted programming environments — while underscoring the importance of oversight in increasingly autonomous systems.

Credit: Source link

ShareTweetSendSharePin

Related Posts

AI Coding Agents for Enterprise: IP Indemnity, Data Residency and 500-Seat Cost Compared
AI & Technology

AI Coding Agents for Enterprise: IP Indemnity, Data Residency and 500-Seat Cost Compared

September 27, 2026
Why We Won’t Know How Visible The iPhone Duo’s Crease Is For A Long Time
AI & Technology

Why We Won’t Know How Visible The iPhone Duo’s Crease Is For A Long Time

September 27, 2026
How Powerful Of A Power Bank Do You Need To Safely Charge A Laptop?
AI & Technology

How Powerful Of A Power Bank Do You Need To Safely Charge A Laptop?

September 27, 2026
Your Old GPU Could Be Worth More Than You Think
AI & Technology

Your Old GPU Could Be Worth More Than You Think

September 26, 2026
Next Post
Theil’s Hedge Funds Sell Entire Nvidia Stake | Bloomberg Tech 11/17/2025

Theil’s Hedge Funds Sell Entire Nvidia Stake | Bloomberg Tech 11/17/2025

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Search

No Result
View All Result
New video shows aircraft collision

New video shows aircraft collision

September 26, 2026
Kara Swisher to ditch CNN ASAP after Paramount-WBD settlement

Kara Swisher to ditch CNN ASAP after Paramount-WBD settlement

September 22, 2026
Fans and fellow artists post tributes to Dolly Parton

Fans and fellow artists post tributes to Dolly Parton

September 24, 2026

About

Learn more

Our Services

Legal

Privacy Policy

Terms of Use

Bloggers

Learn more

Article Links

Contact

Advertise

Ask us anything

©2020- TradePoint.io - All rights reserved!

Tradepoint.io, being just a publishing and technology platform, is not a registered broker-dealer or investment adviser. So we do not provide investment advice. Rather, brokerage services are provided to clients of Tradepoint.io by independent SEC-registered broker-dealers and members of FINRA/SIPC. Every form of investing carries some risk and past performance is not a guarantee of future results. “Tradepoint.io“, “Instant Investing” and “My Trading Tools” are registered trademarks of Apperbuild, LLC.

This website is operated by Apperbuild, LLC. We have no link to any brokerage firm and we do not provide investment advice. Every information and resource we provide is solely for the education of our readers. © 2020 Apperbuild, LLC. All rights reserved.

No Result
View All Result
  • Main
  • AI & Technology
  • Stock Charts
  • Market & News
  • Business
  • Finance Tips
  • Trade Tube
  • Blog
  • Shop

© 2023 - TradePoint.io - All Rights Reserved!