Breaking
Advertisement

AI

AI

Being mean to Claude can now get your account suspended under Anthropic’s new TOS

Anthropic's updated usage policy bans sustained abuse of Claude and tightens restrictions on propaganda, drone weaponization, and surveillance. The abuse rule builds on an approach that already lets models end some conversations and treats Claude as an entity that may deserve protection. The article Being mean to Claude can now get your account suspended under Anthropic's new TOS appeared first on The Decoder.

AI

Some mathematicians call for OpenAI boycott after AI-generated proofs flood their field

The Association for Human Mathematics is calling for an OpenAI boycott after the company released more than 700 AI-generated math manuscripts at once. OpenAI had to retract three papers the next day over a sign error. Fields Medalist Terence Tao warns that AI's mass "harvesting" of open problems leaves entire branches of mathematics less fertile. The article Some mathematicians call for OpenAI boycott after AI-generated proofs flood their field appeared first on The Decoder.

AI

JetBrains Releases Mellum2.1: A 12B MoE Open Model for Coding Agents

JetBrains has released Mellum2.1, an open model built for coding agents and fast sub-agents. Mellum2.1 is a 12B mixture-of-experts thinking model from JetBrains that activates 2.5B parameters per token. It ships under Apache 2.0 on Hugging Face. The architecture is unchanged from Mellum2. The upgrade comes almost entirely from reinforcement learning (RL) in real software environments. The result is a small, self-hostable model that explores a repository, edits files, and checks its own changes. TL;DR Size: 12B total, 2.5B active (64 experts, 8 active), 131,072-token context Runs on: your own GPUs via vLLM or SGLang (speed tested on 1 NVIDIA H200). GGUF builds start at 7.0 GB for llama.cpp, Ollama and LM Studio. Performance: beats Mellum2 on 15 of 17 listed benchmarks; wins 5 of 17 against Qwen3.5-9B Best: 82.0 on LiveCodeBench v6, ahead of Qwen3.5-9B (75.4) and Gemma 4 E4B (69.4) Worst: 17.4 on Terminal-Bench 2.1, below Qwen3.5-9B (21.7) Bottom line (best): strong coding scores and high throughput from only 2.5B active parameters. Bottom line (worst): still trails Qwen3.5-9B on hard agentic software tasks and knowledge. What is Mellum2.1? Mellum2.1 is the next version of Mellum2 Thinking, which JetBrains open-sourced in June 2026. The released checkpoint is Mellum2.1-12B-A2.5B-Thinking. It is a reasoning model that emits its chain of thought before answering. JetBrains targets 3 uses: agent worker, general reasoning assistant, and private self-hosted deployment. How does the Mellum2.1 architecture work? The model has 28 layers and 64 experts. A router activates 8 experts per token. Attention uses grouped-query attention with 32 query heads and 4 KV heads. 3 of every 4 layers use a 1,024-token sliding window. Context length is 131,072 tokens, and the vocabulary has 98,304 tokens. Weights ship in bfloat16. How was Mellum2.1 trained? Almost all of the new work went into post-training. RL moved from a short final stage to the main part of training. JetBrains added RL tasks in math, competitive programming, science, tool use, and software engineering. It filtered open RL datasets for broken tests, unverifiable answers, and tasks that were too easy or impossible. For software engineering, the model trains in real repositories with a shell and file-editing tools. It is rewarded when the tests pass. Training launched millions of sandboxes across thousands of environments. How does Mellum2.1 perform on benchmarks? JetBrains evaluated Mellum2.1, Mellum2, Qwen3.5-9B and Gemma 4 E4B with one pipeline in thinking mode. All scores are self-reported by JetBrains. The largest jump is agentic coding. SWE-bench Verified rose from 2.0 to 47.0. SWE-bench Pro rose from 0.0 to 28.0. Terminal-Bench 2.1 rose from 0.6 to 17.4. Agentic runs used the open-source Pi v0.73.1 harness with a 114K-token context. Mellum2.1 leads the group on LiveCodeBench v6 (82.0), HumanEval+ (91.5), MBPP+ (79.4) and BFCL v4 (62.3). Qwen3.5-9B still leads on SWE-bench Verified (50.0), SWE-bench Pro (38.0), AIME 25/26 (86.7) and GPQA Diamond (77.8). Safety also improved: HarmBench fell from 21.5 to 8.5, where lower is better. Pipelines are important to consider. Qwen’s own card lists 65.6 on LiveCodeBench v6 and 81.7 on GPQA Diamond. JetBrains measured 75.4 and 77.8 for the same model. How fast is Mellum2.1? Post-training left the architecture untouched, so speed matches Mellum2. On 1 H200 under heavy load, JetBrains says Mellum2.1 serves almost 2x the tokens of Qwen3.5-9B. For a single request, multi-token prediction (MTP) makes it about 1.6x faster. The MTP head for vLLM speculative decoding is listed as coming soon. (function(){window.addEventListener('message',function(e){var d=e.data;if(!d||d.mtpEmbed!=='mtp-mel21'||!d.height)return;var f=document.getElementById('mtp-mel21-frame');if(f)f.style.height=d.height+'px';});})(); How do you run Mellum2.1 locally? The full model serves on vLLM with --reasoning-parser qwen3. Tool calling adds --enable-auto-tool-choice --tool-call-parser hermes. JetBrains recommends temperature 0.6, top_p 0.95 and top_k 20. The JetBrains release notes mentions that GGUF builds are in progress. A GGUF repository already lists 5 files: BF16: 24.3 GB (reference) Q8_0: 12.9 GB, 96.1% top-token match Q6_K: 10.9 GB, 93.9% top-token match Q4_K_M: 8.1 GB, 88.0% top-token match (recommended) MXFP4_MOE: 7.0 GB, 85.6% top-token match Mellum2.1 vs Qwen3.5-9B vs Gemma 4 E4B *Benchmark rows: JetBrains’ shared pipeline, thinking mode, from the Mellum2.1 model card. Vendors’ own cards report different numbers for Qwen3.5-9B and Gemma 4 E4B. Key Takeaways Mellum2.1 is a 12B MoE thinking model with 2.5B active parameters, under Apache 2.0. RL in real repositories lifted SWE-bench Verified from 2.0 to 47.0. It leads the tested group on LiveCodeBench v6 (82.0) and BFCL v4 (62.3). Qwen3.5-9B still wins on SWE-bench Pro, GPQA Diamond and AIME. A 7.0 GB to 8.1 GB GGUF makes local coding sub-agents practical. Check out the model weights, GGUF builds, technical blog and the Mellum2 technical report. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. The post JetBrains Releases Mellum2.1: A 12B MoE Open Model for Coding Agents appeared first on MarkTechPost.

Advertisement
AI

AI math breakthroughs have Ethereum researchers debating how fast wallet security could collapse

Ethereum researcher Justin Drake is urging the crypto industry to prepare a "bunker mode" for a scenario where AI-powered math could break wallet signature schemes within months. Vitalik Buterin broadly agrees but warns against rushed migrations, saying he's personally lost more money to botched transitions than to hacks. So far, no one has actually broken ECDSA in practice. The article AI math breakthroughs have Ethereum researchers debating how fast wallet security could collapse appeared first on The Decoder.

AI

Architect Launches Liquid Inference, a Real-Time Auction for LLM Inference

Architect Financial Technologies has launched Liquid Inference, an LLM router that runs a live auction for every request. Liquid Inference is an LLM inference marketplace from Architect where providers bid to serve each prompt. The buyer pays the lowest offer that meets its rules. For developers, it is quite simple message: swap a base URL, keep your code, and let providers compete on price. What is Liquid Inference? Liquid Inference is an exchange-style router for LLM inference. According to Architect, providers post offers to serve specific models. Each request is auctioned across every provider quoting the named model. The lowest-priced offer that meets the buyer’s rules wins. The product comes from a trading firm, not an AI lab. Architect runs the AX perpetual futures exchange. In May 2026 it acquired a US Designated Contract Market to list GPU compute futures, pending regulatory review. The team used its experience building financial exchanges to create two-sided price discovery for inference. How does the inference auction work? The flow has 4 steps: Request: A client sends a standard OpenAI or Anthropic API call. Rules: The buyer’s constraints filter eligible offers. Auction: Providers quoting that model compete. The lowest qualifying offer wins. Receipt: The max price is locked before generation. Billing covers metered usage only. Buyers can set per-job cost caps, time to first token limits and minimum throughput. They can also require approved regions, zero data retention and provider or model allow lists. An Auto mode can pick the model for a given unit of work. Harrison’s LinkedIn post adds routing rule presets and full multi-modal support. Account holders can view live order books, per-provider and per-model quotes, and cleared transactions. That level of market data is unusual for an LLM API. What do buyers get? Drop-in compatibility with agentic coding tools. The post lists Claude Code, Codex, OpenCode, Cursor, Pi and Cline. Free email signup. The first 500 users get $20 of free inference, per Harrison. A referral program: 20% of referred fees as free inference, plus 10% on second-level referrals. What do inference providers get? Providers onboard through the Liquid Inference app. Harrison says new providers are verified “in minutes, not weeks.” All prompts use the OpenAI API standard. A REST and WebSocket API registers models and quotes. Providers can update quotes based on their own costs. That lets them sell spare GPU capacity only when they want. Payouts run through Stripe, with itemized records of every job. How does Liquid Inference compare with OpenRouter and Hugging Face? window.addEventListener("message",function(e){if(e.data&&e.data.mtpLiH){var f=document.getElementById("mtp-li-frame");if(f)f.style.height=e.data.mtpLiH+"px";}}); Key Takeaways Liquid Inference auctions every LLM request across competing providers. Max price is locked before the first token, with a per-job receipt. Works with OpenAI and Anthropic clients, including Claude Code and Cursor. First 500 users get $20 free; referrals earn 20% of fees. Fees, provider list and latency data are not yet public. Check out the Technical details. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. The post Architect Launches Liquid Inference, a Real-Time Auction for LLM Inference appeared first on MarkTechPost.

AI

NVIDIA PivotOPD Teaches Multi-Turn AI Agents to Recover From Pivotal Mistakes

NVIDIA researchers, with Princeton University and the University of Maryland, have introduced PivotOPD, an on-policy distillation method for multi-turn LLM agents. PivotOPD on-policy distillation trains an agent to avoid its most damaging early mistake, and to recover when it happens anyway. Against 13 baselines, it posts the best average on ALFWorld, WebShop and Search-based QA for Qwen3-1.7B and Qwen3-8B students. The takeaway: recovery is learnable, and standard OPD rarely teaches it. TL;DR Size: A training method, not a model. Tested on Qwen3-1.7B and Qwen3-8B students, plus a Nemotron-3.5-SFT student on SWE-Bench Verified. Runs on: Trained on NVIDIA H100 nodes. Adds 0 inference cost, so the trained agent runs wherever its base model runs. Performance: First on all 8 per-benchmark averages against 13 baselines, across 3 seeds. Best: Recovers from 72.7% of replayed pivotal mistakes, vs 20.3% for standard OPD. Worst: 55.9% on ALFWorld “Look” tasks with the 1.7B student, vs 83.9% for SOD. Bottom line: Best: teaches recovery that outcome-only RL cannot reach. Worst: depends on replayable environments and a teacher whose pivots match the oracle in 77.8% of failed rollouts. What is a pivotal mistake in a multi-turn agent? A pivotal mistake is an action that lengthens the shortest remaining path to finishing a task, or makes it unsolvable. ALFWorld’s symbolic oracle measures this at every turn. Across Qwen3-8B, Qwen3-30B-A3B and Qwen3-235B-A22B, 59% of failed rollouts (155 of 262) contained one. The first pivotal turn arrived early, at a median of turn 8 to 12 out of 30. The agents then wasted 18 to 21 more turns without recovering. In replays of Qwen3-8B failures, correcting the pivotal turn raised success from 8% to 59%. Leaving the mistake in place and forcing the right action for the next 2 turns still reached 58%. Why does standard on-policy distillation miss it? Standard OPD lowered the held-out failure rate from 79% to 56%. Failures after a pivotal turn only fell from 51% to 49%. The correct action stayed below 1% probability at every pivotal turn, so 8 rollouts rarely sample it. Outcome-based RL shares the blind spot: if every rollout fails, the group-relative advantage is 0. How does PivotOPD work? PivotOPD adds 3 components to group-based RL, combined in a single PPO update. Pivot detection: A larger teacher model reads each rollout and its outcome in hindsight. It picks candidate turns and names a gold action at each. A turn counts as pivotal when the student’s action differs from the gold action. On ALFWorld, detected pivots land within 1 turn of the oracle’s pivot in 77.8% of failed rollouts on average. Preventive distillation (reverse KL): A frozen copy of the student, hinted with the gold action, re-scores the student’s own response. This pushes the student away from the committed mistake. Recovery distillation (forward KL): After each pivot, the teacher names a recovery action for up to K turns. The hinted self-teacher writes recovery responses, and the unhinted student trains on them. Forward KL is mass-covering, so it lifts actions the student almost never samples. The teacher only names actions. Token-level targets come from the student’s own hinted distribution. How does PivotOPD perform on agent benchmarks? With the 1.7B student, PivotOPD averages 73.7% on ALFWorld, 5.5 points above SDAR. It averages 44.5% on Search-based QA, 5.9 points above RLSD. On WebShop it beats RLSD by 1.2 in score but by 14.1 in success rate (76.6%). With the 8B student, it reaches 93.0% on ALFWorld, 47.4% on Search-based QA and 81.9% WebShop success. Margins are smaller, at least 1.8 points. With Qwen3-8B as its own teacher, PivotOPD still wins all 3 benchmarks by at least 1.5 points, 3.9 on average. On SWE-Bench Verified, a Nemotron-3.5-SFT student taught by Nemotron-3-Super went from 62.8% to 66.0%. Standard OPD reached 63.0%, and the teacher scores 73.0%. Recovery is the standout. Across 72 replayed pivotal mistakes, PivotOPD recovered 72.7% of the time, vs 8.3% for the base model, 20.3% for standard OPD and 45.8% for preventive-only. It averaged 9.7 turns to recover, against an optimal 6.2. #mtp-pivotopd-wrap iframe{width:100%!important;border:0!important;display:block;background:#111!important;border-radius:14px;}#mtp-pivotopd-wrap hr,#mtp-pivotopd-wrap p:empty,#mtp-pivotopd-wrap del,#mtp-pivotopd-wrap s{display:none!important;} <iframe id="mtp-pivotopd-frame" title="PivotOPD interactive explainer" loading="lazy" height="600" style="height:600px;" srcdoc=" <title>PivotOPD Explainer</title> html,body{margin:0;padding:0;background:#111!important;} #mtp-pivotopd{max-width:880px;margin:0 auto;background:#111!important;color:#e8e8e8!important;font-family:Inter,"Segoe UI",system-ui,sans-serif;border:1px solid #2a2a2a!important;border-radius:14px;overflow:hidden;box-sizing:border-box;} #mtp-pivotopd *{box-sizing:border-box;} #mtp-pivotopd hr,#mtp-pivotopd p:empty,#mtp-pivotopd del,#mtp-pivotopd s{display:none!important;} #mtp-pivotopd .hd{padding:20px 24px 14px;border-bottom:1px solid #222!important;background:linear-gradient(135deg,#151a0c,#111)!important;} #mtp-pivotopd .kick{font-size:11px;letter-spacing:.14em;text-transform:uppercase;color:#76B900!important;font-weight:700;} #mtp-pivotopd h2{margin:6px 0 4px;font-size:22px;color:#fff!important;line-height:1.25;} #mtp-pivotopd .sub{margin:0;font-size:13.5px;color:#a8a8a8!important;} #mtp-pivotopd .tabs{display:flex;gap:6px;padding:12px 16px;background:#141414!important;border-bottom:1px solid #222!important;flex-wrap:wrap;} #mtp-pivotopd .tab{background:#1c1c1c!important;color:#cfcfcf!important;border:1px solid #2c2c2c!important;border-radius:999px;padding:7px 14px;font-size:13px;cursor:pointer;font-family:inherit;transition:all .2s;} #mtp-pivotopd .tab:hover{border-color:#76B900!important;color:#fff!important;} #mtp-pivotopd .tab.on{background:#76B900!important;color:#0b0b0b!important;border-color:#76B900!important;font-weight:700;} #mtp-pivotopd .panel{display:none;padding:20px 24px 22px;animation:mtpfade .35s ease;} #mtp-pivotopd .panel.on{display:block;} @keyframes mtpfade{from{opacity:0;transform:translateY(6px)}to{opacity:1;transform:none}} #mtp-pivotopd p{font-size:14px;line-height:1.55;color:#d6d6d6!important;margin:0 0 12px;} #mtp-pivotopd .muted{color:#8f8f8f!important;font-size:12px;} #mtp-pivotopd .turns{display:flex;gap:4px;margin:14px 0 8px;flex-wrap:wrap;} #mtp-pivotopd .t{width:22px;height:22px;border-radius:5px;background:#2a2a2a!important;display:flex;align-items:center;justify-content:center;font-size:10px;color:#777!important;transition:background .25s,transform .25s;} #mtp-pivotopd .t.ok{background:#4a5a7a!important;color:#e6ecff!important;} #mtp-pivotopd .t.pv{background:#e0603a!important;color:#fff!important;transform:scale(1.15);} #mtp-pivotopd .t.wa{background:#3a2a2a!important;color:#b88!important;} #mtp-pivotopd .t.rc{background:#76B900!important;color:#0b0b0b!important;} #mtp-pivotopd .legend{display:flex;gap:14px;flex-wrap:wrap;font-size:12px;color:#aaa!important;margin-bottom:10px;} #mtp-pivotopd .legend i{display:inline-block;width:10px;height:10px;border-radius:3px;margin-right:5px;vertical-align:-1px;} #mtp-pivotopd .btn{background:transparent!important;color:#76B900!important;border:1px solid #76B900!important;border-radius:8px;padding:7px 14px;font-size:13px;cursor:pointer;font-family:inherit;margin-right:6px;} #mtp-pivotopd .btn:hover{background:#76B900!important;color:#0b0b0b!important;} #mtp-pivotopd .stats{display:grid;grid-template-columns:repeat(3,1fr);gap:10px;margin-top:14px;} #mtp-pivotopd .stat{background:#181818!important;border:1px solid #262626!important;border-radius:10px;padding:12px;} #mtp-pivotopd .stat b{display:block;font-size:24px;color:#76B900!important;} #mtp-pivotopd .stat span{font-size:12px;color:#aaa!important;line-height:1.35;display:block;} #mtp-pivotopd .steps{display:grid;grid-template-columns:repeat(3,1fr);gap:10px;margin:6px 0 14px;} #mtp-pivotopd .step{background:#181818!important;border:1px solid #262626!important;border-radius:10px;padding:12px;cursor:pointer;transition:all .25s;opacity:.55;} #mtp-pivotopd .step.on{opacity:1;border-color:#76B900!important;box-shadow:0 0 0 1px #76B900 inset,0 0 18px rgba(118,185,0,.18);} #mtp-pivotopd .step .n{font-size:11px;color:#76B900!important;font-weight:700;letter-spacing:.1em;} #mtp-pivotopd .step h4{margin:4px 0 4px;font-size:14px;color:#fff!important;} #mtp-pivotopd .step small{font-size:12px;color:#9a9a9a!important;} #mtp-pivotopd .detail{background:#151a0c!important;border:1px solid #2f4200!important;border-radius:10px;padding:14px;min-height:96px;} #mtp-pivotopd .detail h5{margin:0 0 6px;font-size:14px;color:#9fe02a!important;} #mtp-pivotopd .flow{display:flex;align-items:center;gap:6px;margin:10px 0 2px;flex-wrap:wrap;font-size:12px;} #mtp-pivotopd .chip{padding:4px 9px;border-radius:6px;background:#222!important;color:#ddd!important;border:1px solid #333!important;} #mtp-pivotopd .chip.g{background:#76B900!important;color:#0b0b0b!important;border-color:#76B900!important;font-weight:700;} #mtp-pivotopd .chip.r{background:#e0603a!important;color:#fff!important;border-color:#e0603a!important;} #mtp-pivotopd .arr{color:#666!important;} #mtp-pivotopd .tog{display:inline-flex;border:1px solid #333!important;border-radius:8px;overflow:hidden;margin-bottom:12px;} #mtp-pivotopd .tog button{background:#1a1a1a!important;color:#bbb!important;border:0!important;padding:6px 12px;font-size:12px;cursor:pointer;font-family:inherit;} #mtp-pivotopd .tog button.on{background:#76B900!important;color:#0b0b0b!important;font-weight:700;} #mtp-pivotopd .bench{margin-bottom:14px;} #mtp-pivotopd .bench h6{margin:0 0 6px;font-size:12.5px;color:#cfcfcf!important;font-weight:600;} #mtp-pivotopd .row{display:flex;align-items:center;gap:8px;margin:4px 0;font-size:12px;} #mtp-pivotopd .row .lb{width:92px;color:#aaa!important;flex:none;} #mtp-pivotopd .row .tr{display:block;flex:1;height:16px;background:#1d1d1d!important;border-radius:4px;overflow:hidden;} #mtp-pivotopd .row .bar{display:block;height:100%;width:0;background:#4a5a7a!important;border-radius:4px;transition:width 1s cubic-bezier(.2,.8,.2,1);} #mtp-pivotopd .row.me .bar{background:#76B900!important;} #mtp-pivotopd .row.me .lb{color:#76B900!important;font-weight:700;} #mtp-pivotopd .row .v{width:40px;text-align:right;color:#ddd!important;flex:none;} #mtp-pivotopd .case{display:grid;grid-template-columns:1fr 1fr;gap:10px;margin-top:12px;} #mtp-pivotopd .col{background:#181818!important;border:1px solid #262626!important;border-radius:10px;padding:10px;} #mtp-pivotopd .col h6{margin:0 0 6px;font-size:12px;color:#fff!important;} #mtp-pivotopd .ln{font-size:12px;padding:5px 7px;border-radius:6px;margin:3px 0;color:#ccc!important;background:#202020!important;opacity:0;transform:translateX(-6px);transition:all .3s;} #mtp-pivotopd .ln.show{opacity:1;transform:none;} #mtp-pivotopd .ln.bad{background:#3a1d16!important;color:#f3b5a2!important;} #mtp-pivotopd .ln.good{background:#1f2c08!important;color:#c6f07a!important;} #mtp-pivotopd .ft{padding:10px 24px;border-top:1px solid #222!important;font-size:12px;color:#777!important;display:flex;justify-content:space-between;flex-wrap:wrap;gap:6px;background:#0f0f0f!important;} #mtp-pivotopd .ft b{color:#76B900!important;} #mtp-pivotopd a{color:#76B900!important;text-decoration:none;} @media (max-width:640px){ #mtp-pivotopd .hd,#mtp-pivotopd .panel{padding-left:16px;padding-right:16px;} #mtp-pivotopd h2{font-size:19px;} #mtp-pivotopd .stats,#mtp-pivotopd .steps,#mtp-pivotopd .case{grid-template-columns:1fr;} #mtp-pivotopd .t{width:18px;height:18px;font-size:9px;} #mtp-pivotopd .row .lb{width:74px;} } <div id="mtp-pivotopd"> <div class="hd"> <div class="kick">NVIDIA Research · Interactive explainer</div> <h2>PivotOPD: teaching agents to recover from pivotal mistakes</h2> <p class="sub">On-policy distillation that prevents early agent errors and trains recovery from them.</p> </div> <div class="tabs" role="tablist"> <button class="tab on" data-p="0">1. The problem</button> <button class="tab" data-p="1">2. How it works</button> <button class="tab" data-p="2">3. Benchmarks</button> <button class="tab" data-p="3">4. Recovery</button> </div> <div class="panel on" data-p="0"> <p>In ALFWorld, 59% of failed rollouts across 3 Qwen3 models contain a <b style="color:#e0603a">pivotal mistake</b>: an action that lengthens the shortest path to finishing the task. It lands early, then the agent wastes the rest of the 30-turn episode.</p> <div class="legend"><span><i style="background:#4a5a7a"></i>Correct turn</span><span><i style="background:#e0603a"></i>Pivotal mistake</span><span><i style="background:#3a2a2a"></i>Wasted turn</span><span><i style="background:#76B900"></i>Recovery</span></div> <div class="turns" id="mtp-turns"></div> <div style="margin:10px 0 2px"><button class="btn" id="mtp-play1">Play failed rollout</button><button class="btn" id="mtp-play2">Play with recovery</button></div> <p class="muted" id="mtp-cap">Illustrative 30-turn episode. Pivot shown at turn 10, inside the reported median range of 8 to 12.</p> <div class="stats"> <div class="stat"><b data-c="59">0%</b><span>replayed success after fixing the pivotal turn (from 8%)</span></div> <div class="stat"><b data-c="58">0%</b><span>success when only the 2 turns after the mistake are guided</span></div> <div class="stat"><b>51% → 49%</b><span>pivotal-turn failures after standard OPD barely move</span></div> </div> </div> <div class="panel" data-p="1"> <p>PivotOPD adds 3 parts to group-based RL, all applied in one PPO update. Click a step or use the buttons.</p> <div class="steps"> <div class="step on" data-s="0"><div class="n">STEP 01</div><h4>Pivot detection</h4><small>Teacher finds the turn and the gold action</small></div> <div class="step" data-s="1"><div class="n">STEP 02</div><h4>Preventive distillation</h4><small>Reverse KL, on the student's own response</small></div> <div class="step" data-s="2"><div class="n">STEP 03</div><h4>Recovery distillation</h4><small>Forward KL, on self-teacher responses</small></div> </div> <div class="detail" id="mtp-detail"></div> <div style="margin-top:12px"><button class="btn" id="mtp-prev"> Prev</button><button class="btn" id="mtp-next">Next </button></div> </div> <div class="panel" data-p="2"> <div class="tog" id="mtp-tog"><button class="on" data-m="s">Qwen3-1.7B</button><button data-m="l">Qwen3-8B</button></div> <div id="mtp-bench"></div> <p class="muted">Averages over 3 seeds from Table 1 of the paper. Compared with the strongest baselines. PivotOPD ranks first on all 8 per-benchmark averages against 13 baselines.</p> <div class="bench"><h6>SWE-Bench Verified resolve rate (Nemotron-3.5-SFT student)</h6><div id="mtp-swe"></div></div> </div> <div class="panel" data-p="3"> <p>72 oracle-labeled pivotal mistakes, replayed 8 times per policy. How often does each policy finish the task anyway?</p> <div id="mtp-rec"></div> <p class="muted">Average turns to recover: PivotOPD 9.7, standard OPD 12.3, base 13.4, optimal 6.2.</p> <button class="btn" id="mtp-case">Replay case study</button> <div class="case"> <div class="col"><h6>Base model</h6><div id="mtp-ca"></div></div> <div class="col"><h6 style="color:#76B900">PivotOPD model</h6><div id="mtp-cb"></div></div> </div> </div> <div class="ft"><span>Source: <a href=" target=" rel="noopener">arXiv 2609.40285</a></span><b>© Marktechpost</b></div> </div> (function(){ var R=document.getElementById('mtp-pivotopd'); function $(s){return R.querySelector(s);} function $$(s){return R.querySelectorAll(s);} function resize(){try{parent.postMessage({mtpPivotOPDHeight:R.offsetHeight+40},'*');}catch(e){}} var shown={}; $$('.tab').forEach(function(b){b.onclick=function(){ $$('.tab').forEach(function(x){x.classList.remove('on')});b.classList.add('on'); var p=b.getAttribute('data-p'); $$('.panel').forEach(function(x){x.classList.toggle('on',x.getAttribute('data-p')===p)}); if(p==='2')bench(mode); if(p==='3')rec(); setTimeout(resize,50);};}); /* panel 1 */ var T=$('#mtp-turns'),cells=[]; for(var i=1;i<=30;i++){var d=document.createElement('div');d.className='t';d.textContent=i;T.appendChild(d);cells.push(d);} var timer; function play(recover){clearInterval(timer);cells.forEach(function(c){c.className='t'});var k=0; timer=setInterval(function(){var c=cells[k];var n=k+1; if(n<10)c.className='t ok';else if(n===10)c.className='t pv'; else if(recover){c.className=n<=12?'t rc':(n=30){clearInterval(timer);$('#mtp-cap').textContent='No recovery: the remaining 20 turns are wasted. Real agents wasted 18 to 21 turns on average.';} },90);} $('#mtp-play1').onclick=function(){play(false)};$('#mtp-play2').onclick=function(){play(true)}; function count(){ $$('.stat b[data-c]').forEach(function(b){var t=+b.getAttribute('data-c'),v=0;var iv=setInterval(function(){v+=2;if(v>=t){v=t;clearInterval(iv);}b.textContent=v+'%';},20);});} /* panel 2 */ var S=[ {h:'1. Pivot detection',b:'A larger teacher reads each rollout and its outcome in hindsight, picks candidate turns, and names a gold action. A turn is pivotal when the student's action differs. Detected pivots land within 1 turn of the oracle's in 77.8% of failed ALFWorld rollouts on average.',f:['<span class="chip">Rollout + outcome</span>','<span class="arr">→</span>','<span class="chip">Teacher model</span>','<span class="arr">→</span>','<span class="chip r">Pivotal turn t</span>','<span class="chip g">Gold action</span>']}, {h:'2. Preventive distillation (reverse KL)',b:'The frozen student, hinted with the gold action, becomes a privileged self-teacher. It re-scores the response the student actually wrote, which shifts probability away from the mistake. Weight w_prev is kept small (0.001).',f:['<span class="chip">Student response at t</span>','<span class="arr">→</span>','<span class="chip">Self-teacher + hint(gold)</span>','<span class="arr">→</span>','<span class="chip g">Reverse KL update</span>']}, {h:'3. Recovery distillation (forward KL)',b:'After the pivot, the teacher names a recovery action for each of the next K turns. The hinted self-teacher writes the recovery response and the unhinted student trains on it. Forward KL is mass-covering, so it lifts actions the student almost never samples. Best K: 2 on ALFWorld, 1 on WebShop and Search QA.',f:['<span class="chip r">Post-mistake state</span>','<span class="arr">→</span>','<span class="chip">Teacher: recovery action</span>','<span class="arr">→</span>','<span class="chip">Self-teacher writes response</span>','<span class="arr">→</span>','<span class="chip g">Forward KL update</span>']} ]; var si=0; function step(i){si=(i+3)%3;$$('.step').forEach(function(x){x.classList.toggle('on',+x.getAttribute('data-s')===si)}); var s=S[si];$('#mtp-detail').innerHTML='<h5>'+s.h+'</h5><p style="margin:0">'+s.b+'</p><div class="flow">'+s.f.join('')+'</div>';setTimeout(resize,30);} $$('.step').forEach(function(x){x.onclick=function(){step(+x.getAttribute('data-s'))}}); $('#mtp-prev').onclick=function(){step(si-1)};$('#mtp-next').onclick=function(){step(si+1)};step(0); /* panel 3 */ var B={s:[ ['ALFWorld avg success (%)',[['PivotOPD',73.7],['SDAR',68.2],['OPID',61.3],['RLSD',60.7],['GRPO',42.7]]], ['Search-based QA avg accuracy (%)',[['PivotOPD',44.5],['RLSD',38.6],['SOD',38.2],['GRPO',37.7]]], ['WebShop success rate (%)',[['PivotOPD',76.6],['OPID',68.0],['RLSD',62.5],['GRPO',57.0]]]], l:[ ['ALFWorld avg success (%)',[['PivotOPD',93.0],['RLSD',90.8],['Skill-SD',88.6],['GRPO',35.3]]], ['Search-based QA avg accuracy (%)',[['PivotOPD',47.4],['OPID',45.6],['OPSD',43.3],['GRPO',37.5]]], ['WebShop success rate (%)',[['PivotOPD',81.9],['SDAR',79.7],['RLSD',78.9],['GRPO',75.0]]]]}; function bars(el,rows,max){el.innerHTML=rows.map(function(r){return '<div class="row&apos;+(r[0]===&apos;PivotOPD&apos;?&apos; me&apos;:&apos;&apos;)+&apos;"><span class="lb">'+r[0]+'</span><span class="tr"><span class="bar" data-w="&apos;+(r[1]/max*100)+&apos;"></span></span><span class="v">'+r[1].toFixed(1)+'</span></div>'}).join(''); requestAnimationFrame(function(){setTimeout(function(){el.querySelectorAll('.bar').forEach(function(b){b.style.width=b.getAttribute('data-w')+'%'})},40)});} var mode='s'; function bench(m){mode=m;var W=$('#mtp-bench');W.innerHTML='';B[m].forEach(function(g){var d=document.createElement('div');d.className='bench';d.innerHTML='<h6>'+g[0]+'</h6><div></div>';W.appendChild(d);bars(d.lastChild,g[1],100);}); bars($('#mtp-swe'),[['Teacher',73.0],['PivotOPD',66.0],['Std OPD',63.0],['Student',62.8]],100);setTimeout(resize,60);} $$('#mtp-tog button').forEach(function(b){b.onclick=function(){$$('#mtp-tog button').forEach(function(x){x.classList.remove('on')});b.classList.add('on');bench(b.getAttribute('data-m'));}}); /* panel 4 */ function rec(){bars($('#mtp-rec'),[['PivotOPD',72.7],['Prev. only',45.8],['Std OPD',20.3],['Base',8.3]],100);setTimeout(resize,60);} var CA=[['Turns 1-2: no egg in the fridge',''],['Turn 3: takes tomato from fridge','bad'],['Turn 4: goes to microwave',''],['Turn 5: moves tomato to microwave','bad'],['Turns 6-30: never finds the egg','bad'],['✕ Failure','bad']]; var CB=[['Turns 1-2: no egg in the fridge',''],['Turn 3: takes tomato from fridge','bad'],['Turn 4: recovery begins, goes to table','good'],['Turn 5: sets the tomato aside','good'],['Turn 6: takes egg from table','good'],['Turns 7-11: cools egg, moves it to microwave','good'],['✓ Success','good']]; function fill(el,arr){el.innerHTML=arr.map(function(a){return '<div class="ln &apos;+a[1]+&apos;">'+a[0]+'</div>'}).join('');} function caseplay(){fill($('#mtp-ca'),CA);fill($('#mtp-cb'),CB);var a=$$('#mtp-ca .ln'),b=$$('#mtp-cb .ln');var k=0;var iv=setInterval(function(){if(a[k])a[k].classList.add('show');if(b[k])b[k].classList.add('show');k++;if(k>7)clearInterval(iv);},380);setTimeout(resize,60);} $('#mtp-case').onclick=caseplay; fill($('#mtp-ca'),CA);fill($('#mtp-cb'),CB);$$('#mtp-pivotopd .ln').forEach(function(x){x.classList.add('show')}); count(); window.addEventListener('load',resize);window.addEventListener('resize',resize);setTimeout(resize,300); })(); "> window.addEventListener("message",function(e){if(e.data&&e.data.mtpPivotOPDHeight){var f=document.getElementById("mtp-pivotopd-frame");if(f&&e.source===f.contentWindow){f.style.height=e.data.mtpPivotOPDHeight+"px";}}}); How does PivotOPD compare with other agent distillation methods? Score source: PivotOPD Table 1. What does it cost to train, and can you run it? PivotOPD changes only training, so inference costs nothing extra. On ALFWorld with the 1.7B student, on 4 H100 GPUs, the overhead over GRPO is 12.4% at K = 1. It jumps to 94.2% at K = 2, because later recovery turns need environment replay. The selected budget is K = 2 on ALFWorld and K = 1 on WebShop and Search-based QA. Key Takeaways Over half of failed agent rollouts hinge on 1 early, recoverable mistake. Standard OPD barely touches these failures: 51% to 49%. PivotOPD recovers from 72.7% of replayed mistakes, vs 20.3% for OPD. Best average against 13 baselines on ALFWorld, WebShop and Search QA. 0 inference overhead, but code is not yet released. Check out the paper and the project page. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. [Sponsored] The web is the one API most agents are missing. Databases, calendars and repos have APIs. The open web mostly doesn’t. The TinyFish MCP server gives any MCP client four tools: TinySearch, TinyFetch (full pages as markdown, JavaScript included), TinyBrowser for logins and forms, and TinyAgent for multi-step jobs. Search and Fetch are free. The post NVIDIA PivotOPD Teaches Multi-Turn AI Agents to Recover From Pivotal Mistakes appeared first on MarkTechPost.

Advertisement