Anthropic's updated usage policy bans sustained abuse of Claude and tightens restrictions on propaganda, drone weaponization, and surveillance. The abuse rule builds on an approach that already lets models end some conversations and treats Claude as an entity that may deserve protection. The article Being mean to Claude can now get your account suspended under Anthropic's new TOS appeared first on The Decoder.
The Association for Human Mathematics is calling for an OpenAI boycott after the company released more than 700 AI-generated math manuscripts at once. OpenAI had to retract three papers the next day over a sign error. Fields Medalist Terence Tao warns that AI's mass "harvesting" of open problems leaves entire branches of mathematics less fertile. The article Some mathematicians call for OpenAI boycott after AI-generated proofs flood their field appeared first on The Decoder.
JetBrains has released Mellum2.1, an open model built for coding agents and fast sub-agents. Mellum2.1 is a 12B mixture-of-experts thinking model from JetBrains that activates 2.5B parameters per token. It ships under Apache 2.0 on Hugging Face. The architecture is unchanged from Mellum2. The upgrade comes almost entirely from reinforcement learning (RL) in real software environments. The result is a small, self-hostable model that explores a repository, edits files, and checks its own changes. TL;DR Size: 12B total, 2.5B active (64 experts, 8 active), 131,072-token context Runs on: your own GPUs via vLLM or SGLang (speed tested on 1 NVIDIA H200). GGUF builds start at 7.0 GB for llama.cpp, Ollama and LM Studio. Performance: beats Mellum2 on 15 of 17 listed benchmarks; wins 5 of 17 against Qwen3.5-9B Best: 82.0 on LiveCodeBench v6, ahead of Qwen3.5-9B (75.4) and Gemma 4 E4B (69.4) Worst: 17.4 on Terminal-Bench 2.1, below Qwen3.5-9B (21.7) Bottom line (best): strong coding scores and high throughput from only 2.5B active parameters. Bottom line (worst): still trails Qwen3.5-9B on hard agentic software tasks and knowledge. What is Mellum2.1? Mellum2.1 is the next version of Mellum2 Thinking, which JetBrains open-sourced in June 2026. The released checkpoint is Mellum2.1-12B-A2.5B-Thinking. It is a reasoning model that emits its chain of thought before answering. JetBrains targets 3 uses: agent worker, general reasoning assistant, and private self-hosted deployment. How does the Mellum2.1 architecture work? The model has 28 layers and 64 experts. A router activates 8 experts per token. Attention uses grouped-query attention with 32 query heads and 4 KV heads. 3 of every 4 layers use a 1,024-token sliding window. Context length is 131,072 tokens, and the vocabulary has 98,304 tokens. Weights ship in bfloat16. How was Mellum2.1 trained? Almost all of the new work went into post-training. RL moved from a short final stage to the main part of training. JetBrains added RL tasks in math, competitive programming, science, tool use, and software engineering. It filtered open RL datasets for broken tests, unverifiable answers, and tasks that were too easy or impossible. For software engineering, the model trains in real repositories with a shell and file-editing tools. It is rewarded when the tests pass. Training launched millions of sandboxes across thousands of environments. How does Mellum2.1 perform on benchmarks? JetBrains evaluated Mellum2.1, Mellum2, Qwen3.5-9B and Gemma 4 E4B with one pipeline in thinking mode. All scores are self-reported by JetBrains. The largest jump is agentic coding. SWE-bench Verified rose from 2.0 to 47.0. SWE-bench Pro rose from 0.0 to 28.0. Terminal-Bench 2.1 rose from 0.6 to 17.4. Agentic runs used the open-source Pi v0.73.1 harness with a 114K-token context. Mellum2.1 leads the group on LiveCodeBench v6 (82.0), HumanEval+ (91.5), MBPP+ (79.4) and BFCL v4 (62.3). Qwen3.5-9B still leads on SWE-bench Verified (50.0), SWE-bench Pro (38.0), AIME 25/26 (86.7) and GPQA Diamond (77.8). Safety also improved: HarmBench fell from 21.5 to 8.5, where lower is better. Pipelines are important to consider. Qwen’s own card lists 65.6 on LiveCodeBench v6 and 81.7 on GPQA Diamond. JetBrains measured 75.4 and 77.8 for the same model. How fast is Mellum2.1? Post-training left the architecture untouched, so speed matches Mellum2. On 1 H200 under heavy load, JetBrains says Mellum2.1 serves almost 2x the tokens of Qwen3.5-9B. For a single request, multi-token prediction (MTP) makes it about 1.6x faster. The MTP head for vLLM speculative decoding is listed as coming soon. (function(){window.addEventListener('message',function(e){var d=e.data;if(!d||d.mtpEmbed!=='mtp-mel21'||!d.height)return;var f=document.getElementById('mtp-mel21-frame');if(f)f.style.height=d.height+'px';});})(); How do you run Mellum2.1 locally? The full model serves on vLLM with --reasoning-parser qwen3. Tool calling adds --enable-auto-tool-choice --tool-call-parser hermes. JetBrains recommends temperature 0.6, top_p 0.95 and top_k 20. The JetBrains release notes mentions that GGUF builds are in progress. A GGUF repository already lists 5 files: BF16: 24.3 GB (reference) Q8_0: 12.9 GB, 96.1% top-token match Q6_K: 10.9 GB, 93.9% top-token match Q4_K_M: 8.1 GB, 88.0% top-token match (recommended) MXFP4_MOE: 7.0 GB, 85.6% top-token match Mellum2.1 vs Qwen3.5-9B vs Gemma 4 E4B *Benchmark rows: JetBrains’ shared pipeline, thinking mode, from the Mellum2.1 model card. Vendors’ own cards report different numbers for Qwen3.5-9B and Gemma 4 E4B. Key Takeaways Mellum2.1 is a 12B MoE thinking model with 2.5B active parameters, under Apache 2.0. RL in real repositories lifted SWE-bench Verified from 2.0 to 47.0. It leads the tested group on LiveCodeBench v6 (82.0) and BFCL v4 (62.3). Qwen3.5-9B still wins on SWE-bench Pro, GPQA Diamond and AIME. A 7.0 GB to 8.1 GB GGUF makes local coding sub-agents practical. Check out the model weights, GGUF builds, technical blog and the Mellum2 technical report. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. The post JetBrains Releases Mellum2.1: A 12B MoE Open Model for Coding Agents appeared first on MarkTechPost.
Ethereum researcher Justin Drake is urging the crypto industry to prepare a "bunker mode" for a scenario where AI-powered math could break wallet signature schemes within months. Vitalik Buterin broadly agrees but warns against rushed migrations, saying he's personally lost more money to botched transitions than to hacks. So far, no one has actually broken ECDSA in practice. The article AI math breakthroughs have Ethereum researchers debating how fast wallet security could collapse appeared first on The Decoder.
Architect Financial Technologies has launched Liquid Inference, an LLM router that runs a live auction for every request. Liquid Inference is an LLM inference marketplace from Architect where providers bid to serve each prompt. The buyer pays the lowest offer that meets its rules. For developers, it is quite simple message: swap a base URL, keep your code, and let providers compete on price. What is Liquid Inference? Liquid Inference is an exchange-style router for LLM inference. According to Architect, providers post offers to serve specific models. Each request is auctioned across every provider quoting the named model. The lowest-priced offer that meets the buyer’s rules wins. The product comes from a trading firm, not an AI lab. Architect runs the AX perpetual futures exchange. In May 2026 it acquired a US Designated Contract Market to list GPU compute futures, pending regulatory review. The team used its experience building financial exchanges to create two-sided price discovery for inference. How does the inference auction work? The flow has 4 steps: Request: A client sends a standard OpenAI or Anthropic API call. Rules: The buyer’s constraints filter eligible offers. Auction: Providers quoting that model compete. The lowest qualifying offer wins. Receipt: The max price is locked before generation. Billing covers metered usage only. Buyers can set per-job cost caps, time to first token limits and minimum throughput. They can also require approved regions, zero data retention and provider or model allow lists. An Auto mode can pick the model for a given unit of work. Harrison’s LinkedIn post adds routing rule presets and full multi-modal support. Account holders can view live order books, per-provider and per-model quotes, and cleared transactions. That level of market data is unusual for an LLM API. What do buyers get? Drop-in compatibility with agentic coding tools. The post lists Claude Code, Codex, OpenCode, Cursor, Pi and Cline. Free email signup. The first 500 users get $20 of free inference, per Harrison. A referral program: 20% of referred fees as free inference, plus 10% on second-level referrals. What do inference providers get? Providers onboard through the Liquid Inference app. Harrison says new providers are verified “in minutes, not weeks.” All prompts use the OpenAI API standard. A REST and WebSocket API registers models and quotes. Providers can update quotes based on their own costs. That lets them sell spare GPU capacity only when they want. Payouts run through Stripe, with itemized records of every job. How does Liquid Inference compare with OpenRouter and Hugging Face? window.addEventListener("message",function(e){if(e.data&&e.data.mtpLiH){var f=document.getElementById("mtp-li-frame");if(f)f.style.height=e.data.mtpLiH+"px";}}); Key Takeaways Liquid Inference auctions every LLM request across competing providers. Max price is locked before the first token, with a per-job receipt. Works with OpenAI and Anthropic clients, including Claude Code and Cursor. First 500 users get $20 free; referrals earn 20% of fees. Fees, provider list and latency data are not yet public. Check out the Technical details. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. The post Architect Launches Liquid Inference, a Real-Time Auction for LLM Inference appeared first on MarkTechPost.