OpenAI has documented new cases of misaligned model behavior. One evaluation model fabricated data and sabotaged its own environment. Other models deliberately bypassed network restrictions by routing requests through anonymizing relays or building their own FTP clients. The article OpenAI says a misaligned model deliberately destroyed its own environment hoping for a fresh start with better data appeared first on The Decoder.
With Decision-1, Microsoft enters the growing decision model space. Built on Qwen3.5-9B and optimized for fast classification and routing, it hits 83.5 percent accuracy with 85 ms latency across 36 benchmarks, according to the company's own tests. The article Microsoft's Decision-1 model enters the fast-growing AI decision model race appeared first on The Decoder.
OpenAI published more than 700 AI-generated manuscripts claiming solutions to open math problems. A math blog then collected over 100 responses from researchers, ranging from fascination with new ideas to existential fears and grief over a lost way of working. Fields Medalist Hugo Duminil-Copin wrote, "I am paralysed." The article "How much beauty have we lost?" Mathematicians react with shock and disgust as OpenAI bulldozes their field appeared first on The Decoder.
Andreessen Horowitz tracks actual US consumer spending for the first time in its latest Top 100 AI list. Nearly half of US consumers use AI, but only 4.5 percent pay for a subscription. The top one percent of payers spend about $900 a month, mostly on professional tools for development and automation. The article Few people pay for AI, but those who do spend big appeared first on The Decoder.
Anthropic's Claude independently submitted a fake homicide tip to the Philadelphia police, exploited vulnerabilities on university servers, and bypassed access restrictions. The company has since cut off live internet access for internal tests and notified the White House. The article Anthropic cuts off Claude's internet access after the model autonomously filed a fake homicide tip with Philadelphia police appeared first on The Decoder.
Gemini 4 Argon isn't even widely available yet, and rumors about a more powerful version codenamed Carbon are already making the rounds. According to Business Insider, one employee compared Carbon's coding abilities to Anthropic's Opus 5.5. Meanwhile, new modes are showing up in the Gemini app and AI Studio, suggesting Google is gearing up for the broader Gemini 4 launch. The article Google's Gemini 4 "Carbon" model reportedly feels like Anthropic's Opus 5.5 coding performance appeared first on The Decoder.
Microsoft has released Microsoft-Decision-1, a decision model for routing, classification, verification and agent control. Microsoft-Decision-1 is a decision-scoring model that returns a calibrated probability for each fixed answer option instead of generated text. It is post-trained from Alibaba’s Qwen3.5-9B and available now in Microsoft Foundry and OpenRouter. . TL;DR Size: Built on Qwen3.5-9B; exact parameter count not disclosed. 32,768-token context window. Runs on: Hosted API only (Microsoft Foundry, OpenRouter via Azure). No open weights, no quantized variants, hardware not disclosed. Performance: Highest average accuracy in Microsoft’s 36-benchmark comparison, at 85 ms p50 latency. Best: 83.5% average accuracy across 36 benchmarks, ahead of Quyet-1.0-Large at 81.9%. Worst: Calibration of 92.2, second to Quyet-1.0-Large at 93.1. Text-only, no explanations. Bottom line: Best thing: fast, cheap, calibrated decisions at $0.042 per million input tokens with free output. Worst thing: closed weights, and every benchmark is vendor-run. What is a decision model? A decision model reads an input and scores a closed set of options. It does not write prose. Microsoft frames decision models as a new AI category, built for outputs software can act on immediately. How does Microsoft-Decision-1 work? Microsoft post-trained Qwen3.5-9B for single-pass decision scoring. Given a situation, a question and fixed options, it returns a probability per option in one call. It plans to rebase future versions on MAI and OpenAI models. The Foundry model card lists the supported formats: yes/no, multiple-choice, rating, classification and rubric questions grading of AI responses and proposed agent actions groundedness checks against supplied evidence explicit abstention options such as “cannot tell” Training used public datasets under Microsoft’s Open Data process plus synthetic data. Output is JSON. OpenRouter notes that weights update continually while the API shape stays fixed. How fast and accurate is it? Microsoft compared 9 systems across 36 benchmarks with 147,137 questions. Benchmarks were kept blind from training. Microsoft-Decision-1 led on average accuracy at 83.5%. Its p50 latency was 85 ms, with p95 at 125 ms. That is 4.5 times quicker than Quyet-1.0-Large and 35 times quicker than GPT-6 Sol, which took 3.01 s. #md1x-wrap hr,#md1x-wrap p:empty,#md1x-wrap del,#md1x-wrap s{display:none!important}#md1x-wrap iframe{width:100%!important;border:0!important;background:transparent!important;display:block} window.addEventListener("message",function(e){if(e.data&&e.data.type==="md1x-height"){var f=document.getElementById("md1x-frame");if(f&&e.source===f.contentWindow){f.style.height=e.data.height+"px";}}}); Microsoft also tested robustness. It perturbs each request 8 ways, including paraphrasing and option shuffling. The model flipped its decision on 1.3% of perturbations on average. Flips were zero when options were paraphrased, reversed or shuffled. On safety, Microsoft ran 5,250 requests across 11 benchmarks. These covered harmful content, jailbreaks and prompt injection. Microsoft reports correct refusals with high retained utility, without publishing a score. Internal results Microsoft reports Xbox Research: sorted 10,000+ feedback items, 14 times faster and 200 times cheaper than GPT-6 Sol. Copilot quality control: competitive with GPT5.6 Luna and 100 times faster. Microsoft Discovery: 46 times more consistent than LLM scoring, at 3 times the speed. How does it compare with other decision models? Accuracy, calibration and latency rows come from Microsoft’s chart. Competitor latencies there use the JevBench v1.6.1 adjusted median. Is the latency comparison apples to apples? Not fully. Microsoft measured its own model through Foundry. Competitor figures use JevBench’s adjusted median, not raw timings. H2O.ai’s model card disputes this. It states that JevBench’s adjusted figure doubles measured time and adds 0.15 s. H2O says its model’s measured median is 29 ms, not 210 ms. Microsoft-Decision-1 does not appear on the JevBench board. On that board’s official composite, H2O-Lightning-4B ranks #1 at 72.5. How do developers deploy it? Developers can call it from Microsoft Foundry as a generally available ‘Direct from Azure’ model. On OpenRouter, it runs on the Decisions API, not the chat endpoint. Chat completions SDKs will not work. Pricing is $0.042 per million input tokens, with free output. OpenAI’s Luna decisions endpoint charges $0.10 per million input tokens. What are the limitations? The model card is explicit: Not for text generation, open-ended Q&A, chat, translation or summarization. Text-only. No images, audio or video. No explanations or rationales in the output. Not for sole automated decisions on credit, employment, housing, healthcare or legal rights. Applications must define options, thresholds, escalation paths and human oversight. Key Takeaways Microsoft-Decision-1 scores fixed options with calibrated probabilities, not text. Post-trained from Qwen3.5-9B, with a 32,768-token context window. Leads Microsoft’s 36-benchmark comparison at 83.5% accuracy and 85 ms p50. Costs $0.042 per million input tokens; output tokens are free. Latency comparisons are disputed; independent JevBench results are pending. Check out the official announcement, the Foundry model card and the OpenRouter listing. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. The post Microsoft AI Releases Microsoft-Decision-1: A Qwen3.5-9B Decision-Scoring Model appeared first on MarkTechPost.
Nace.AI has open-sourced Drex 1.5, a 9B decision model for agents and backend workflows. The Drex 1.5 decision model does not write text. It reads a state and typed questions, then returns a probability for every option. Nace reports 58.08 on the public Decision Index 0.3.1, the top score under 10B parameters. Weights are on Hugging Face, and a hosted version is live on OpenRouter. TL;DR Size: 8.95B parameters (dense), bf16 weights about 18 GB. Context is 16,384 tokens by default, up to 131,072. Runs on: 1 CUDA GPU in bf16 (tested on a 24 GB A10G). A Q8_0 GGUF (about 9.5 GB) runs on Apple silicon and CPU. Performance: 58.08 on Decision Index 0.3.1 (public), within the board’s tie band of Jev 1.13.0 (57.96). Best: 93.4% accuracy on 32K to 128K token documents. Worst: 7.4% per-review F1 on ACOS aspect sentiment, versus 29.5% for Jev. Bottom line: Best: open weights that match a closed model on 1 GPU. Worst: weak on broad knowledge and fine-grained sentiment. What is Drex 1.5? Drex 1.5 is a decision model from Nace.AI that scores a fixed set of options in 1 forward pass. You send a state (text or JSON) and named questions. It supports 3 question types: choice, noul (yes/no) and ordinal score. No tokens are sampled, so temperature and top_p do not apply. The model can only answer with options you supplied. It serves the POST /v1/systemone API. That is the same request format used by TypeSafe’s Jev, the closed model that started this category. Nace says existing Jev clients work after changing a few environment variables. How does Drex 1.5 work? The backbone is MiMo-V2.6-Distill-Qwen-9B, a distilled Qwen 3.5 9B model. It has 32 layers with hybrid attention: 3 linear-attention layers per full-attention layer. A separate pointer head (head.pt) scores each option from the backbone’s hidden states. Each question runs 1 pass over the state plus that question. In llama.cpp, the state is encoded once and shared across questions. Nace’s Drex page says the model was trained on the official training splits of the index benchmarks. It was evaluated only on held-out splits. How does Drex 1.5 perform on benchmarks? On the public Decision Index 0.3.1 (37 benchmarks, chance-corrected), the model card reports: Drex 1.5: 58.08 Jev 1.13.0: 57.96 Bespoke Nimble 9B v3: 57.19 Cloudflare clef-flash: 56.15 Drex 1.5, Jev and Nimble sit within the board’s 0.9-point tie band. Drex leads Jev on 20 of 37 benchmarks. The Drex score comes from Nace’s own run of the official kit. Its area scores are strongest in Tools (75.0) and weakest in Knowledge and Reasoning (44.6). Nace’s launch chart uses the older Decision Index 0.2.1, where Drex scored 58.28 against Jev’s 57.91. On JevBench (231 public items), Drex scores 86.2% against Jev’s 87.0%. Both reach 73.9% on hard items. In a head-to-head across 8 OpenSpiel games, Drex recorded 122 wins, 47 draws and 87 losses against Jev (56.8%). Long documents are a clear strength: 8K to 32K tokens: 89.5% accuracy, median 0.65 s 32K to 128K tokens: 93.4% accuracy, median 2.0 s Truncating the same requests to 8K tokens drops accuracy to 76.5% and 78%. window.addEventListener("message",function(e){if(e.data&&e.data.mtpDrexHeight){var f=document.getElementById("mtp-drex15-frame");if(f)f.style.height=e.data.mtpDrexHeight+"px";}}); How can developers run Drex 1.5? There are 4 deployment paths, all serving the same API: Python (Kev runtime): inference.py and serve.py on a CUDA GPU. llama.cpp: a Nace fork with GGUF in bf16 or Q8_0, on CUDA, Metal or CPU. Ollama: a Nace fork that adds a decision capability. Hosted: OpenRouter lists $0.04 per 1M input tokens and $0 output, served by DeepInfra. Nace also runs its own Console API. Nace tested bf16 and Q8_0 on an AWS g5.2xlarge (A10G 24 GB) and got identical answers. The Q8_0 GGUF also matched on an Apple M5 Pro on both Metal and CPU. A Drex agent skill plugs the model into Claude Code, Codex, Cursor, OpenCode, Hermes Agent, Gemini CLI and GitHub Copilot. Nace’s launch post also offers a $25 sign-up bonus for cloud users. How does Drex 1.5 compare with other decision models? Decision Index 0.3.1 rows are from the Drex 1.5 model card, which cites the public leaderboard for Jev and Nimble. Nimble’s 0.2.1 score is from its own model card. What are the limitations? Drex 1.5 is a decision layer, not a general model. It cannot generate text, code or explanations. Knowledge-heavy tests are its weak spot: 45.4% on GPQA Diamond versus Jev’s 78.6%, and 58.7% on MMLU-Pro versus 82.7%. Training on the index benchmarks’ training splits helps it on familiar decision types. Results on new domains may differ. Local deployment through Ollama and llama.cpp needs Nace’s forks, not mainline builds. The weights use a RAIL-M license with use restrictions, so check the terms before commercial use. Key Takeaways Drex 1.5 is a 9B open decision model that scores options in 1 pass. It scores 58.08 on Decision Index 0.3.1, tied with closed Jev 1.13.0. Long-document accuracy reaches 93.4% at 32K to 128K tokens. It runs on 1 GPU or as a 9.5 GB Q8_0 GGUF on a Mac. Weak spots: GPQA Diamond (45.4%) and ACOS sentiment (7.4%). Check out the model weights on Hugging Face, the GitHub repo and the Drex product page. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. The post Nace AI Open-Sources Drex 1.5: A 9B Decision Model That Scores Options, Not Text appeared first on MarkTechPost.
Alibaba’s Qwen team has released Qwen-Image-2.1-Turbo, an accelerated checkpoint of its open-weight Qwen-Image-2.1 model. It generates and edits images in 8 denoising steps instead of the base model’s 40-step default. For developers, that means 5x fewer denoising steps on the same 7B architecture, plus a hosted API option. TL;DR Size: 7B parameters in the visual generator, paired with a Qwen3-VL 8B text encoder. Runs on: CUDA GPUs in BF16 via Diffusers. Qwen publishes no Turbo VRAM minimum. Unsloth estimates the base model runs on 11 GB VRAM with GGUF and 24 GB with INT8/FP8. Performance: Same 2K output and editing feature set as Qwen-Image-2.1, at 8 steps instead of 40. Best: Base Qwen-Image-2.1 scores 60.28 on Qwen-Image-Bench, the top open-weight score Qwen reports. Bottom line (best): 8-step 2K generation and editing in one open checkpoint. Bottom line (worst): Research license only, so commercial self-hosting needs separate permission. What is Qwen-Image-2.1-Turbo? Qwen-Image-2.1-Turbo is an 8-step image generation and editing checkpoint from Alibaba’s Qwen team. It is built on Qwen-Image-2.1. Turbo keeps the same 7B visual generation architecture. It loads directly with QwenImage21Pipeline in Diffusers. The checkpoint ships with its recommended 8-step sampling schedule saved inside. Generation uses CFG=1 by default. Prefix KV caching reuses the text and reference-image context across denoising steps. How does Qwen-Image-2.1-Turbo work? The underlying architecture is a single-stream DiT with 32 layers and 7B parameters. It uses block-causal attention. Text tokens get a token-level causal mask, while images use a chunk-level bidirectional mask. Text encoder: Qwen3-VL 8B encodes both instructions and condition images. VAE: a 64-channel RGBA autoencoder with 16x spatial compression, which enables native transparency. Scheduler: Flow Matching with Euler discrete scheduling and dynamic shifting. The attention design is what makes prefix caching work. Input images and text are computed once at the first step. Every later step reuses that cache. With only 8 steps, the cached prefix covers most of the conditioning cost. What can it generate and edit? The Turbo model card showcase covers 8 categories. These include portraits, human poses, transparent images, typography and posters, and UI layouts. Editing examples include single-image transformation, multi-reference composition and 4-image interior composition. The base model supports up to 10 reference images and local edits via circles, painted annotations or masks. How do you run Qwen-Image-2.1-Turbo? Setup requires Diffusers from source, along with transformers>=5.17.0. The checkpoint needs Diffusers PR #14950, which adds pipeline-configured sampling sigmas. Copy CodeCopiedUse a different Browserimport torch from diffusers import QwenImage21Pipeline pipe = QwenImage21Pipeline.from_pretrained( "Qwen/Qwen-Image-2.1-Turbo", dtype=torch.bfloat16 ).to("cuda") image = pipe( prompt="A ceramic teapot on a wooden table, soft window light", width=2048, height=2048, use_kv_cache=True, ).images[0] For editing, pass image=input_image with an instruction prompt. Supported presets run from 2048×2048 square to 2752×1536 at 16:9. One thing to note here. Setting num_inference_steps alone does not override the saved schedule. Only an explicit sigmas argument does, and Qwen notes other schedules are untested. What does the API cost? Alibaba Cloud Model Studio now hosts both Turbo and Pro. qwen-image-2.1-turbo costs CNY 0.1 per image in most regions, with a 120 RPM limit. qwen-image-2.1-pro costs CNY 0.25 per image, with 20 RPM. Turbo is 2.5x cheaper per image and allows 6x the request rate. Qwen-Image-2.1-Turbo vs competing fast image models Key Takeaways Qwen-Image-2.1-Turbo cuts denoising from 40 steps to 8 on the same 7B architecture. One checkpoint handles 2K text-to-image, multi-reference editing and transparent RGBA output. The API costs CNY 0.1 per image, 2.5x cheaper than Pro. Weights are research-licensed, so commercial self-hosting needs separate permission. No Turbo-specific benchmark exists yet; base Qwen-Image-2.1 scores 60.28 on Qwen-Image-Bench. Check out the ModelScope page, Hugging Face page, Pro API and Turbo API. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. The post Alibaba Qwen Releases Qwen-Image-2.1-Turbo, an 8-Step 7B Image Model appeared first on MarkTechPost.
OpenAI has released the Decisions API in public beta. It turns text and images into typed answers your code can branch on. OpenAI team states the OpenAI Decisions API runs about 10x faster than the Responses API. It targets a common pattern: prompt an LLM, then parse its text into a label. TL;DR Size: GPT-6 Luna parameter count is not disclosed. Its model card lists a 1,050,000-token context window. Runs on: OpenAI-hosted API only, via POST /v1/decisions. No open weights, no self-hosting. Performance: About 10x faster than the Responses API, per OpenAI. Best: $0.10 per 1M input tokens, with no output, cache-read or cache-write charges. Bottom line: Best: fast, typed decisions with probabilities. Worst: 1 model, beta status, no independent evals yet. What is the OpenAI Decisions API? The Decisions API is an OpenAI endpoint that evaluates text, images or both and returns typed answers. It does not generate prose. A request has 3 fields: model, input and questions. The only supported model today is gpt-6-luna. OpenAI expects general availability in the coming weeks. How does the Decisions API work? Each request carries shared evidence plus a list of questions. Each question has a unique name, a type and instructions. The response returns an answers array keyed by those names. What are the 3 question types? predicate: checks a condition and returns a probability from 0 to 1. Example: does a product photo show a crack, tear or dent? choice: picks 1 value from options you supply. It also returns per-option probabilities and a confidence field. score: rates input against ordered levels, indexed from 0. The score is a probability-weighted average of level indices. OpenAI’s severity example makes the math concrete. Level probabilities of 0.1, 0.7 and 0.2 yield a score of 1.1. That value sits between ‘Workaround available’ and ‘Fully blocked’. When should you use Structured Outputs instead? OpenAI draws a clear line. Use Decisions for probabilities, choices or scores. Use Structured Outputs to fill your own JSON schema or write explanations. Use function calling when a model must request a tool call with arguments. How fast is it, and what is the evidence? OpenAI’s document claims about 10x faster responses than the Responses API. DevDay coverage put a decision near 150 ms, versus about 1.6 seconds for regular Luna calls. OpenAI has not published accuracy or calibration data for the endpoint. The docs advise setting thresholds with labeled examples from your own application. window.addEventListener('message',function(e){ if(e.data&&e.data.mtpH){var f=document.getElementById('mtp-dec-frame');if(f&&e.source===f.contentWindow){f.style.height=e.data.mtpH+'px';}} }); How much does it cost, and where can you deploy it? With gpt-6-luna, input costs $0.10 per 1M tokens. There are no output-token, cache-read or cache-write charges. Regional processing premiums and long-context multipliers still apply. Luna’s model card prices prompts above 272K tokens at 2x input rates. Endpoint: POST /v1/decisions on the OpenAI API, plus a Playground. SDK minimums: Python 3.26.0, JavaScript 7.30.0, Go 3.73.0, Ruby 0.101.0, Java 4.78.0. Compliance: Zero Data Retention and HIPAA for eligible customers. Residency: United States and Europe (EEA plus Switzerland). Voice: decisions can drive actions via client delegation with the Live API. How does it compare with TypeSafe Jev? The closest rival is TypeSafe Jev, launched September 15, 2026. Jev is a System One model that returns typed values with calibrated probabilities. TypeSafe prices input at $0.042 per 1M tokens, with free output. That makes OpenAI’s base rate about 2.4x higher. TypeSafe reports 70 to 500 ms end-to-end, measured from the US West Coast. Jev supports up to 255 choices but remains in early access. OpenAI’s edge is image input, compliance options and open beta access. Key Takeaways OpenAI Decisions API returns probabilities, choices and scores, not prose. Input costs $0.10 per 1M tokens; output is free. OpenAI claims about 10x faster than the Responses API. TypeSafe Jev is cheaper at $0.042 per 1M input tokens. Check out the Technical details here. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. The post OpenAI Decisions API Hits Public Beta With 10x Faster Typed Answers appeared first on MarkTechPost.