Breaking
Advertisement

AI

AI

Anthropic’s Claude can now orchestrate up to 1,000 AI agents in parallel through dynamic workflows

Anthropic is adding dynamic workflows to Claude Managed Agents, letting a lead agent distribute tasks across up to 1,000 sub-agents at once. In testing, a single agent found at most 27 of 70 hidden bugs in a codebase, while the multi-agent workflow consistently caught 66. The article Anthropic's Claude can now orchestrate up to 1,000 AI agents in parallel through dynamic workflows appeared first on The Decoder.

AI

Anthropic launches a free AI scanner for open-source projects

Anthropic has launched "Cyber Mission," a program to protect critical infrastructure and open-source software from cyberattacks. Partners like CrowdStrike and Palo Alto Networks will help secure power grids and water systems, while a free AI scanner automatically checks open-source projects for vulnerabilities with an expected accuracy above 90 percent. The article Anthropic launches a free AI scanner for open-source projects appeared first on The Decoder.

AI

OpenAI revenue keeps surging as company seeks $30 billion in fresh capital

OpenAI's annualized revenue rate sits at about $50 billion, well below the initially reported $70 billion figure that was based on a different accounting method. The correction sent chip stocks sliding. Meanwhile, OpenAI is negotiating at least $30 billion in fresh capital at a $1.4 trillion valuation. The article OpenAI revenue keeps surging as company seeks $30 billion in fresh capital appeared first on The Decoder.

Advertisement
AI

OpenAI’s safety crisis keeps getting worse and the company keeps making it worse

OpenAI fired three safety researchers who helped investigate the Hugging Face hack. In an open letter, they warn that the firings are scaring remaining staff and eroding safety culture. OpenAI claims they violated policies but won't say how. The article OpenAI's safety crisis keeps getting worse and the company keeps making it worse appeared first on The Decoder.

AI

Anthropic’s Claude Science creates the first complete ultraviolet map of the sky

Astrophysicist Brice Ménard of Johns Hopkins University used Anthropic's Claude Science to map the entire sky in ultraviolet light for the first time. AI agents downloaded data from multiple space missions, calibrated it, and filled in gaps using inpainting. Predictions averaged about ten percent deviation from actual measurements. Ménard sees the project as an example of research that simply wouldn't have gotten done without AI. The article Anthropic's Claude Science creates the first complete ultraviolet map of the sky appeared first on The Decoder.

AI

OpenAI uncovers Russian and Iranian influence ops that planted fake stories in real news outlets

OpenAI exposed a Russian and an Iranian influence operation and banned the accounts involved. The Russian "Dark Clark" campaign spread disinformation across Latin America and drew reactions from politicians, earning the first category 5 out of 6 rating in OpenAI's reporting history. The Iranian "Bogus Bylines" operation used seven fake journalists to place nearly 100 articles in online outlets worldwide. The article OpenAI uncovers Russian and Iranian influence ops that planted fake stories in real news outlets appeared first on The Decoder.

Advertisement
AI

Meet the Underdog Saluki 27B: A 2-bit Qwen3.8-27B That Beats the Original at Tool Calling

Underdog, the on-device assistant from Conway Research, has released Saluki 27B under Apache 2.0. Underdog Saluki 27B is a 2-bit GGUF of Qwen3.8-27B that fits in 7.89 GB. The full BF16 model needs 54 GB. Underdog tuned the compression to protect tool calling, the skill that turns a chat model into an agent. For developers, that means a 27B-class agent model that runs in stock llama.cpp. TL;DR Size: 27B dense parameters. 7.89 GB GGUF versus 54 GB for BF16. Runs on: stock llama.cpp and apps built on it, with full GPU offload. Optional 629 MB or 928 MB vision add-on. Performance: 96% average retention across 9 benchmarks versus full Qwen3.8-27B. Best: Parallel tool calls, 42 versus 35 for the full model (120% retention). Worst: AIME 2025, 79.2 versus 96.7 (about 82% retention). Bottom line: Best: beats the 54 GB original at tool calling in a sub-8 GB file. Worst: competition math and multi-step reasoning drop 12 to 18 points. What is Underdog Saluki 27B? Saluki 27B is a 2-bit, mixed-precision GGUF of Qwen3.8-27B built for local agents. It stacks 3 layers of work: The base is Qwen3.8-27B, a dense 27B model from the Qwen team. It has 64 layers, mixes Gated DeltaNet linear attention with gated attention, and supports 262,144 tokens natively. The second layer is ISTA-DASLab’s Qwen3.8-27B-GSQ-RCO-GGUF. GSQ learns accurate low-bit scalar grids per tensor. RCO assigns a quantization type to each tensor under a fixed size budget. ISTA’s smallest file, IQ2_XS, is 8.4 GB at 2.50 bits per weight. The third layer is Underdog’s own pass. It shrank the file to 7.89 GB and targeted tool calling. The file is named IQ2-mix and carries an imatrix tag. Underdog has not published the full recipe for this pass. How does Saluki perform on benchmarks? Underdog splits its results into 2 groups. The first group ran both models in the same harness: Underdog Bench: 120 tasks from BFCL v4, frozen before testing. Thinking off, temperature 0. Saluki scores 88, the full model 84, and PrismML’s Bonsai 2 scores 70. Parallel tool calls: 100 BFCL v4 parallel tasks with the official checker. Saluki 42, full model 35. SWE-bench Verified: 50 issues. Saluki fixes 30, the full model 33. The second group compares Saluki with public full-size scores: window.addEventListener("message",function(e){var f=document.getElementById("mtp-saluki-frame");if(f&&e.source===f.contentWindow&&e.data&&e.data.mtpH){f.style.height=e.data.mtpH+"px";}}); How does Saluki compare with other compact Qwen3.8-27B builds? Bonsai 2 is smaller and reports 98.2% retention across 14 thinking-mode benchmarks. It also posts stronger math, including 95.00 on AIME25. However, it needs PrismML’s llama.cpp fork, since stock llama.cpp rejects its packing formats. Saluki runs on stock llama.cpp. Each vendor uses its own harness, so cross-vendor scores are not directly comparable. Key Takeaways Saluki 27B fits Qwen3.8-27B into 7.89 GB, down from 54 GB. It beats the full model on tool calling: 88 versus 84. Parallel tool calls rise to 42 from 35. Math and reasoning take the biggest hit, down to about 82 to 85%. It runs in stock llama.cpp under Apache 2.0. Check out the Model Card on Hugging Face, the Underdog launch page and the announcement on X. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. The post Meet the Underdog Saluki 27B: A 2-bit Qwen3.8-27B That Beats the Original at Tool Calling appeared first on MarkTechPost.

AI

Google Cloud Launches Gemini Agent, One Universal Agent for Enterprise Work

Google Cloud has introduced the Google Cloud Gemini agent, a single agent for enterprise work. The Gemini agent is a cloud-hosted agent from Google Cloud that answers questions, does knowledge work, creates media, and writes and runs code. It does all of this from 1 prompt box and 1 API. For developers, the agent is the product and the model is a routing decision. TL;DR Runs on: Google Cloud (AI Hypercomputer). Reached from web, mobile, desktop, CLI, Workspace, Microsoft 365, Slack, or headless. Best: Bloomberg Media lifted SQL query accuracy by 63% by grounding data agents in Knowledge Catalog. Bottom line: Best: 1 governed agent with cloud memory, sub-agents and hard spend caps. Worst: buyers must evaluate it on customer anecdotes, not reproducible numbers. What is the Google Cloud Gemini agent? It is a delegation layer, not a chatbot. You give it objectives, not instructions. It plans the work, picks skills and tools, connects to company systems, and returns finished output. Google lists 6 architectural principles: Unified agent: chat, autonomous objectives and code generation share 1 interface. Omnipresent access: any device or channel, plus embedding in third-party apps. Persistent execution: it runs in the cloud with 1 set of memories and 1 personalization graph. Jobs lasting hours or days keep running after you close the laptop. Multi-agent orchestration: it creates temporary sub-agents, each with its own identity, and runs parallel or sequential steps. Deeply contextual: it learns your tools, data and work history over time. Model choice flexibility: each job runs on the best-fit model. How does the Gemini agent remember and reason? It keeps 4 kinds of memory. Session memory covers the current task, even across days. Semantic memory is a structured knowledge base it builds from documents and people. Procedural memory stores how jobs get done, including skills it writes for itself. Episodic memory records everything it has done before. Skills are modular prompts stored in a shared company registry. Tools come from an enterprise tools registry. Connectors cover Slack, Jira, Salesforce, ServiceNow, BigQuery, Snowflake, desktop files and any Model Context Protocol server. Which models does it use? Today it orchestrates across Google’s Gemini models and Claude models from Anthropic, with other private and open models planned. Google’s own lineup is Argon for frontier reasoning, Flash for speed and volume, Omni for generative media, and Gemma for open-weights edge work. What are coworker agents? A coworker agent is a persistent teammate with a defined role. In Workspace it receives its own account: email address, calendar, Drive and a directory entry. Colleagues @mention it in Chat or Docs, and its edits appear under its own name in version history. It sees only what is shared with it. The agent also works inline across Gmail, Docs, Sheets, Slides, Chat and Calendar, and offers 1-click delegation of tasks it spots. What does it add for data teams? Data and ML engineers describe outcomes in plain language. The agent then writes PySpark code, provides notebooks, trains models and fixes pipeline issues. Business users get saved BigQuery reports that rerun without token costs. Three services ground the answers. Knowledge Catalog maps business definitions once for every agent. Smart Storage enriches unstructured objects in place; Google says 90% of enterprise data is unstructured. Borderless Lakehouse queries Amazon S3 and Azure Data Lake with no variable egress fees. window.addEventListener('message',function(e){if(e.data&&e.data.type==='mtp-resize'&&e.data.mtpHeight){var f=document.getElementById('mtp-gem-frame');if(f&&e.source===f.contentWindow){f.style.height=e.data.mtpHeight+'px';}}}); How is the agent governed? Google frames governance as 4 questions: who, what it may do, what it did, and what it must never touch. Identity: each agent gets a cryptographically attested identity with least-privilege permissions. Authorization: role-based access, mapped to external systems through standards such as OAuth. Auditing: every action is logged to the agent, not a person. Policy: agents run in an Agent Sandbox. All traffic passes through Agent Gateway, an AI network firewall that applies 1 policy to every agent. How does Google control cost? Three levers: multi-model orchestration, Smart Routing that triages workloads to the cheapest capable model, and real-time spend caps. Teams set a hard project limit in the Cloud Billing Console. When it triggers, that project’s agent pauses until someone resumes it. Underneath, Google team states its TPU 8i system delivers 80% better price-performance than the prior generation. How does it compare? Key Takeaways 1 agent, 1 API: Q&A, knowledge work, media and code. Routes jobs across Gemini and Claude models today. Coworker agents get their own Workspace identity. Hard per-project spend caps pause runaway agents. No benchmarks, price or GA date yet. Check out the Google Cloud announcement. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. The post Google Cloud Launches Gemini Agent, One Universal Agent for Enterprise Work appeared first on MarkTechPost.

AI

Google Research RRSI Guide: Mastering Self-Improving AI Agents

In this tutorial, we implement RRSI (Regularized Recursive Self-Improvement), a method that lets an LLM agent rewrite its own harness, prompts, tools, memory, control flow, and sub-agents around a frozen model, without the harness overfitting to the tasks it evolves on. The full RRSI loop drafts edits with Claude Opus on Vertex AI and scores them inside Docker benchmarks, which is not something a free notebook can run. The part of RRSI that actually carries the paper’s idea, the rules that decide which proposed edits to keep, is plain Python, and that is what we drive directly. We install the package from the official repository, walk through its estimator, its calibrated noise band, both branches of its selection algorithm, its annealed edit budget, its deterministic leakage screen, and its edit history, and then plug a simulated agent into RRSI’s own Domain interface. Because we built the simulated environment ourselves, we know the true effect of every edit, which lets us audit RRSI’s decisions against ground truth and compare them with an unregularized search that simply keeps whatever scores highest. Copy CodeCopiedUse a different Browserimport os import sys import json import math import copy import random import tempfile import textwrap import traceback import subprocess import statistics as st from pathlib import Path RESULTS = {} def banner(title): print("n" + "=" * 78) print(title) print("=" * 78) def section(name): def wrap(fn): def run(*a, **kw): banner(name) try: out = fn(*a, **kw) RESULTS[name] = out if isinstance(out, str) else "ok" return out except Exception as e: RESULTS[name] = f"SKIPPED / FAILED -> {type(e).__name__}: {e}" print(f"n[!] {name} did not complete: {type(e).__name__}: {e}") traceback.print_exc(limit=3) return None return run return wrap banner("1. Install RRSI and map the paper onto the code") subprocess.run([sys.executable, "-m", "pip", "install", "-q", "git+"], check=True) from importlib.metadata import version from rrsi.config import RRSIConfig from rrsi.evaluate import TaskResult, EvalResult, aggregate, evaluate from rrsi.calibrate import calibrate from rrsi.selection import Candidate, cost_rule, judge, select_round from rrsi.schedule import edit_budget, budget_table from rrsi.history import History, stall_flag, exploration from rrsi.components import K, K_STR, normalize, novelty from rrsi.critic import precheck, review from rrsi.domain import Domain print(f" rrsi {version('rrsi')} | anthropic {version('anthropic')} | Python {sys.version.split()[0]}") print("n RRSI evolves an agent's HARNESS (prompts, tools, memory, control flow, sub-agents) around a") print(" frozen model. The full loop drafts edits with Claude Opus on Vertex AI and scores them in Docker") print(" benchmarks. The part that decides which edits to KEEP is plain Python, and that is what we drive:") for symbol, where in [ ("S_hat, C_hat Eq. (estimate)", "rrsi.evaluate.aggregate"), ("delta noise band", "rrsi.calibrate.calibrate"), ("Algorithm 2 floor + cost rule", "rrsi.selection.judge / select_round"), ("b_t Eq. (anneal)", "rrsi.schedule.edit_budget"), ("Critic leakage screen", "rrsi.critic.precheck / review"), ("L_t, g_t, B_t history, yield, prune", "rrsi.history.History"), ("sigma_t, U_t stall + exploration", "rrsi.history.stall_flag / exploration"), ("nu structural novelty", "rrsi.components.novelty"), ]: print(f" {symbol:40s} -> {where}") CFG = RRSIConfig() print(f"n paper defaults: T={CFG.T} rounds, k={CFG.k} trials/task, m={CFG.m} candidates/round," f" b in [{CFG.b_min},{CFG.b_max}]") print(f" beta0={CFG.beta0} beta1={CFG.beta1} w_s={CFG.w_s} w_c={CFG.w_c} w_n={CFG.w_n}" f" delta_z={CFG.delta_z}") print("n Nothing below needs an API key, a GPU or a dataset download.") We install RRSI from the google-research repository, pinned to the commit this notebook was written against, since the package is not on PyPI. Its only dependency is the Anthropic client, which the search roles use to call Claude and which we never exercise. We then print the mapping the repository itself documents between the paper’s symbols and the functions that implement them: the empirical score and cost estimate in evaluate, the noise band in calibrate, Algorithm 2 in selection, the annealed edit budget in schedule, the leakage screen in critic, and the edit history with its yield, prune, stall and exploration summaries in history. RRSIConfig holds the paper’s hyperparameters, and every function below receives it exactly as the real loop does. Copy CodeCopiedUse a different Browser@section("2. Evaluate(H): a score and a cost, and why a crash counts as zero") def estimator(): base = { "task_000": TaskResult(rewards=[1, 1], tokens=[11_800, 12_400]), "task_001": TaskResult(rewards=[1, 0], tokens=[15_100, 14_600]), "task_002": TaskResult(rewards=[0, 0], tokens=[21_000, 19_500]), } ev = aggregate("H0", 2, base) print(f" three tasks x k=2 trials -> S_hat = {ev.S:.3f} C_hat = {ev.C:,.0f} tokens/trial" f" ({ev.n_expected} trials expected, {ev.missing} missing)") crashy = dict(base) crashy["task_002"] = TaskResult(rewards=[0.0, 0.0], tokens=[None, None], missing=2) ev_crash = aggregate("crashy", 2, crashy) dropped = {t: r for t, r in crashy.items() if not r.missing} naive = sum(sum(r.rewards) for r in dropped.values()) / sum(len(r.rewards) for r in dropped.values()) print("n A candidate crashes on the hardest task instead of failing it:") print(f" an estimator that drops missing trials reports {naive:.3f} <- looks like a gain") print(f" RRSI's aggregate (missing = 0, full denominator) {ev_crash.S:.3f} <- no reward for crashing") print(f" ...and C_hat uses only recorded token counts: {ev_crash.C:,.0f}") rubric = {"memo": TaskResult(rewards=[0.5, 1.0], weights=[10, 10]), "brief": TaskResult(rewards=[0.0, 0.0], weights=[90, 90])} ev_w = aggregate("rubric", 2, rubric) print("n Weighted rewards (Harvey LAB style: weight = number of rubric criteria):") print(f" mean of per-task means = {st.mean(r.mean for r in rubric.values()):.3f}" f" vs RRSI's S_hat = {ev_w.S:.3f} (fraction of all criteria passed)") return f"crash scored {ev_crash.S:.3f} under RRSI vs {naive:.3f} if dropped" estimator() RRSI measures two numbers per harness: S, the reward averaged over every trial of every task, and C, the mean policy tokens per trial. TaskResult records one task’s trials and aggregates them. The detail worth copying into any agent evaluation is how missing trials are handled. When a candidate crashes on the hardest task, an estimator that drops the missing trials reports 0.750 and makes the crash look like an improvement. At the same time, RRSI counts each missing trial as zero reward with the full denominator and reports the same 0.500 as before, so a candidate cannot look better by destroying the trials it finds hard. Weighted rewards cover rubric-graded suites such as Harvey LAB, where S becomes the fraction of all criteria passed rather than the mean of per-task means. Copy CodeCopiedUse a different BrowserFAMILIES = ["parse", "search", "edit", "test"] H0 = {"skill": {f: 0.0 for f in FAMILIES}, "memo": [], "cost": 1.0} def make_world(seed, n_evolve, n_heldout=80): """Tasks with a family and a difficulty. Evolve ids are task_NNN, held-out ids held_NNN.""" r = random.Random(seed) world = {f"task_{i:03d}": (FAMILIES[i % 4], r.gauss(0, 1)) for i in range(n_evolve)} world.update({f"held_{i:03d}": (FAMILIES[i % 4], r.gauss(0, 1)) for i in range(n_heldout)}) return world def p_success(h, world, task): """The frozen policy: a logistic in harness skill minus task difficulty, or 0.97 if memorised.""" if task in h["memo"]: return 0.97 family, difficulty = world[task] return 1 / (1 + math.exp(-(0.3 + h["skill"][family] - difficulty))) def run_trials(h, world, ids, k, rng): return {t: TaskResult(rewards=[float(rng.random() <p>12s} {'trials':>7s} {'delta':>8s} method") deltas = {} for n, k in [(40, 2), (100, 4), (400, 8)]: cal, _ = calibrated_delta(make_world(0, n), k, seed=0) deltas[f"{n}x{k}"] = cal["delta"] print(f" {f'{n} x k={k}':>12s} {n * k:>7d} {cal['delta']:>8.4f} {cal['method']}") print("n The paper's calibrated bands are 0.017 (coding), 0.004 (workspace) and 0.020 (engineering).") print(" A gain smaller than delta is indistinguishable from re-running the same harness, and") print(" Algorithm 2 treats it that way. Remember the first row: it becomes the lesson of step 11.") return "delta " + ", ".join(f"{k}={v:.3f}" for k, v in deltas.items()) noise_band() Before any rule can separate a real gain from luck, it needs to know how far one harness’s score moves on its own. We build a small simulated agent, whose success on each task is a logistic function of harness skill minus task difficulty, and evaluate the unchanged starting harness six times on forty tasks with two trials each: the scores spread by 0.113 although nothing changed. calibrate turns repeated evaluations of the same harness into delta, twice the standard deviation of the difference between two runs. With 80 trials delta is about 0.108; with 3,200 trials it falls to about 0.013, in the range the paper reports for its instances (0.004 to 0.020). For selection purposes, any gain smaller than delta is indistinguishable from re-running the same harness. Copy CodeCopiedUse a different Browserdef ev_at(S, C, job, n=200): """An EvalResult with exactly score S (in steps of 1/n) and C tokens per trial.""" hits = round(S * n) return aggregate(job, 1, {f"task_{i:03d}": TaskResult(rewards=[1.0 if i 6,} {'ADMIT ' if dec.admissible else 'reject'}") print(textwrap.indent(textwrap.fill(dec.reason, 88), " " * 6)) print("n Three rules, in the order RRSI applies them:") print(" 1. never fall below the best score ever seen, minus the noise band (A)") print(" 2. a gain bigger than delta must pay for any extra tokens: dC 2d}" for t in range(CFG.T))) print(" b_t : " + " ".join(f"{b:>2d}" for b in table)) print(f" round {CFG.T} (after the run) -> {edit_budget(CFG.T, CFG.T, CFG.b_min, CFG.b_max)}") print("n Early candidates may bundle up to 4 coordinated edits. Note the ceil(): the cosine term is") print(f" only exactly zero at t = T, so inside a {CFG.T}-round run the budget bottoms out at" f" {min(table)}, not {CFG.b_min}.") print(" Every edit in a bundle inherits the bundle's single measurement, so the shrinking budget is") print(" what makes late history attributable to fewer components. It caps how MANY edits ride") print(" together, never WHICH mechanisms the harness may eventually contain.") return "budget " + "".join(str(b) for b in table) edit_budget_schedule() The proposal side regularizes how edits are drafted, not which ones are kept. edit_budget implements the annealed L0 budget from the paper: a cosine schedule from b_max to b_min, and by default it allows up to four coordinated edits per candidate for the first eight rounds, three for the next five, and two for the last seven. The ceiling in the formula means the budget only reaches its minimum of one at t = T, one step after the run ends, which is easy to miss when reading the equation. Because every edit in a bundle inherits the bundle’s one measurement, the shrinking budget is what makes late-run history attributable to fewer components. The budget limits how many edits travel together and never restricts which mechanisms the harness may eventually contain. Copy CodeCopiedUse a different BrowserCRITIC_PATTERNS = [ (r"btask_d{3}b", "hard-codes an evolve-set task id"), (r"expected_output|grader|rubric[", "reads the grader or the expected answer"), ] COMPONENT_SIGNALS = [ ("control_flow", [r"bretry(", r"max_attempts"]), ("context_mgmt", [r"compress_context", r"keep_last"]), ("config", [r"CONFIG["]), ] class ScreenOnly(Domain): name = "screen" critic_patterns = CRITIC_PATTERNS briefs = {"critic": "A coding agent harness."} @section("7. The critic's deterministic layer, and how edits are tagged") def critic_and_tags(): dom = ScreenOnly() diffs = { "memorise answers": "+ memory = Memory('answers')n+ memory.remember('task_007', cached_patch)", "peek at the grader": "+ if os.path.exists('/grader/expected_output.txt'): return read_it()", "leaked credential": "+ api_key = 'sk-live-0123456789abcdefghijkl'", "empty diff": " ", "general retry rule": "+ for attempt in range(max_attempts): result = retry(step)", } for label, diff in diffs.items(): try: verdict = review(dom, diff, summary=label, targets_mode="evolve") print(f" {label:20s} -> {verdict['verdict']:6s} {verdict['reasons']}") except ZeroDivisionError as e: print(f" {label:20s} -> passed the deterministic layer; the LLM layer raised" f" ZeroDivisionError: {e}") print("n Gotcha: with RRSI_VERTEX_PROJECTS unset, rrsi.llm.generate computes `x % len(projects)`") print(" outside its try block, so the helpful 'set RRSI_VERTEX_PROJECTS' error is never reached.") print(" A clean diff is supposed to go to Claude for an intent review; here there is no Claude.") print("n Every edit is tagged with the component it touches, and a tag needs evidence in the diff:") tag_cases = [ ("skill", "+ "Remember to run the tests before finishing.""), ("control_flow", "+ for attempt in range(max_attempts): result = retry(step)"), (None, "+ review = subcall('reviewer', transcript)"), ("memory", "+ context = compress_context(context, keep_last=8)"), ] for declared, diff in tag_cases: tag = normalize(declared, diff, COMPONENT_SIGNALS) print(f" declared {str(declared):13s} -> tagged {tag:13s} {diff[2:60]!r}") print(" A proposer cannot label a prompt tweak as a new 'skill' to look novel: without evidence") print(" the tag falls back to what the diff actually is.") counts = {"prompt": 3, "subagent": 1} print(f"n novelty(nu) counts STRUCTURAL components {K_STR} the incumbent has never accepted.") print(f" against an incumbent with accepted edits {counts}:") for comps in (["memory"], ["subagent"], ["prompt", "client_tool", "memory"]): print(f" {str(comps):38s} nu = {novelty(comps, counts)}") return "precheck rejected 4/5 diffs without an LLM call" critic_and_tags() The critic screens every candidate diff before any evaluation is spent, in two layers. The first is a deterministic precheck against a generic credential pattern plus the domain’s own denylist; with patterns for evolve-set task ids and grader paths it rejects a diff that memorises the answer to task_007, one that reads the expected output, one that leaks an API key, and an empty diff, all without calling a model. A clean diff falls through to the second layer, an intent review by Claude, and here the notebook surfaces a real gotcha: with RRSI_VERTEX_PROJECTS unset, rrsi.llm.generate computes an index modulo the number of projects outside its try block and raises ZeroDivisionError, so the helpful configuration error in the code is never reached. We also look at how edits are tagged: normalize keeps a declared component only when the diff carries evidence for it, so a prompt tweak cannot pose as a new skill to win the novelty bonus, and a mislabelled context-management change is tagged as what it is. Copy CodeCopiedUse a different Browser@section("8. The edit history: what was tried, what paid off, what to prune") def edit_history(): with tempfile.TemporaryDirectory() as tmp: h = History(Path(tmp) / "history.jsonl") log = [ (0, "A", [("prompt", "tell the agent to read the failing test first")], "ACCEPTED", 0.030, 0.02, True, 0.66), (1, "A", [("prompt", "ask for a plan before editing")], "REJECTED", -0.010, 0.05, False, 0.65), (1, "B", [("subagent", "add a reviewer sub-agent"), ("memory", "persist lint rules")], "REJECTED", -0.020, 0.40, False, 0.64), (2, "A", [("config", "raise the step limit")], "ACCEPTED", 0.005, -0.03, True, 0.665), (3, "A", [("prompt", "shorter system prompt")], "LOST", -0.001, -0.10, False, 0.664), (4, "B", [("memory", "cache task_014 solution")], "critic_reject", None, None, False, None), (5, "A", [("prompt", "stricter output format")], "REJECTED", -0.004, 0.00, False, 0.661), ] for t, v, edits, outcome, dS, dC, acc, S in log: h.append_candidate(t, v, [{"id": f"C{i + 1}", "component": c, "hypothesis": hyp} for i, (c, hyp) in enumerate(edits)], outcome, dS, dC, acc, S, 12_000 if S else None, diff=None) t_now = 6 print(f" {len(h.records())} per-edit records from {len(log)} candidates" f" (the two-edit bundle in round 1 wrote two records with ONE measurement)") print(f" T_t, tried components : {sorted(h.tried())} (the critic-rejected edit is not 'tried')") g_t = h.yield_g(t_now, CFG.n_prune) print(" g_t, best gain in the last n_prune rounds: " + ", ".join(f"{c} {'none measured' if g == -math.inf else f'{g:+.3f}'}" for c, g in sorted(g_t.items()))) prune = h.prune_set(t_now, CFG.n_prune) print(f" B_t, prune set : {[p['component'] for p in prune]}") for p in prune: if p["accepted_edits_in_incumbent"]: print(f" {p['component']}: still in the incumbent but no recent gain ->" f" {[e['hypothesis'] for e in p['accepted_edits_in_incumbent']]}") trajectory = [0.630, 0.660, 0.660, 0.665, 0.665, 0.665, 0.665] sigma = stall_flag(trajectory, t_now, CFG.w, DELTA) ex = exploration(t_now, sigma, h.tried(), CFG.m_draft) print(f"n S over rounds {trajectory}: moved {trajectory[t_now] - trajectory[t_now - CFG.w]:+.3f}" f" in the last w={CFG.w} rounds -> sigma_t = {sigma}") print(" what the proposer is told next round:") print(textwrap.indent(textwrap.fill(ex["text"], 84), " ")) print("n The proposer sees this history, so a falsified hypothesis ('ask for a plan first', -1pp)") print(" is not redrawn, and a stalled run is pushed toward components it has never touched.") return f"prune set {[p['component'] for p in prune]}, stall flag {sigma}" edit_history() History writes one JSONL record per edit, and the loop derives four summaries from it. The tried set excludes edits the critic dropped, since they were never measured. The recent yield g_t is the best measured gain per component within the last n_prune rounds, and every component whose recent yield is not positive enters the prune set B_t, together with any machinery from that component that is still in the incumbent; in our history that flags a prompt edit accepted in round zero that has not paid off since. The stall flag fires when the score has moved less than delta over the last w rounds, and exploration then writes the directive the proposer receives, reserving a candidate slot for components the run has never exercised. Because the proposer is conditioned on all of this, a falsified hypothesis is not drawn again. Copy CodeCopiedUse a different Browserdef write_harness(root, h): root = Path(root) root.mkdir(parents=True, exist_ok=True) (root / "harness.json").write_text(json.dumps(h)) return root class SimulatedAgentDomain(Domain): """A Domain adapter over the simulated agent: the same contract RRSI's coding, workspace and engineering instances implement.""" name = "simulated" critic_patterns = CRITIC_PATTERNS component_signals = COMPONENT_SIGNALS briefs = {"critic": "A coding agent harness evaluated on parse/search/edit/test tasks."} def __init__(self, world, seed): self.world, self.rng = world, random.Random(seed) def evolve_ids(self): return [t for t in self.world if t.startswith("task_")] def heldout_ids(self): return [t for t in self.world if t.startswith("held_")] def smoke_ids(self, incumbent_per_task=None): return self.evolve_ids()[:3] def run(self, root, runs_dir, job, ids, k, log_prefix=""): out = Path(runs_dir) / "jobs" / job if (out / "trials.json").exists(): return # resume-safe, as the contract requires h = json.loads((Path(root) / "harness.json").read_text()) per = run_trials(h, self.world, ids, k, self.rng) out.mkdir(parents=True, exist_ok=True) (out / "trials.json").write_text(json.dumps( {t: {"rewards": r.rewards, "tokens": r.tokens} for t, r in per.items()})) def score(self, runs_dir, job, ids, k): d = json.loads((Path(runs_dir) / "jobs" / job / "trials.json").read_text()) return {t: TaskResult(rewards=d[t]["rewards"], tokens=d[t]["tokens"]) for t in ids}, {} @section("9. A Domain adapter: plugging an environment into RRSI's own evaluate()") def domain_adapter(): dom = SimulatedAgentDomain(make_world(0, 40), seed=3) with tempfile.TemporaryDirectory() as tmp: root, runs = write_harness(Path(tmp) / "wt_H0", H0), Path(tmp) / "runs" ev = evaluate(dom, root, runs, "H0", dom.evolve_ids(), k=2) files = sorted(str(p.relative_to(tmp)) for p in Path(tmp).rglob("*") if p.is_file()) print(f" evaluate(domain, worktree, runs_dir, 'H0', 40 ids, k=2) -> S={ev.S:.3f} C={ev.C:,.0f}") print(f" files: {files}") again = evaluate(dom, root, runs, "H0", dom.evolve_ids(), k=2) print(f" evaluate() again on the same job -> S={again.S:.3f} (read back, not re-run)") print("n The harness lives in files under a worktree root, exactly as RRSI's real instances keep") print(" one git worktree per candidate. RRSI never runs an agent or grades anything itself; the") print(" adapter's run() and score() do, and everything in steps 2-8 consumes what they return.") return f"adapter evaluated H0 at S={ev.S:.3f} through rrsi.evaluate.evaluate" domain_adapter() RRSI never runs an agent or grades a deliverable itself; a Domain adapter does, and the same interface backs the paper’s coding, workspace and engineering instances. We implement one for the simulated agent: evolve and held-out task splits, a run method that reads the harness from files under a worktree root and writes trial results under a runs directory, a score method that reads them back as TaskResult objects, and the domain’s critic patterns and component signals. RRSI’s own evaluate function then scores the starting harness through the adapter, and calling it a second time on the same job reads the stored trials back rather than re-running them, the resume-safety the contract requires. Copy CodeCopiedUse a different Browserdef propose_edit(r, evolve_ids): """The scripted proposer. Each draw is one edit whose TRUE effect we know.""" u = r.random() if u < 0.35: fam = r.choice(FAMILIES) comp = r.choice(["prompt", "control_flow", "context_mgmt"]) diff = {"prompt": f"+ "On {fam} tasks, check the edge cases before finishing."", "control_flow": f"+ for attempt in range(max_attempts): result = retry({fam}_step)", "context_mgmt": f"+ context = compress_context(context, keep_last=12) # {fam}"}[comp] return {"kind": "general", "component": comp, "family": fam, "effect": r.gauss(0.20, 0.30), "dcost": 0.02, "diff": diff} if u < 0.55: ids = r.sample(evolve_ids[:40], 3) return {"kind": "leaky", "component": "memory", "ids": ids, "dcost": 0.03, "diff": "+ memory = Memory('solutions')n" + "n".join( f"+ memory.remember('{i}', cached_patch)" for i in ids)} if u < 0.70: return {"kind": "inert", "component": "config", "dcost": r.uniform(-0.02, 0.05), "diff": "+ CONFIG['log_level'] = 'debug'"} if u inc.S] winner = max(live, key=lambda c: c.ev.S) if live else None for c, (hc, edits) in zip(cands, drafts): audit += [(e["kind"], c is winner) for e in edits] if winner is not None: h, inc = drafts[cands.index(winner)][0], winner.ev S_star = max(S_star, inc.S) return {"evolve_measured": inc.S, "evolve_true": true_score(h, world, ids), "heldout_true": true_score(h, world, dom.heldout_ids()), "tokens": h["cost"], "memorised": len(h["memo"]), "evals": n_evals, "audit": audit} MODES = ["greedy", "critic only", "rrsi"] def compare(n_evolve, k, seeds): out = {mode: [] for mode in MODES} deltas = [] for s in seeds: world = make_world(s, n_evolve) cal, _ = calibrated_delta(world, k, seed=s) deltas.append(cal["delta"]) for mode in MODES: out[mode].append(search(world, mode, seed=s, k=k, delta=cal["delta"])) return out, st.mean(deltas) def fmt(xs, d=3): return f"{st.mean(xs):.{d}f}±{st.pstdev(xs):.{d}f}" @section("10. Greedy vs RRSI in a world where we know the truth") def miniature(): seeds = range(8) out, delta = compare(40, 2, seeds) globals()["MINIATURE"] = (out, delta) # step 11 reuses this row h0_held = st.mean(true_score(H0, make_world(s, 40), [f"held_{i:03d}" for i in range(80)]) for s in seeds) print(f" 40 evolve tasks x k=2, 80 held-out tasks, T={CFG.T}, m={CFG.m}, {len(seeds)} seeds," f" mean calibrated delta {delta:.3f}") print(f" H0 held-out (true) = {h0_held:.3f}n") print(f" {'mode':12s} {'evolve meas':>12s} {'evolve TRUE':>12s} {'held-out TRUE':>14s}" f" {'tokens':>11s} {'memorised':>10s} {'evals':>6s}") for mode in MODES: rs = out[mode] print(f" {mode:12s} {fmt([r['evolve_measured'] for r in rs]):>12s} {fmt([r['evolve_true'] for r in rs]):>12s}" f" {fmt([r['heldout_true'] for r in rs]):>14s} {fmt([r['tokens'] for r in rs], 2) + 'x':>11s}" f" {st.mean(r['memorised'] for r in rs):>10.1f} {st.mean(r['evals'] for r in rs):>6.1f}") print("n Ground-truth audit: share of proposed edits of each kind that ended up in the incumbent") kinds = ["general", "expensive", "compress", "inert", "leaky"] print(f" {'mode':12s}" + "".join(f"{k:>11s}" for k in kinds)) for mode in MODES: tally = {k: [0, 0] for k in kinds} for r in out[mode]: for kind, accepted in r["audit"]: tally[kind][0] += 1 tally[kind][1] += accepted print(f" {mode:12s}" + "".join(f"{tally[k][1]:>5d}/{tally[k][0]:14s} {'memorisation':>14s} (evolve meas - evolve true | evolve true - held-out true)") for mode in MODES: curse = st.mean(r["evolve_measured"] - r["evolve_true"] for r in out[mode]) memo = st.mean(r["evolve_true"] - r["heldout_true"] for r in out[mode]) print(f" {mode:12s} {curse:>+14.3f} {memo:>+14.3f}") print(" Selecting the best of noisy scores inflates every mode about equally; no rule here removes") print(" that - only re-measuring on tasks the search never saw does. The critic removes almost all") print(" of the memorisation, and it does so before evaluation, which is why its runs cost fewer evals.") g, c, r_ = (st.mean(x["heldout_true"] for x in out[m_]) for m_ in MODES) tg, tc, tr = (st.mean(x["tokens"] for x in out[m_]) for m_ in MODES) return f"held-out {g:.3f} / {c:.3f} / {r_:.3f}, tokens x{tg:.2f} / x{tc:.2f} / x{tr:.2f} (greedy / critic / rrsi)" miniature() Now we run the whole selection side as a search: twenty rounds, two candidates a round, bundles sized by the annealed budget, edits tagged by normalize, candidates screened by precheck and scored through the adapter, winners chosen by select_round, and every outcome written to History in the order RRSI’s loop uses. The scripted proposer draws five kinds of edit whose true effects we know: usually helpful general changes, leaky edits that memorize evolve-task answers, inert configuration changes, expensive sub-agents that help a little everywhere at 1.5 times the tokens, and cheaper context compression. We compare three acceptance rules on the same candidate stream over eight seeds: greedy keeps the best score if it rose, critic only adds the leakage screen, and RRSI adds Algorithm 2 with a calibrated delta. Greedy reaches the highest held-out score, 0.681 against 0.616 for RRSI, while memorizing about ten answers and nearly tripling its token cost; RRSI memorizes none, ends at about half the token cost (1.53 times the starting harness against 2.97 times), and uses a third fewer evaluations. Splitting the evolve-to-held-out gap is the most useful result here: selecting the best of noisy scores inflates every mode by about nine points, which no rule in this loop removes. At the same time, the critic eliminates almost all of the memorization component. Copy CodeCopiedUse a different Browser@section("11. Turn the evaluator up: RRSI's caution is calibrated, not configured") def noise_sweep(): rows = [("40x2", MINIATURE[1], MINIATURE[0], 8)] # from step 10 for n, k in [(100, 4), (400, 8)]: out, delta = compare(n, k, range(5)) rows.append((f"{n}x{k}", delta, out, 5)) print(f" {'evaluator':>10s} {'seeds':>5s} {'delta':>7s} " + "".join(f"{m_ + ' held / tokens':>26s}" for m_ in MODES)) for label, delta, out, n_seeds in rows: cells = "".join(f"{st.mean(r['heldout_true'] for r in out[m_]):>13.3f} /" f" x{st.mean(r['tokens'] for r in out[m_]):10s} {n_seeds:>5d} {delta:>7.3f} {cells}") first, last = rows[0], rows[-1] rr_first = st.mean(r["heldout_true"] for r in first[2]["rrsi"]) rr_last = st.mean(r["heldout_true"] for r in last[2]["rrsi"]) tok_ratio = (st.mean(r["tokens"] for r in last[2]["greedy"]) / st.mean(r["tokens"] for r in last[2]["rrsi"])) print(f"n As the evaluator sharpens, delta falls from {first[1]:.3f} to {last[1]:.3f} (the paper: 0.004-0.020)," f" and RRSI's held-out score rises from {rr_first:.3f} to {rr_last:.3f}.") print(f" At the sharpest setting the unregularized search is spending {tok_ratio:.1f}x RRSI's tokens.") print("n Read the table honestly: this world has no diminishing returns, so every sub-agent the") print(" greedy search stacks keeps buying accuracy. That is the most favourable world possible for") print(" spending, and the greedy search does score higher. RRSI trades some of that score for a") print(" bounded token bill, zero memorised answers and fewer wasted evaluations - and the size of") print(" the trade is set by delta, which it measures from your evaluator rather than taking from you.") return (f"delta {first[1]:.3f} -> {last[1]:.3f}; RRSI held-out {rr_first:.3f} -> {rr_last:.3f};" f" greedy uses {tok_ratio:.1f}x the tokens") noise_sweep() Our first evaluator was noisy, and delta was calibrated from it, so RRSI’s caution in step 10 reflected the evaluator rather than the method. We repeat the comparison with larger evolve sets and more trials. As the evaluator sharpens, delta falls from 0.096 to 0.025, and RRSI’s held-out score rises from 0.616 to 0.759, while the unregularized search at the sharpest setting uses 6.5 times as many tokens as RRSI. Greedy still scores higher, and the notebook says why: this simulated world has no diminishing returns, so every sub-agent the greedy search stacks keeps buying accuracy, which is the most favorable world possible for spending. RRSI trades part of that score for a bounded token bill, no memorized answers, and fewer wasted evaluations, and the size of the trade is set by delta, which it measures rather than asks for. Copy CodeCopiedUse a different Browserbanner("SUMMARY") for name, res in RESULTS.items(): print(f" {name:<76s} {res}") print(""" What the miniature does not model - Edits that help the evolve suite but hurt a different suite. The paper's LLM critic and its held-out and out-of-distribution splits exist for those; a regex screen cannot catch them. - A proposer that reads the history. Ours is scripted, so it redraws falsified ideas freely. Where to go next - Run a real instance: `python3 rrsi.py --domain coding smoke` after configuring Claude on Vertex AI (RRSI_VERTEX_PROJECTS) and harbor; see domains/coding/README.md. - Add a domain: implement rrsi.domain.Domain in domains//adapter.py, as step 9 did in the notebook. - Re-adjudicate a stored round under a different delta or cost rule without re-running anything: `python3 rrsi.py --domain readjudicate --t `. - Paper: arxiv.org/abs/2609.24972 Code: github.com/google-research/rrsi """) The summary prints the one-line result each section returned, states plainly what the miniature does not model, namely edits that help the evolve suite while hurting a different suite, which is what the paper’s LLM critic and held-out splits exist for, and a proposer that reads the history, and then points at running a real instance, adding a domain, and re-adjudicating a stored round under a different delta without re-running anything. In conclusion, we took RRSI apart along the line where the paper’s idea actually lives, the rules that decide which edits a self-improving agent keeps, and ran them offline on a CPU without a model or an API key. The estimator refuses to reward a crash, the noise band is measured rather than chosen, and Algorithm 2 turns out to be subtler than ‘reject noisy gains’: it never lets the score fall below the best seen minus the noise band, it makes a real gain pay for its extra tokens, and inside the band it deliberately prefers the cheaper harness, even one that scored slightly lower. Auditing those rules against a world whose ground truth we controlled was the most informative part: the deterministic critic removed almost all of the memorization before any evaluation was spent, Algorithm 2 held the token bill several times below an unregularized search, and neither removed the inflation that comes from picking the best of noisy scores, which only re-measurement on unseen tasks can do. The honest caveat is equally clear: in a world where tokens always buy accuracy, the unregularized search scores higher, so the value of RRSI’s regularization depends on how expensive tokens are and how much a leaked answer would cost you, and the notebook gives you the machinery to measure both on your own evaluator. Check out the FULL CODES here. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. The post Google Research RRSI Guide: Mastering Self-Improving AI Agents appeared first on MarkTechPost.

Advertisement
AI

Claude can now generate animated explainer videos and live data dashboards from text prompts

Anthropic launched two new beta features for Claude. Dashboards turns data sources like BigQuery and Snowflake into live dashboards from text prompts. Motion generates animated explainer videos from text and images. Docs, Slides, and Design now work across all plans, including free accounts. The article Claude can now generate animated explainer videos and live data dashboards from text prompts appeared first on The Decoder.