AI agent populations get their own ecology, complete with a tipping point (ecological safety), code worlds let agents keep evolving (AgentGarten), and flow matching learns physics that refuses to be one-to-one (Bi-FORK). Also: TypeSafe's $7.5B decision models, SynthID Detector opens to everyone, and Claude Haiku 5.5.
TypeSafe AI, the startup behind Jev, a non-text 'decision model' that outputs calibrated probabilities instead of generated text, has raised $870 million at a $7.5 billion valuation in a Series A led by Andreessen Horowitz, with participation from Sequoia and DCVC. The company says Jev runs faster and uses far fewer tokens than LLMs, and that a third of Fortune 500 companies are already using it.
Google has opened its SynthID Detector portal to everyone, globally and in English: anyone can now upload an image, video or audio file to synthid.com and check whether it carries an invisible SynthID watermark. The detector also reads watermarks from partners OpenAI, NVIDIA and Kakao, with Apple coming soon. Google says it has watermarked more than 180 billion images and videos, plus 240,000 years of audio, since 2023.
Anthropic has launched Claude Haiku 5.5, the third model in its Claude 5.5 family, with API prices starting at $0.10 per million input tokens and $0.50 per million output tokens, which Anthropic says makes it about 75% cheaper to run on average than Haiku 4.5. It ships with a 1M-token context window, is the first Haiku with an adjustable effort setting, and scored 72.4% on OSWorld 2.1 in the company's own benchmarks.
DeepSeek is close to raising at least $12 billion in a new funding round backed by Tencent and battery maker CATL, far beyond its original target, Bloomberg reports, with the total possibly approaching $15 billion. The round would rank among the largest private AI fundraises on record, and comes ahead of a planned early 2027 IPO.
AgentGarten builds 'code worlds': interactive environments whose physics and rules live in editable simulator code, while a shared neural renderer turns structured conditions into photorealistic observations in real time. New worlds can be authored by coding agents from a text or image prompt, and agents improve across rounds by distilling each round's experience into playbooks that later agents inherit and refine. The study reports agents learning in 4 rounds what takes a conventional RL counterpart millions of episodes.
What agents can learn is bounded by the worlds they practice in, and building diverse, faithful worlds has been the bottleneck for evolving agents. Code worlds that scale in number and difficulty alongside their agents are the missing training ground for open-ended self-improvement.
The paper builds an ecological theory of AI-agent populations: fitness (growth rate) depends on cybersecurity capability, and with collaboration, collective capability grows with population size. That creates a critical population threshold, a strong Allee effect: below it the population declines, above it the population takes off, even when individual-agent capability stays fixed. Red-teaming a small group, the authors argue, cannot guarantee safety at larger populations.
Agent safety has been studied one agent, or a few agents, at a time. If growth itself can be a phase transition driven by population size, then deployment plans that scale up agent counts are changing the risk regime, not just the workload. The paper's call for ecological red teaming and population pacing gives labs and regulators a concrete, testable frame for the population era.
Bi-FORK is a generative framework for learning one-to-many solution maps at symmetry-breaking bifurcations, where a single input admits multiple equally valid outcomes and standard surrogates average them into nonphysical mush. It generates complete trajectories with latent flow matching, keeping branch choices coherent across space and time, and uses repulsion-guided sampling to recover distinct solution branches in one amortized pass. On buckling beams, mechanical metamaterials and Allen-Cahn phase separation, it scales to discretizations of up to 260,000 points, orders of magnitude beyond prior work.
Generative models are moving into scientific simulation, but physics is full of bifurcations where determinism breaks into multiple valid futures. A method that respects this one-to-many structure, rather than smoothing it away, is what turns generative surrogates from statistical interpolators into instruments scientists can reason with.
PathLang is a language-centered, clinically grounded zero-shot benchmark for pathology vision-language models. Instead of fixing the prompt and varying images, it fixes the slides and labels and varies only the diagnostic language along the axes pathologists actually use: terminology, specificity and reporting style. All prompts and candidate pools were validated by six board-certified pathologists. Across nine VLMs and five datasets spanning four organs, performance proved highly sensitive to clinically equivalent paraphrases, a failure mode canonical-prompt benchmarks cannot see.
Pathology VLMs are being proposed for clinical workflows like fine-grained subtyping and report-level differential diagnosis, where a clinician will phrase the same diagnosis in many ways. A model that passes only under a benchmark designer's chosen prompt is not ready for that world. PathLang makes the language axis a first-class evaluation dimension, which is exactly what clinical deployment requires.
PureVision tackles multi-phenotype lesions in medical vision-language models, whose diagnosis needs joint assessment of several pathological phenotypes. Its PureEyes module supervises visual encoders with ideal spatial distributions that encode anatomical hierarchies and phenotype relationships, replacing text-only semantic supervision that lacks geometric structure. A second module, PureNeurons, selectively aggregates evidence from sparse lesion patches so normal-tissue patches do not drown out the signal. On LIDC-IDRI, CBIS-DDSM and 3DReasonKnee, it improves lesion grounding and phenotype characterization in VQA and radiology report generation.
Real lesions rarely present as textbook single phenotypes, and current medical VLMs smear the distinctions that matter for diagnosis. Giving the visual embedding space a geometry that mirrors anatomy, instead of hoping text embeddings carry the structure, is a principled fix for one of the field's most practical failure modes.
The paper evaluates agents on epistemic humility: the willingness to recognize, act on and communicate uncertainty when retrieved evidence contradicts the agent's prior beliefs. Three trajectory-level behaviors are scored: Identify the gap, Solve it with bounded tool use, and Escalate residual uncertainty to the user. Across four agent classes, higher task accuracy does not imply greater humility: some accurate agents detect conflicts early but never surface the uncertainty in their final answers. All trajectories and judgments are released.
Outcome-based benchmarks reward agents that are right, even when they are right by luck and silent about doubt. In high-stakes deployment, that silence is the danger. A benchmark that grades whether agents say 'I am not sure' turns humility from a virtue into a measurable engineering requirement.
The paper equips agents with an explicit latent mental model of their counterpart: a learned, amortized recursive Theory-of-Mind representation that infers hidden beliefs, intentions and likely reactions from interaction history. A belief-conditioned reward model then evaluates candidate actions against the inferred partner state, and the resulting signal trains a policy that deploys standalone at inference time. On SOTOPIA with Qwen2.5-7B it gains 11.5% relative over the base model, and on BigToM it improves paired accuracy by 26.7 percentage points.
Current multi-agent systems coordinate through prompts and orchestration, but they do not actually model the minds they are interacting with. Making partner modeling an explicit, reusable latent variable, rather than a pile of prompt text, is how multi-agent interaction graduates from scripted cooperation to genuinely adaptive behavior.
MASS is a recursive self-improvement method for open-ended research tasks that have no verifiable answers. It alternates between evolutionary optimization of multi-agent workflows and fine-tuning the shared model on self-generated trajectories. After two cycles, a 27B model's score per output token rose 1.2 to 1.6 times on four research benchmarks, and a student trained on multi-agent traces beat a single-agent student trained on 1.4 times more tokens.
Recursive self-improvement hits a supervision bottleneck exactly when model outputs exceed what human experts can reliably judge. Showing that a model can become its own better optimizer and evaluator by reorganizing itself as a team, with quantified efficiency gains, is a concrete step toward AI systems that keep improving after training.
A mixed-method survey of 239 practitioners from 31 countries asks how reviewers govern AI-authored pull requests: who bears the review cost when generated code must be reviewed, explained and maintained. Reviewers do not reject AI-authored PRs categorically; they apply a conditional model where review effort follows accountable signals like bounded scope, project-grounded rationale, validation beyond CI and identifiable post-merge ownership. The paper names the gap 'reviewability debt': the work reviewers absorb when generated code lacks human grounding.
As coding agents flood repositories with generated contributions, reviewer attention becomes the scarcest resource in software engineering. Framing the problem as reviewability debt, and asking what makes a generated contribution deserve human attention, gives teams a vocabulary for governing AI-authored code before review queues collapse.
Diffusion models can represent multimodal trajectory distributions, but extracting uncertainty from them usually needs costly Monte Carlo sampling, too slow for real-time control. SCOPE (Score-Curvature for Online Precision Estimation) learns a structured precision matrix around each nominal trajectory by distilling the score model's curvature, producing calibrated Gaussian tubes with per-timestep covariance in a single pass. It was evaluated on pedestrian forecasting, crowd navigation, Maze2D control and a real Franka Panda arm, improving closed-loop performance throughout.
Robots cannot wait for a hundred samples before deciding whether to yield to a pedestrian. Turning a diffusion model's implicit uncertainty into explicit, control-ready tubes in one pass is what moves diffusion policies from impressive demos to machines that can act safely in real time.