Game theory gives multi-agent systems a way to police LLM hallucination (G-Frame), a radioactive watermark follows stolen protein models back to their source (MARCO), and agents learn open-ended medical reasoning by accumulating knowledge (MedZERO). Also: Biohub's $1.8B bet on virtual cells, a consortium to authenticate your personal agent, and Claude Haiku 5.5's price war.
Zuckerberg-backed Biohub said its Virtual Biology Initiative has grown to $1.8 billion, with Meta, Google DeepMind and Isomorphic Labs jointly investing $300 million and the US Department of Energy committing more than $500 million over five years. The money funds open datasets for training predictive 'virtual cell' models of biology; the datasets will be released publicly, though commercial funders get an embargo head start.
Meta and Sierra, the startup led by Bret Taylor, announced an open Personal Agent Protocol for authenticating personal AI agents with businesses, with founding partners including Walmart, Stripe, Shopify and Genesys. The move follows Amazon's blocking of Meta's Muse shopping agent; a v0.1 specification is due later in October.
SAP agreed to acquire TechWolf, whose AI platform maps enterprise work and skills, to anchor work intelligence in its SuccessFactors portfolio. The deal is expected to close in Q4 2026 with terms undisclosed; TechWolf will stay independent under CEO Andreas De Neve in Ghent.
OpenAI began rolling out GPT-6 in ChatGPT's Chat tab on October 7 with 'Intelligent UI', letting answers render as interactive charts, forms, calculators and games instead of prose alone. Paid tiers run on GPT-6 Sol starting October 7, with Free and Go tiers on GPT-6 Luna from October 8.
Anthropic launched Claude Haiku 5.5, its cheapest and fastest small model, priced at $0.10/$0.50 per million input/output tokens for prompts under 100K, roughly 75% cheaper than Haiku 4.5. It is the third Claude 5.5 model in a month, arriving as the company reportedly prepares for a planned IPO.
Anthropic merged Project Glasswing into an expanded Cyber Verification Program with three access tiers (Defense, Red Team, Specialized), giving vetted security professionals reduced-safeguard access to Claude Opus 5.5, Sonnet 5.5 and Mythos 5.1. Glasswing partners uncovered 129,000 verified vulnerabilities between April and July 2026, and the new program formalizes that pipeline.
Researchers at Dalian University of Technology and HKUST introduce G-Frame, an adaptive multi-agent framework inspired by Bayesian and team game theory that generates, evaluates and trains on its own data: a specialized corpus of 363,045 chain-of-thought examples and 199,589 question-answer pairs. Domain pre-training on this corpus lifts their 7-billion-parameter OmniChem model to competitive benchmark performance and cuts its ChemJudge hallucination rate from 41.95% to 12.91%, with case studies in molecular design and synthesis planning.
Hallucination is the tax every serious LLM application pays, and chemistry is a domain where a confident wrong answer can be dangerous. A framework that manufactures its own training and evaluation data, then uses game theory to keep agents honest, points toward a future where specialized models are grown rather than hand-built.
The first radioactive watermarking framework for protein generative models: watermarks are embedded during the diffusion reverse-denoising process while the protein model stays frozen, optimized to preserve biophysical fidelity such as C-alpha distances and torsion angles. Crucially, the watermark transfers to any pirate model trained on watermarked outputs, countering model extraction and enabling forensic tracing of biosecurity misuse.
Protein design models are dual-use technology: the same model that designs a cure can design a threat, and stolen models leave no trace. A watermark that propagates into copies gives labs and regulators a way to prove provenance after theft, which is the missing piece for responsible release of powerful biological models.
A self-evolving agent framework for open-ended medical reasoning: an Examiner generates frontier medical question-option pairs while a Reasoner answers with evidence-grounded, multi-turn reasoning and external tools. Controlled knowledge accumulation separates exploratory from persistent knowledge, and the system beats prior self-evolving baselines by up to 13.7 accuracy points across five medical reasoning benchmarks.
Medical knowledge never stops changing, but medical models are frozen at training time; an agent that safely accumulates new knowledge from its own practice narrows that gap. The controlled separation matters: it is what keeps a self-updating medical system from drifting into confident nonsense.
A sobering audit of how we evaluate explanations in pathology AI: in hierarchical compact-evidence explanations of whole-slide multiple-instance learning models, candidate filtering silently changes the prediction being explained, which can reverse comparative conclusions about explanation fidelity. CHARTER is an audit protocol (declare the target and reference, quantify prediction shift, check conclusion stability) demonstrated on cases where rankings actually flip.
Explainability is what stands between a pathology model and clinical use, and this paper shows a common evaluation setup can certify the wrong explanation as the better one. An audit protocol that catches reference substitution is infrastructure the whole field needs before these explanations reach patients.
A survival prediction framework fusing whole-slide images with transcriptomic profiles through learnable semantic anchors that enforce regularized cross-modal alignment, plus a hierarchical mixture-of-experts that mimics the pathologist workflow: filtering salient regions within each magnification and routing across resolution levels.
Prognosis is where pathology meets the question patients actually ask: what happens to me. Aligning what the slide shows with what the genes say, in a model that reasons the way pathologists zoom through magnifications, is how multimodal AI earns a place in treatment decisions.
A multimodal knowledge distillation framework trains a teacher on fused whole-slide images plus clinical text using low-rank multimodal fusion, then distills it into a student that needs only the image at inference. On the PatchGastric benchmark it gains at least 3.35% mean accuracy over the state of the art, without transformers or LLMs.
Clinics cannot always run giant models or fetch paired clinical text at diagnosis time; a lightweight image-only model that learned from richer data during training fits how pathology actually deploys. Beating the state of the art without transformers is also a useful reminder that architecture fashion is not destiny.
A benchmark for the continual self-evolution of AI-for-science agents spanning 23 disciplines: the framework repairs executable workflows through multi-turn interaction, converts verified failure-to-success trajectories into Skill and Operator candidates, and keeps updates only when replay and transfer both improve.
Science agents that cannot learn from their own failed experiments will keep repeating them; a benchmark that measures whether evolution actually transfers, rather than just whether it happens, is how the field separates real self-improvement from self-congratulation.
A training scheme that co-evolves a task curriculum, an injection adversary and the web agent itself inside a frozen web world model. Training lifts a 4-billion-parameter agent's task completion with and without attacks, transfers to a real browser, and improves completion under an unseen frontier-model adversary by 33.6% relative across 150 web tasks.
Web agents live in hostile territory: every page they read might be trying to hijack them. Training the attacker alongside the agent, rather than bolting defenses on afterward, is the more honest way to build agents that survive the open web.
Hindsight Meta-Experience Distillation builds reusable 'Meta-Experience' for self-improving agents by re-executing incumbent and revised meta-skills from the same restored discovery state, so what a revision changed is observed under shared conditions. Even rejected revisions contribute learning signal, and the method improves results across three interactive agent benchmarks.
Self-improving agents usually learn only from what worked; learning from the revisions that did not is a deeper trick, and it mirrors how human experts improve by studying their own near misses. Making rejected attempts useful doubles the value of every experiment.
A single-encoder framework combining dual-level decorrelation with prototype-guided cross-group contrastive learning to disentangle demographic signals from disease representations in chest X-ray diagnosis. On 34,809 CheXpert test images across eight intersectional groups, it cuts the mean equalized-odds gap from 15.41% to 10.86%, and introduces DRAR, a new metric for demographic-structure reduction in disease representations.
Chest X-rays are among the most common medical images in the world, and a model that works better for some demographics than others quietly rations care. Disentangling who the patient is from what the disease looks like is fairness at the representation level, where it is hardest to fake.
A careful look at activation steering for debiasing finds the linear 'debiasing direction' is dominated by model confidence rather than bias: much of the apparent bias reduction is a side effect of the model abstaining more. The authors argue bias directions disentangled from confidence are hard to isolate, and steering-based debiasing results need careful interpretation.
Debiasing claims are easy to make and hard to check; a result that looks like less bias can just be a model that says less. This paper gives the fairness community a concrete confound to control for, which makes every future steering claim more honest.
A fairness recipe for speech recognition: fine-tune only the connector of a speech-LLM ASR model on demographic subsets, merge the subgroup-adapted connectors, then apply intersection-specific correction vectors for conflicting cross-axis demographic pairs. On Fair-Speech, TIES merging with WER-based correction cuts word error rate from 7.38% to 5.13% while improving disparities, though lower average error does not always mean lower subgroup disparity.
Voice assistants that misunderstand some accents more than others quietly exclude their speakers; fixing it at the model-merging level, with explicit attention to intersectional groups, is fairness engineering rather than fairness theater.
H-CDLM diffuses tokens and coarser embedding-cluster modalities in parallel, each with its own sampler and schedule. Applied to CoBit it cuts generative perplexity by 24.2 and 20.7 points on LM1B and OWT respectively and reaches 27.4% on GSM8K; the same framework also improves the flow matching model FLM, showing the idea generalizes across continuous generative paradigms.
Diffusion language models are the most credible challenger to autoregressive text generation, promising speed and controllability; a hierarchical recipe that works for both diffusion and flow matching suggests the field is converging on shared machinery rather than competing religions.
DireSMC guides a population of weighted samples toward rare events in diffusion models via sequential Monte Carlo, producing both the samples and a calibrated rare-event probability estimate. On a score-based climate emulator it runs 9x to 1413x faster than plain Monte Carlo for event rarities from 10^-3 to 10^-5.
The events that matter most, extreme weather and system failures, are the ones models see least; a principled way to steer generative models toward the tails, with honest probability estimates attached, turns diffusion models into tools for risk rather than just for pretty pictures.
A feed-forward conditional flow matching model that transports vision-foundation-model hand-object state estimates toward the interaction manifold, correcting translation, rotation and alignment, with physical interaction constraints and 2D evidence steering the generative transport itself. It reaches state of the art on out-of-domain 4D hand-object benchmarks.
Robots and AR glasses both need to understand hands touching things, and that understanding has to work outside the lab's training distribution. Flow matching that respects physical contact constraints is a step toward interaction models the real world cannot easily break.