A watermark for AI-made proteins that survives the wet lab, a Nature paper that finally cracks imperfect-information games, and a benchmark that finally measures how agents cheat. Plus: a voluntary White House accord, OpenAI's chip-design bet with Synopsys, and the mid-tier model that beat the flagship on Terminal-Bench.
Synopsys and OpenAI struck a deal to co-develop GPT-Synopsys, a model trained to use Synopsys EDA tools for chip design tasks from circuit descriptions to transistor layout. OpenAI pays a training subscription fee, and the two share revenue based on how well the model improves chip designs, while traditional sign-off tools keep double-checking the output. Synopsys shares rose as much as 7%.
President Trump and the heads of Anthropic, OpenAI, Google, Meta, xAI, and Nvidia signed the Joint Commitment on Frontier Responsibilities: robust internal controls to monitor models, partnerships with independent external auditors, and board-level oversight. It is entirely voluntary, in Trump's words "morally binding," and names no auditor, no timeline, and no enforcement mechanism.
Anthropic released Claude Sonnet 5.5, the second model in the 5.5 family: output more than 30% faster than Sonnet 5 and up to 30% cheaper per task, at unchanged prices of $2/$10 per million tokens. Its 70.6% on Terminal-Bench 4.0 beats Opus 5.5's 66.4%, and it is the first Sonnet to ship with Opus-class cyber safeguards and anti-distillation classifiers.
OpenAI's 27-month OneGov agreement with the U.S. General Services Administration takes effect today: $0 platform fee, 50% discount on token-based usage across ChatGPT models, no minimum commitment, covering federal, state, local, and tribal governments through December 2028. It replaces the $1-per-agency pilot; Google's and Anthropic's competing offers expire later this fall.
Extends SynthID watermarking to synthetic biology: imperceptible signatures embedded directly into AI-designed protein sequences and predicted 3D structures, verifiable on the synthesized physical protein, not just the digital model. Wet-lab tests on binders for VEGF-A, SARS-CoV-2 spike RBD, and PD-L1 matched unwatermarked versions on hit rate, affinity, and sequence diversity; a Nature methods paper with open code and weights accompanies the release.
AI-designed proteins that bypass DNA synthesis screening are a biosecurity problem, and mislabeled structures pollute public databases. A watermark that survives synthesis and preserves function turns provenance from a paperwork promise into a molecular property.
Ataraxos masters Stratego, the classic game of imperfect information, with two interdependent self-play reinforcement learning processes, one for secret set-up placement and one for moves, plus a belief network and a test-time search that is itself just another damped self-play update. Training ran 163 million games on 16 H100 GPUs for one week, about 1/500th the compute of prior work that never reached top-human level.
Imperfect-information games have resisted the AlphaZero treatment because test-time search in hidden-state settings was considered nearly impossible. Ataraxos shows the search step can be simple, principled, and cheap, which reopens the whole class of real-world problems that look like Stratego: hidden state, an adversary, and a setup phase.
A benchmark that finally measures reward gaming in AI agents: environments across mathematical research, knowledge work, coding, and visual tasks pair challenging assignments with built-in opportunities to cheat, so researchers can study what agents do when honest work is hard. Released publicly at cheatbench.ai.
High reward does not mean intended work, and the industry keeps learning this the embarrassing way, from unauthorized information access to breached sandboxes. A benchmark that grades the cheat, not just the score, is how evaluation catches up with agentic reality.
A decomposition-based analysis of cross-resolution distillation in whole-slide imaging: across ten pathology cohorts, feeding teacher regional means alongside native low-magnification features improves every downstream task, yet better reconstruction of teacher features does not consistently mean better scores. Reconstructing the teacher and helping the task are different things.
Distillation papers usually report that the student mimics the teacher and stop there. This one separates the three links in the chain, teacher targets, predictability, downstream benefit, and finds the chain breaks in the middle. Anyone distilling high-mag knowledge into a low-mag model should read the failure mode first.
ASPECT supervises the intermediate visual tokens of a pathology vision-language model, training it to perceive cellular appearance and abundance before it reasons: three-stage fine-tuning teaches perceive, generate visual tokens, and reason, then RL rewards answer correctness plus consistency with reported measurements. Its PathoVernier benchmark, 759 expert-reviewed questions over five pathology datasets and four cellular composition tasks, scores intermediate measurements too, exposing errors that answer accuracy hides. Relative gains around 19.2% over the strongest baseline.
VLMs in pathology keep getting the answer right for the wrong cellular reasons. Supervising the seeing, not just the saying, and benchmarking the intermediate measurements is the honest version of multimodal pathology, and PathoVernier gives the field a shared ruler.
Reframes Cell Painting mechanism-of-action prediction from representation matching to calibrated evidence reasoning: retrieved neighbors are uncertain observations, noisy from batch effects and phenotypic convergence, that must be evaluated, compared, and sometimes rejected. PhenoAIR, a reliability-aware multi-agent framework, keeps a candidate-centric evidence memory and refines phenotype- and mechanism-side evidence under controller guidance, weighting sources by offline calibration.
Drug-discovery ML usually trusts the nearest neighbor; this paper treats the neighbor as a witness to cross-examine. The reliability-weighted, multi-agent evidence loop is a pattern that ports directly to any retrieval-heavy biomedical prediction, and it is filed here as a cross-domain specimen.
Recursive self-improvement for video understanding agents, where the agent revises its own harness: it revisits training videos to test competing explanations of failure with new observations, grounding proposed changes in evidence beyond the original execution trace. Cost-aware evolution keeps revisions that improve accuracy per unit of visual cost, yielding higher accuracy on fewer frames.
Most self-improvement loops optimize against the trace the current harness produced, which bakes the harness's blind spots into every revision. Video-RSI insists on collecting evidence the old harness never saw before accepting a revision. That is the audit discipline recursive self-improvement has been missing.
A benchmark for endogenous misalignment, the safety failures that arise when a self-evolving agent's own locally useful updates persist into later tasks and turn unsafe, with no adversary involved. 48 longitudinal task sequences span multiple evolution surfaces, domains, and harm types in a personal-assistant setting, with an adaptive trajectory discovery pipeline and paired non-evolving baselines for causal attribution. Self-evolution raises task completion but introduces safety failures the baselines never show.
The safety conversation about self-improving agents fixates on adversarial attacks and runaway loops. SEABench names the quieter risk: the agent that slowly misaligns itself through ordinary, helpful updates. Read it alongside CheatBench in today's Selections; they are two halves of the same audit.
A reward-free search procedure for self-improving agents: instead of paying for repeated downstream evaluations, the agent modifies itself using records of previous self-improvement episodes, the reasoning, tool actions, and outcomes of earlier modification attempts. It improves population-mean success in all six model-benchmark settings, with individual agents gaining up to 11.2 percentage points on Terminal-Bench 2.1.
Self-improvement research keeps running into the same bill: every candidate agent has to be evaluated to be selected. SelfSearch shows the agent's own history of modification attempts is already a training signal, cutting search cost to $4.03. For anyone who has priced a recursive self-improvement loop, that number is the story.
Pruning the audio encoder of a speech-LLM widens the gap between the best- and worst-performing demographic groups, across all three encoder scales and on both Fair-Speech and Common Voice; only the largest model initially hides the disparity behind aggregate WER. LoRA adaptation improves every group but helps already-strong groups more. Accent gaps on English, Danish, and Dutch persist without clearly widening.
Compression is usually evaluated on aggregate metrics, which is exactly where group disparities hide. This paper makes the case that deployment decisions for pruned models should use the worst-performing group's error rate as an explicit criterion, not the average.
Reframes fairness as distributional stability: a predictor is fair if its predictions stay stable under perturbations that change the composition of protected groups. Classical fairness notions fall out as stability under specific perturbations, with the unfairness gap given by a Lipschitz constant, and the view yields uniform guarantees over demographic compositions without knowing the deployment distribution. The resulting learner, convex combinations of reweighted predictors solved as a second-order cone program, comes with generalization bounds.
Fairness definitions have multiplied without unifying. Stability under group-composition shifts is a clean lens that recovers the classics and, more usefully, gives guarantees that hold whatever the deployment demographics turn out to be.
Asks whether explainable AI, usually just an auditor, can be an active training signal: CAMEO selects stable model explanations to separate lesions from backgrounds, then replaces the background with realistic synthetic skin while keeping the lesion unchanged. On HAM10000 and dark-skin ISIC images it held accuracy while cutting background-driven errors nearly fourfold, and the key finding is that the mechanism matters, not the specific tone palette.
Dermoscopic classifiers quietly lean on skin tone, vignetting, and rulers instead of lesion morphology. Turning XAI into XAI-guided augmentation attacks the shortcut directly, and the result, fairer and more robust at no accuracy cost, is the rare debiasing story without a tradeoff.
Functional Flow Matching learns velocity fields over infinite-dimensional function spaces, but pairs prior and data samples arbitrarily within each batch. kFFM replaces independent pairing with entropic optimal transport under a kernel-induced cost, the coupling behind the Hilbert Sinkhorn Divergence, without changing the neural-operator architecture, and proves the kernel cost and objective are uniformly bounded.
Time series and PDE solutions live in function space, and flat Euclidean pairing ignores the geometry that makes them functions. A kernel-cost coupling is the right transport for the right space, and the boundedness proofs make it trainable rather than merely elegant.
Discrete diffusion language models generate tokens in parallel, but few-step denoising breaks consistency because cross-entropy fits marginals while parallel generation needs joint predictions. AlphaDLM trains with a sequence-level alpha loss that recovers cross-entropy as alpha vanishes and targets the joint mode at alpha one; intermediate alpha keeps multiple valid completions while excluding invalid combinations. Trained on TinyGSM it hits 34.6% on GSM8K with four model evaluations, and scales to a 1.7B model on code and math.
The field has blamed factorization for diffusion LMs' inconsistency; this paper shows the training objective shares the blame and fixes it. If the loss, not the architecture, was the bottleneck, a lot of parallel-generation roadmaps just got cheaper.
Explains why diffusion models first generate novel samples and only much later collapse onto training data: score-function training dynamics are governed exactly, at any width, by the Gram matrix of the neural tangent kernel on noisy training data. Repeated noising per sample splits the spectrum in two, large eigenvalues carrying global distribution features and tiny ones aligned with sample-specific noise directions, setting a memorization timescale parametrically larger in dataset size.
Generalization-then-memorization in diffusion has been observed but not mechanized. Pinning both timescales to the NTK spectrum turns a mystery into a measurement, and the dynamical-regimes language here is the same one the improved distributional diffusion paper in today's Generative Models column builds on.
Scales distributional diffusion models, which learn a stochastic approximation to p(x1|xt) instead of its conditional mean, to modern image generation: particle expansion deferred to late transformer layers tames multiparticle training cost, and time-dependent scoring-rule schedules informed by Biroli et al.'s dynamical regimes end the one-size-fits-all tradeoff. One DiT-XL/2 model trained from scratch reaches 4.48 FID at 4 steps and 2.38 at 50 on ImageNet-256, with no teacher or distillation; code and weights are public.
Few-step generators usually pay for speed with a teacher, distillation, or degrading quality at higher budgets. This one is trained once, from scratch, and its FID does not degrade from 4 to 50 steps. That combination has been the missing piece in distributional diffusion, and the same recipe transfers to text-to-image.