A Nature paper takes Stratego superhuman at a fraction of the usual compute (Ataraxos), Cloudflare open-sources decision models that choose without generating text (Clef), and a 385-slide trial shows AI-assisted mitotic counting is far more reproducible (MitPro). Also: SoftBank's $30B OpenAI close and a court pausing Minnesota's AI nudification ban.
SoftBank completed the final $10 billion tranche of its $30 billion OpenAI commitment via Vision Fund 2, bringing its cumulative OpenAI investment to $64.6 billion for a 13% stake. The close follows OpenAI's $122 billion fundraising round valuing the company at $852 billion, anchored by Amazon, Nvidia, and SoftBank.
Synopsys forecast fiscal 2027 revenue of $11.10 to 11.20 billion, above the $10.81 billion consensus, and signed a revenue-sharing deal with OpenAI to build an AI model for chip design. AWS separately signed a multi-year deal worth over $1 billion to license Synopsys chip-design IP.
Microsoft AI released MAI-Transcribe-2-Streaming, its first streaming speech-to-text model covering 60 languages with first hypotheses in just over 100 ms, alongside MAI-Voice-2.1 and a faster Flash variant. Microsoft claims the transcription model ranks No. 1 on Artificial Analysis for final and partial transcript accuracy.
The 8th US Circuit Court of Appeals granted xAI an injunction halting Minnesota's first-in-the-nation ban on AI-generated fake nude images while the company pursues its constitutional challenge. The law, in effect since August 1, bars letting users create realistic intimate imagery of identifiable people; xAI calls it a free-speech violation.
Governor Gavin Newsom signed SB 574, the first US state law governing lawyers' and arbitrators' use of generative AI. Attorneys must personally verify citations, take reasonable steps to check AI outputs, correct hallucinations, and disclose AI use in court filings.
Meta Superintelligence Labs product chief Nat Friedman announced Muse Gadgets: open-source ESP32 firmware and a Linux SDK connecting Meta's Muse personal AI to custom devices, plus a Muse Home Link USB-C adapter for home networks. Meta made 5,000 adapters, free to active US Muse subscribers, shipping within weeks.
Ataraxos combines two interdependent self-play reinforcement learning processes (transformer networks for Stratego set-up and move selection), a belief network that predicts hidden opponent piece types, and a test-time search step built on update equivalence. Its search policy beat world number-one Pim Niemeijer 31-14-5 over 50 games and set new state-of-the-art results in Barrage Stratego, Hanabi, and Dou Dizhu, after training on 16 H100 GPUs for one week.
Imperfect-information games are the closest laboratory we have for real-world strategy under hidden state, from negotiations to multi-agent markets. A 500-fold compute cut with stronger results suggests the bottleneck was algorithmic rather than scale, and the belief-plus-search recipe ports to any multi-agent setting with hidden information.
Clef (27B) and Clef-flash (9B) read an input state plus a schema of typed questions (yes/no, choice, score) and return a probability for every allowed answer in a single forward pass, with no text generation. Cloudflare reports 2.5x median latency speedup over TypeSafe's Jev and top scores on 7 of 10 decision benchmarks; weights are Apache 2.0 on Hugging Face and both run on Workers AI.
Agents spend much of their budget forcing chat models to make structured choices through prose. A model that only decides, fast and calibrated, could become the default gatekeeper inside every agent loop, and open weights mean anyone can self-host the hot path instead of paying per call.
MitPro directs pathologists to regions with the highest predicted mitotic activity and highlights candidate mitotic figures for review. In a paired reader study across 385 whole-slide images, seven tumour types, and three centres, AI assistance lifted the intraclass correlation coefficient from 0.589 to 0.949 and cut median assessment time by about 55%, from 286.4 to 127.8 seconds.
Mitotic counting feeds directly into cancer grading, and its notorious inter-pathologist variability is a patient-safety issue. Evidence that AI assistance more than halves the time while nearly perfecting agreement is the kind of result that moves hospital procurement, not just citations.
NICER reframes whole-slide image condensation as distribution matching under the fixed lens of the pretrained encoder, using a nonparametric prior with slide-adaptive capacity. On five histopathology datasets it beat prior condensation methods by 7.44% average accuracy with a better efficiency-accuracy tradeoff.
A single whole-slide image can span gigabytes, and training on thousands of them is a storage and compute wall. Condensation that preserves the feature distribution the encoder actually sees makes large-scale pathology training cheaper without the usual heuristic hacks.
CPathOGen is a conditional latent-diffusion framework that generates paired H&E counterfactuals with explicit knobs for cell spatial organization, nuclear morphology, and stain appearance. The authors use verified interventions to probe pathology encoders and introduce the Biology-Nuisance Sensitivity Ratio, contrasting sensitivity to real biology against stain-related nuisance.
Pathology models can ace benchmarks by latching onto which lab stained the slide rather than what the tissue shows. Counterfactuals that change one tissue factor at a time give the field a long-missing tool to tell biological signal from batch artifact before deployment.
Under equal time budgets and the same frontier LLM backbone, open-source state-of-the-art harnesses for autonomous ML engineering show no advantage over a single session of a minimal-harness coding agent. Large-scale ablations argue the elaborate machinery layers are redundant and the backbone is the primary performance driver.
The agent field has been adding orchestrators, retrieval subagents, and memory hierarchies on the assumption that scaffolding is the moat. If the backbone does the work, research effort should move from hand-crafted machinery to the models themselves, and a lot of leaderboard engineering needs re-examination.
RSIGame organizes autonomous game development into a local explore-diagnose-improve loop with an evolving checklist plus a global loop that tracks quality and detects saturation, then internalizes successful trajectories into the generator by training. Across 140 GameCraft-Bench tasks, experience internalization let Qwen3.8-27B reach 61.38 on Godot, beating GPT-5.5 one-shot scores with 11x fewer generation tokens.
Recursive self-improvement usually stalls on overfitting to a few test cases; RSIGame shows a concrete loop that avoids that trap and compounds test-time gains into training-time gains. It is one of the clearest demonstrations yet that the two can multiply.
The Independent-Communicate-Revise (ICR) framework audits multi-agent communication as answer revision after independent reasoning, measuring correction rate and preservation rate with a no-message control. Across four reasoning benchmarks, richer messages amplify both beneficial correction and harmful misdirection, challenging the idea that communication quality is an intrinsic property of a channel.
Multi-agent papers routinely credit communication for accuracy gains without showing the messages did the work. ICR separates correction from corruption and gives the field a shared yardstick, which matters as agents move from text to latent channels nobody can read.
A counterfactual audit of ten open-source LLMs on pediatric Emergency Severity Index prediction, varying one demographic, socioeconomic, or system-context factor at a time. Sensitivity did not reliably fall with larger models or medical pretraining; a fine-tuned Qwen2.5-7B had the lowest shift rate at 5.27%.
Triage decides who gets seen first in the emergency department, and open-source models are attractive precisely because hospitals can run them locally. A lightweight audit showing that bigger and medical-tuned do not mean fairer is a practical pre-deployment check the whole field can copy.
Revisiting fairness evaluation for ICU mortality prediction on MIMIC-IV, the authors show interventions can look very different across accuracy, AUROC, sensitivity, and false-positive rate, and that marginal demographic summaries hide heterogeneous errors inside intersections of ethnicity, gender, and insurance.
Fairness reports usually pick one metric and one demographic split, which lets uncomfortable disparities hide in the averaging. This paper makes the case that evaluation resolution is itself a fairness decision, a point regulators and hospital AI committees should take seriously.
An adapter connects a frozen diffusion model to a pretrained vision-language space, steering batch composition toward target attribute proportions for fairness while a disagreement-based score preserves diversity, all without sensitive-attribute annotations. It works for unconditional and text-conditional diffusion models alike.
Diffusion models amplify the demographic imbalances in their training data, and existing fixes either need attribute labels or kill diversity. A supervision-free method that keeps both fairness and variety lowers the bar for shipping fairer image generators.
The Energy-based Feynman-Kac Corrector adapts pretrained diffusion models to new sampling tasks without retraining, correcting the pretrained model's own path-tracking error on the fly given a reference energy. On Gaussian mixtures, particle systems, and alanine peptides, it matched target distributions and free-energy profiles where standard inference-time scaling kept substantial errors.
Inference-time scaling assumes the pretrained model is exact, which it never is, especially out of distribution. Correcting the model's own error instead of just adding particles makes pretrained diffusion models genuinely reusable for scientific sampling tasks like molecular simulation.
CoFlow tackles flow matching's trajectory-crossing bottleneck head-on: crossing points create huge local Lipschitz constants that networks struggle to fit and that wreck few-step integration. A contrastive repulsive drift pushes trajectories apart during training, cutting FID on ImageNet 256 at 20 steps with no added training cost.
Few-step generation is what makes flow models deployable, and crossing trajectories are the quiet reason they still need dozens of steps. Fixing the velocity field itself rather than distilling around it is the more principled fix, and it comes for free.
CEASE (Continual Erasure via Adaptive Subspace Editing) removes new concepts from text-to-image diffusion models over time without undoing prior erasures, using a training-free closed-form solver with two subspace constraints. Across continual erasure of celebrities, artistic styles, and instances, it kept the most consistent erase-preserve tradeoff.
Content governance is not a one-shot task: takedown requests arrive continuously, and each edit risks breaking the model a little more. A method that keeps erasures stable over long edit sequences is infrastructure for the legal reality of deployed image models.