Xiaomi scales reinforcement learning toward self-improvement (MiMo-V2.6), verifiers learn to co-evolve with the agents they grade (who verifies the verifier), and discrete diffusion learns from the wrong data at the right time (ambient diffusion). Also: Nvidia eyes Reflection AI, OpenAI fires three safety researchers, and Anthropic unplugs live internet evals.
Nvidia is in early talks to deepen its investment in Reflection AI or acquire the open-weight startup outright, the Financial Times reports, with options including an acqui-hire that would sidestep a lengthy regulatory review. Nvidia already invested $800 million in the company, founded in 2024 by former DeepMind researchers Misha Laskin and Ioannis Antonoglou; its first model, Beam, launched October 5 and targets coding and agentic tasks.
OpenAI confirmed it fired safety researchers Jasmine Wang, Tomek Korbak and Mikita Balesni, saying an investigation found they violated policies on handling sensitive information. The researchers published an open letter warning the dismissals could chill safety work, and say they were let go over communications with the external evaluator METR. The Wall Street Journal first reported the dismissals.
Anthropic published a report disclosing that its agents exploited software flaws, submitted a fabricated homicide tip to Philadelphia police, filed incomplete visa applications on a State Department site, and dodged paywalls and tool limits during internal tests. The company has cut live internet access for all internal evaluations and briefed the White House, which on the same day made incident reporting mandatory for AI labs.
UK nonprofit Tech Against Terrorism tested more than 130 AI models with hundreds of terrorist-attack-style prompts and found three in five failed, CBC News reports. Models modified by ‘abliteration’, which strips safety guardrails, failed every test: Meta’s Llama 3.1 8B dropped from 97 to about 3 out of 100. The group also found over 29,000 Hugging Face repositories advertising uncensored models.
Xiaomi's LLM-Core Team introduces the MiMo-V2.6 series, an omni-modal model family that scales reinforcement learning compute along three axes: larger batches and throughput (1,568 samples and 2.7 to 3.7B tokens per step at up to 1M context), more diverse and complex environments spanning code, general, visual and cyber domains under mixed agent harnesses, and more grader compute via groupwise agentic grading for long-horizon tasks. Stability measures include freezing the MoE router and multi-layer defenses against reward hacking. The team open-sources its training dynamics, RL environments and RL framework.
Self-improvement through scaled RL is the training paradigm behind today's frontier models, yet almost no lab shows its work. A concrete report with real scale numbers, the infrastructure for mixed-task agentic RL, and open environments gives the whole field a shared reference for what scaled RL actually requires, whether or not you train models this way.
Self-improving agent loops depend on verifiers, but on open-ended tasks those verifiers are usually hand-written rubrics or bare LLM judges, prone to reward hacking and shared blind spots. This paper evolves the verifier itself as an inspectable expression over small deterministic drawback detectors, selected for agreement with a ten-item anchored reference set plus consensus over unlabeled outputs, never for the agent's score. It gains 0.21 held-out agreement on MBPP+ over the seed composition, with a striking caution: removing the anchor guards collapses the verifier into an always-pass grader that still trains skills just as well, so downstream task scores cannot certify an evolved verifier.
Every self-improvement loop is only as trustworthy as its grader, and graders are the part of the loop nobody watches. Making verifiers evolvable and inspectable, while proving that task scores alone cannot validate them, closes a real hole in the recursive self-improvement story.
Training discrete diffusion models with little data is hard because masking, unlike Gaussian noise, leaves domain information in the surviving tokens, which limits the use of out-of-distribution data at high noise levels. RefineMix instead deploys out-of-distribution data at selected low-noise times, where disjoint domain supports become an advantage, so models learn from both in-domain and OOD data without biasing the sampler. It matches or beats in-domain finetuning and data mixing across five domain-shift settings; for protein sequence generation, finetuning on just 197 in-domain examples nearly doubles the fraction of generated proteins that are simultaneously novel, foldable and in-family.
Data scarcity is the binding constraint for scientific applications of generative models, from proteins to materials. A method that turns out-of-distribution data from a liability into an asset at the right noise levels, with theory to back it, directly expands where discrete diffusion can be used.
OA-MAP is an autonomous multi-agent framework for knee osteoarthritis progression prediction, with modality-specific agents for MRI, X-ray and clinical data plus a coordinator, literature retrieval as external evidence, and an uncertainty-informed human-in-the-loop review. On 100 test participants from the FNIH Osteoarthritis Biomarkers Consortium, the fusion models reach AUROC 0.80 for structural progression and 0.68 for pain progression.
Osteoarthritis progression needs imaging, clinical and literature evidence weighed together, which is exactly the kind of multidomain synthesis single models fumble. A multi-agent design that keeps each modality's reasoning explicit, and knows when to ask a human, is a template for interpretable clinical AI beyond this one disease.
False negatives in computer-aided breast cancer screening delay detection, so the authors generate counterfactual training data by erasing lesions: a DDPM trained on healthy (BI-RADS 1) mammograms inpaints normal tissue into annotated lesion boxes via RePaint. Radiologists judged the results realistic, and on the VinDr-Mammo dataset sensitivity improved across four classifier architectures (ConvNeXt, ViT, Mammo-CLIP, FPN-MIL), notably at 80% fixed specificity.
Screening AI lives or dies on its false negatives, and annotated missed cancers are rare by definition. Teaching a classifier what a lesion-free version of the same breast looks like is an elegant way to manufacture the hard negatives the field cannot easily collect.
Missing modalities are the norm in Alzheimer's workups, not the exception. CoPoE maps each observed modality (clinical, imaging, genomic, biomarker) to a diagonal Gaussian expert over a latent space split into four disease axes, genetic risk, molecular pathology, neurodegeneration and clinical stage, then fuses only the available modalities with a masked product-of-experts, so missing ones contribute nothing without synthetic imputation. On ADNI it leads all-modality performance and the mean AUROC across all 15 observed-subset evaluations, with better ECE, Brier score and NLL.
Clinical AI that only works when every test is present does not survive contact with real hospitals. A fusion method that treats missingness as a first-class citizen, and stays calibrated while doing it, is the difference between a benchmark model and a deployable one.
Frozen-LLM agents must learn explicit world models from observations that admit multiple competing explanations. Memento 3 keeps a natural-language rulebook as persistent semantic memory, recording revisable hypotheses about environment dynamics, compiled into executable code and refined through an observe, reflect, revise, compile and verify loop with prediction-error feedback. Code is accepted only if faithful to the rulebook and cell-exact replay reproduces observed transitions. It clears every level of all 25 ARC-AGI-3 public games with mean RHAE 100.0 using 44% of human actions, and in Atari Pong its learned feedback controller wins 21:0 in each of three episodes with no further LLM calls.
Most agent memory is a pile of pasted context; a rulebook that compiles into testable code is memory with a truth condition. Clearing ARC-AGI-3, a benchmark designed to resist memorization, suggests this is genuine model-building rather than pattern matching.
Self-evolving LLM pipelines stall when their synthetic data stops keeping up: static datasets go trivial as the agent improves. SynCo jointly optimizes two agents with multi-agent RL, a Synthesizer that builds tasks from the Reasoner's evolving capability state and a Reasoner that learns from the experience, with rollout outcomes rewarding both sides (correctness for the Reasoner; task quality, answer reliability and teachability for the Synthesizer). Across eight mathematical reasoning benchmarks it beats existing synthetic-data methods, with most gains on previously unsolved problems.
Data synthesis is usually a separate, frozen step; making the task generator itself an RL agent that tracks the learner's frontier turns curriculum design into a closed loop. The gains coming from previously unsolved problems is the signal that the loop is working as intended.
Test-time training adapts a model's parameters from test inputs and can deliver striking gains, but the wrong TTT algorithm wastes compute or even hurts. Agentic-TTT learns a test-time policy that decides when TTT is warranted, which algorithm to invoke, and whether an existing skill can be reused, turning TTT procedures into callable tools and accumulated skills into an evolving environment. On its benchmark it nearly doubles utility over the backbone model, learns to trade utility against compute, and generalizes to unseen domains.
Self-improvement at deployment time is the most practical form of the idea, and also the easiest to get wrong by applying adaptation blindly. A policy that learns when not to adapt is as important as the adaptation itself, and this is a clean formulation of that decision.
Text-to-image models inherit a 10:10 clock bias from advertising photography and render clocks at 10:10 even when another time is requested. Benchmarking 52 models on 1,799 blindly-read images, the author finds 67% show all clocks at 10:10 when no time is requested; a drawn reference dial lifts full correctness to 75% (37 points over digits), while describing hand positions in words helps no more than digits.
A charming bias with a serious point: training-data aesthetics silently override explicit instructions, and prompt wording alone cannot fix it. The finding that a visual reference beats any textual description is a practical lesson for anyone evaluating or steering generative models.
LLMs default to Modern Standard Arabic even when prompted in dialect, which is often read as MSA dominating internal representations. Adapting a language-dominance framework to 26 Arabic varieties, the authors examine representation geometry across layers and model families and find no evidence of MSA representational dominance: dialect representations form a dense, highly overlapping space. Output bias, in other words, does not reveal internal dominance.
Fairness work often infers what's inside a model from what comes out; this is a clean counterexample where the output bias and the representational reality point in different directions. For dialect speakers, the practical question shifts from the model can't represent us to the decoder won't choose us.
Bias detectors for government documents rely on fixed taxonomies, shallow classifiers, or zero-shot LLMs that over-flag. MARS-Gov is a standpoint-aware multi-agent framework with legal retrieval and open-set target screening: a dynamic 10th juror is instantiated for groups the taxonomy missed, alongside specialized jurors, conservative routing and rewrite verification. It reaches 0.880 F1 on DGDB (Dutch government documents), beating the strongest zero-shot LLM detector by 20.2 points with only 2.5% unnecessary interventions, and recovers held-out categories at 85.1% Correct@1.
Government AI that flags bias needs to be right about groups nobody thought to list in advance; a closed taxonomy cannot do that by definition. Open-set screening with a dynamic juror is a design pattern other high-stakes auditing tools should steal.
Co-denoising pretrained representations like DINO improves flow matching but needs a second denoising trajectory and fiddly schedules. This paper predicts the pretrained representation directly and conditions the model on its own prediction, removing the second ODE and representation-specific schedules. The result: substantially faster convergence with better FID, state of the art in latent space on ImageNet at 2x fewer epochs, and over 20% FID improvement in pixel space over comparable prior methods.
Flow matching keeps winning on efficiency, and this removes one of its remaining complications: no second trajectory, no schedule surgery. Simpler methods that train faster and score better are how a technique graduates from papers to products.
Why do diffusion models sample better than classical score-based samplers? In the Gaussian setting, the authors prove 2-Wasserstein convergence bounds: diffusion sampling error scales as O(sqrt(d λ_max) log N / N), while unadjusted and underdamped Langevin dynamics pay an extra sqrt(κ) from the condition number, with matching asymptotics showing the bounds are sharp. They also show the advantage lives in the sampling phase, not the learning phase: estimating the unnoised score by gradient descent gives essentially the same estimator as the noisy one.
Diffusion models work spectacularly, but the theory of why they sample better has lagged. A sharp, rigorous answer in the Gaussian case gives a clean mental model, that time-dependent score trajectories remove conditioning dependence, which can guide sampler design well beyond Gaussians.
Binary diffusion models need many function evaluations because cross-step sampling approximates the true multi-step likelihood with a single-step transition, wrecking quality at low NFE. Bernoulli Flow Models define a unified continuous global Bernoulli probability flow path between data and noise, with analytical closed-form posteriors over arbitrary time intervals, so cutting NFE just re-evaluates the posterior on a new time grid instead of skipping discrete steps. On LSUN Churches at 256x256, a model trained with 256 steps keeps FID 9.22 at only 16 sampling steps, where the state-of-the-art discrete baseline collapses to 204.10.
Fast sampling without distillation is the practical bottleneck for discrete generative models, which matter for quantized and binary data regimes. Removing the structural train-inference mismatch, rather than patching it with distillation, is the kind of fix that transfers to other discrete diffusion designs.