The SynthID protein watermark, now as a peer-reviewed Nature paper, Google's new frontier model Gemini 4 Argon, and the self-distillation loop that collapses when nobody is watching. Plus: Broadcom's $42B lifeline to Anthropic and the FTC's first probe into rogue AI agents.
Anthropic's IPO prospectus discloses that Broadcom agreed to lend it up to $42 billion in convertible notes, covering about a third of Anthropic's $125.2 billion five-year compute-lease commitment. Anthropic is set to become Broadcom's largest compute customer in 2027.
The FTC is preparing civil investigative demands and plans to compel executive testimony from OpenAI, Anthropic, and METR over consumer risks from autonomous AI agents, spurred by the Hugging Face hacking incident. It is the first formal U.S. enforcement action focused on agents acting beyond human control.
Micron reported record fiscal Q4 revenue of $54.23 billion, nearly five times the prior year and above the $51.07 billion consensus, with its data-center division revenue up 11x on HBM and AI demand. It guided next-quarter revenue of $61.5 billion.
Key provisions of Connecticut's Public Act 26-15 entered into force on October 1, including whistleblower protections for frontier-model developers and a rule that using automated employment-decision tools 'shall not be a defense' against discrimination complaints, plus design and risk-management duties for models trained with over 10^26 FLOPs.
At its Inscribe conference, legal AI startup Ivo released Ivo Sage, an open-source model post-trained for long-horizon contract work (built with River AI on DeepSeek V4 Flash; reinforcement learning lifted it from 70% to 91% on Legal Agent Benchmark contract criteria), plus a preview of its Contract Bench benchmark.
The peer-reviewed record of DeepMind's SynthID Bio: imperceptible watermarks embedded in AI-designed protein sequences, via a tournament-sampling adaptation of SynthID-text inside ProteinMPNN, and in predicted 3D structures, via fine-tuning of AlphaFold 3's diffusion module. Wet-lab tests on binders for the SARS-CoV-2 spike RBD, VEGF-A, and PD-L1 showed watermarked designs matching unwatermarked ones on binding affinity, hit rate, and sequence diversity.
AI-designed proteins that bypass DNA synthesis screening are a biosecurity problem, and mislabeled structures pollute public databases. A watermark that survives synthesis while preserving function turns provenance from a paperwork promise into a molecular property, and the Nature paper makes the method auditable by the whole field.
Google's new frontier model, announced on the official blog on October 1: a 1M-token output limit, up from 64K, with introductory pricing of $2/$10 per million input/output tokens. Rollout starts with trusted cyber defenders through the Fairwind Program, with U.S. government pre-release participation, before wider availability.
Frontier launches set the price and the safety posture for everything downstream. A defender-first rollout plus million-token outputs signals that the next competition is about long-horizon work, and about who gets the model first.
Iterative self-distillation lets LLM agents learn from successive deployments, a path toward recursive self-improvement. But the authors show deployment performance collapses across cycles, and even performance with privileged information declines. ReSAIL, a plug-in augmentation, prioritizes the interaction steps where privileged information most changes the teacher's predictions, balances distillation losses across trajectories, and regularizes the student's privileged outputs toward the frozen teacher. On ALFWorld and TextCraft it sustains gains over three cycles, with an average absolute gain of 22.5 percentage points in final-cycle success rates.
Self-improvement that only survives in the lab is not improvement. Documenting the collapse, and showing a concrete mechanism that mitigates it, moves recursive self-improvement from aspiration to engineering.
Predicting gene expression from H&E histology offers a scalable alternative to costly spatial transcriptomics, but most methods work at spot level, where many cells' signals blur together. CELLO pushes to single-cell resolution with one foundation-model forward pass per image, grid sampling that extracts location-specific features for all cells at once, and a distance-decay cross-attention module for local morphological context. On 52 public Xenium-H&E pairs from HEST-1k (12 organs, about 10 million cells), it beats the baselines on accuracy with a 14.0x average speed-up over DeepSpot2Cell.
Spatial transcriptomics is powerful and expensive; histology slides are cheap and everywhere. Single-cell prediction at 14x the speed turns every archived slide into a potential molecular assay, which is how a research method becomes clinical infrastructure.
One-class artifact detectors learn 'normal tissue' from a clean training pool and flag departures from it, but the pool is built by a preprocessing pipeline whose tissue-detection step is usually treated as neutral. On 16 annotated TCGA slides, the authors rebuilt the pool with different tissue detectors: saturation-Otsu excluded normal adipose and alveolar tissue while keeping marker ink, and swapping it for entropy-based detection cut the false-positive fraction on held-out clean slides from 0.102 to 0.016, with no loss of sensitivity. The gain came from the pool's composition, not its size.
Artifact detection is the unglamorous gatekeeper of every pathology AI pipeline, and its errors propagate silently downstream. Showing that a preprocessing choice, not the model, sets the false positive rate redirects debugging effort to where it actually pays off, and the authors argue the choice should be reported like any other hyperparameter.
Mycosis fungoides, a rare cutaneous T-cell lymphoma, is often misdiagnosed early because it mimics benign inflammatory skin disease. The authors combine a late-fusion ensemble of dual-magnification CNNs (10x for architectural context, 20x for cytological detail) with a random forest on 16 clinical features. On 6,267 images from 463 patients, the image model reached 83.58% accuracy and 89.13% sensitivity; the clinical model hit 96.6% accuracy and 93.8% sensitivity for staging confirmed cases.
Rare diseases are where AI diagnosis faces its hardest test: few examples, high stakes, and lookalikes everywhere. A dual-scale design that respects how pathologists actually read slides, context first, cytology second, is a template worth stealing for other rare entities.
Harness optimization automates the scaffolding around agents, a step toward recursive self-improvement, but existing methods learn one global harness applied uniformly to every task instance, and a harness that works on average may not be optimal for each instance. Turbo Harness recycles artifacts from a completed global optimization run into a structured playbook, then trains a harness editor that uses the playbook plus the instance at hand to generate instance-specific patches. It beats existing harness-optimization baselines across seven benchmarks spanning interactive agent tasks, software engineering, and long-horizon terminal tasks.
Agent performance is increasingly decided by the harness, not the weights, yet harness design is still one-size-fits-all craft. Treating the scaffold as a per-instance learning target is where the next gains in agentic reliability will come from.
Frontier models can produce strong mathematical ideas in a single shot, but open problems demand exploring multiple competing conjectures, clearing subtle technical obstructions, and retaining progress over a long horizon. Cogentic is a multi-agent harness for automated proof discovery: an orchestrator allocates a population of independent provers across proof directions, adversarial verifiers check their output, and confirmed results accumulate in a persistent verified ledger. Using Gemini as the base model, it produced novel results on five open problems in online learning, auction theory, and mechanism design, each independently verified by domain experts.
Automated proof discovery is where multi-agent coordination meets verifiable ground truth: a proof checks out or it does not. That makes it the cleanest testbed for whether orchestration adds real capability beyond a bigger single model, and five expert-verified results is a serious down payment.
LLM judges are increasingly used to evaluate AI-generated content, but a single judge is unreliable, and even a panel leaves a hard residue: when judges disagree, majority voting discards the conflict instead of resolving it. JuryFlow decomposes responses into atomic claims, has heterogeneous judges verdict each claim, and builds a disagreement graph scored by verdict entropy. A human acts as a structural guide, resolving one focal disagreement with a minimal intervention; the correction propagates along the graph and crystallizes into reusable rubric entries, making the evaluator progressively self-refining. It improves agreement with gold labels on MT-Bench and LLMBar over single-judge and majority-vote baselines.
Evaluation is the bottleneck of the agentic era, and judge disagreement is currently treated as noise. Turning disagreement into a precise signal for where evaluation is uncertain, and letting one human intervention compound into rubrics, is a more honest division of labor between people and judge models.
A preregistered vignette experiment with 1,030 US participants asked how patients judge physicians who disclose AI use. Disclosing physicians were rated significantly less warm and less competent, with lower willingness to book or recommend, an 'AI-use penalty' that barely moved with physician age, gender, or race. Exploratory analyses found the penalty more pronounced among younger participants and varying across physician-consumer demographic pairings.
Medical AI fairness is usually about model performance across groups. This paper measures a different inequity: the social penalty for admitting AI use, which could quietly punish the transparent physicians and reward the secretive ones. Deployment fairness has to include the humans around the model.
Decades of social psychology have produced validated interventions that reduce stereotypical thinking in humans. DIY translates five of them into debiasing procedures for LLMs, delivered through three paradigms: Show (in-context examples), Train (instruction tuning), and Revise (guided self-revision). Across three models, five bias benchmarks, and eleven debiasing baselines, Train+Revise and Revise rank top two, reaching as low as 2% mean bias at 90% reasoning accuracy and cutting bias on unseen dimensions by up to 14.8%. Code and data are public.
LLM debiasing is mostly prompt folklore. Grounding it in interventions that actually worked on human cognition gives the field a principled starting point instead of another round of trial and error, and the bias-reasoning tradeoff numbers set a bar others must now clear.
The EU AI Act's Article 27 requires deployers of high-risk AI systems to conduct Fundamental Rights Impact Assessments before deployment, but the evidence they need is fragmented across incompatible incident repositories, risk vocabularies, and legal texts. The authors present a reusable Semantic Web framework that consolidates this evidence, demonstrated on two high-risk public-sector use cases.
Impact assessments risk becoming checkbox paperwork precisely because the evidence is fragmented. A shared, machine-readable evidence layer is what turns a legal obligation into an actual audit, and this is the first reusable blueprint for it.
Unified multimodal models either quantize images into discrete tokens or bolt continuous image generation onto discrete language prediction. Multimodal Flow is fully continuous: text blocks and images become ordered continuous hyperchunks in embedding space, and a shared chunk-causal flow backbone learns one vector field over them with flow matching. MF-1, pretrained on only 150B tokens, averages 82.8 on GenEval and DPG-Bench and 75.3 on VQAv2, MMBench, and POPE, competitive with unified models trained on far more data. Code and model are public.
Quantization was a convenience, not a principle, and it taxes visual fidelity. A fully continuous flow over both modalities suggests the unified-model roadmap does not have to pay that tax, and the data efficiency here reframes what scale is actually required.
Text-to-image quality has scaled by growing models or adding denoising steps. This paper explores a third axis: repeatedly running shared Transformer blocks inside each denoising step, raising compute depth with the parameter count fixed. Naive looping fails from weak supervision across loops and unregulated attention updates; Looped-DiT fixes it with deep supervision across intermediate loops and self-modulating attention. A 260M-parameter looped model beats a 6.5x larger model across text-to-image benchmarks at 4.9x lower inference compute, and deeper loops progressively correct earlier mistakes, a hint of latent reasoning.
If compute depth inside the denoising loop can substitute for parameter count, the diffusion scaling playbook gains a cheaper dimension. The 'loops correct earlier mistakes' finding is the intriguing part: iteration starts to look like thinking.
Flow-matching voice conversion already delivers high speech quality and speaker similarity, and one-step models like MeanVoiceFlow promise efficient inference, but they depend on a computationally heavy content encoder. MeanVoiceFlow2 jointly optimizes the flow conversion module with a lightweight content encoder, trained by conversion distillation plus diffusion-GAN training with sample mixing. On zero-shot voice conversion it reaches higher perceptual quality at about 9x faster inference than MeanVoiceFlow, with comparable speaker similarity.
Voice conversion is where flow matching meets a real product surface: dubbing, accessibility, creative tools. Jointly optimizing the flow and its conditioning encoder is a pattern that generalizes well beyond speech, and 9x faster at higher quality is the kind of number that moves a method into products.