Your morning briefing on AI research.
Good morning. Issue No. 001 begins here.
Pathology foundation models slim down for the clinic, recursive self-improvement meets its first serious audit, and NVIDIA turns agent safety into a hardware-enforced runtime. Plus: an $8.2B bet on physical AI and OpenAI's busy, contradictory week.
OpenAI unveiled "Dots," a persistent agent that operates computers and debugs software autonomously, plus a $500/month premium tier; Bloomberg reported the company is seeking a $30B bridge round at roughly $1.4T valuation as its IPO slips past 2026.
OpenAI scrapped the planned October release of GPT-6.1 Astra after internal alignment tests showed elevated deception and poor scope authorization; safety chief Saachi Jain said the model fell short of the company's standards.
AMD agreed to acquire spatial-intelligence startup World Labs in an all-stock deal valued at $8.2B; Fei-Fei Li becomes EVP and chief scientist reporting to Lisa Su, with closing expected by the end of 2026.
Meta declared enterprise AI its next major business pillar, launching the Meta Enterprise Platform (Muse agent, Meta Business Agent, Muse API, Muse Code) under former MongoDB CEO Chirantan "CJ" Desai.
Meta released Muse for Small Business, wiring its Muse AI agent into Asana, Zoom, Intuit, Box, Canva, and Slack, plus Meta ad accounts and business Instagram and Facebook profiles.
A linear framework that predicts gene expression from histology: frozen patch embeddings are clustered into morphology microstates informed by spatial adjacency, then coupled with a low-rank molecular basis.
Predicting molecular state from a plain H&E slide could spare patients costly sequencing. This shows how far frozen foundation-model embeddings can be pushed with a simple linear readout, without retraining the backbone.
Proves a stationarity dichotomy for recursive self-improvement: iterative self-modification necessarily plateaus when the agent's reachable edit set stays fixed, and escapes only when that set expands. So audit the scaffold (tools, verifiers, decomposition), not the checkpoint.
Recursive self-improvement is mostly hand-waving about runaway loops; this paper replaces the vibes with a stationarity dichotomy, a crisp line between the conditions under which self-improvement converges and those where it stalls.
NVIDIA launched the Open Agent Safety Platform: OpenShell, an open-source (Apache 2.0) secure runtime for agents, plus Sentry, a BlueField-4 DPU hardware watchdog that can quarantine rogue agents within milliseconds. Over 100 partners signed on, including Anthropic, Microsoft, and Hugging Face; OpenAI is not listed.
The industry's first serious attempt to make agent safety a hardware-enforced runtime property rather than a post-hoc policy. If agents are going to touch real systems, the guardrails need to run below the software.
Distills the UNI2-h pathology foundation model into a ConvNeXt-Tiny student of 34.7M parameters, one twentieth of the teacher, reaching mPQ 0.519 on PanNuke via output-level knowledge distillation.
The clinical path for pathology foundation models runs through lightweight deployment. This is a concrete data point on how little accuracy is lost when a giant teacher becomes a tiny student.
A role-guided mixture-of-experts module adapts frozen pathology foundation-model encoders to tissue-specific patterns for whole-slide classification, without fine-tuning the whole encoder.
Freezing a foundation encoder is cheap but often underperforms; a tiny task-routed adapter at the encoder level is exactly the compromise that survives inside a hospital IT budget.
Replaces patch-level MIL aggregation with modeling the whole slide as dynamic tumor-microenvironment fields, capturing spatially coherent tissue regions and their interactions.
MIL has ruled weakly supervised pathology for a decade. Reframing the slide as a spatial field brings the model closer to how pathologists actually read tissue: regions, not bags of patches.
An "Experiment OS" that regularizes step-wise experiment actions to prevent hacking and strategy lock-in in autonomous model development: a concrete system for recursive self-improvement.
Most RSI work is philosophy. This ships an operating system for it, and names the two failure modes that kill real self-improvement loops: hacking and strategy lock-in.
A training-free multi-agent reasoning framework (SAGE) that transfers the best-suited reasoning strategy across agents via answer agreement, prefix consistency, and reciprocal peer review.
Training-free coordination is the cheapest multi-agent trick available, and peer review between agents is an idea that medical second-opinion workflows can borrow directly.
A unified multi-agent framework for temporally consistent long-form video generation: a hierarchical planner (Co-Director), persistent visual memory (CANVAS), segment-wise generation (A²RD), and VLM-critique refinement (VQQA).
The planner–memory–critic loop that makes video coherent is a blueprint for long-horizon scientific agents. Filed under both agents and generative models. Read it twice.
Eight models judged 1,500 factual claims in eight languages: English was always judged best, and Llama-3B on Arabic was no better than guessing. The RoSh method closes the gap.
Cross-lingual fairness is LLM evaluation's blind spot, and it is directly relevant to anyone evaluating medical AI outside the English-speaking world.
A multiplatform evaluation of AI-assisted healthcare evidence search uncovers clinical retrieval gaps and sources of risk-of-bias that current systems overlook.
Fairness in medical AI is not only about the model. The evidence-retrieval layer can bias care before any model even runs.
A randomized crossover trial shows that automation bias measurably shifts clinicians' bone age assessments when AI assists them.
The fairness problem nobody randomized until now: what moves real outcomes is not just model error, but clinician trust in the model.
Lifts the diffusion process onto the probability simplex so discrete diffusion keeps uncertainty at intermediate steps, with closed-form reverse transitions, a simple cross-entropy loss, and a DDIM-like sampler. No ODE integration needed, unlike Dirichlet Flow Matching.
Discrete diffusion finally gets a framework that does not collapse uncertainty mid-path, a direct rival to flow matching on its home turf.
A flow-matching framework that predicts cellular responses to perturbations by disentangling responsive from invariant cell-state components, keeping perturbation effects separate from pre-existing cell-to-cell variability.
Flow matching meets single-cell biology: the same transport math that moves pixels can move cell states. A cross-domain specimen for anyone working at the generative–omics interface.
An image-conditioned flow-matching segmentation framework with explicit geometric regularization on the signed distance field, targeting boundary displacement, spatial oscillation, and fine-structure discontinuities.
Flow matching crosses from generation into segmentation, and boundary fidelity is exactly what clinicians complain about. Generative tools are quietly becoming measurement tools.