Taiwan · Est. 2026
Back Issues

The Daily Feed

No. 010 Friday, October 9, 2026

Thousands of LLM agents left to debate politics polarize just like humans do (polarized agents), diffusion models get put on a physical clock (analytical variance schedules), and a new gym measures whether agents can really improve themselves (RSIGym). Also: Google's agent for work, lawmakers vs. Google's data grab, and Gemini's free tier shrinks.

01 · Daily News

The last 24 hours in AI.

N1
Google Cloud launches a Gemini agent for work across Workspace, Microsoft 365 and Slack
October 8, 2026 · via Reuters

Google Cloud launched a single AI agent for work that answers questions, handles tasks, creates content and writes code inside Google Workspace, Microsoft 365 and Slack, picking the best model per task across Gemini and Anthropic Claude. So-called coworker agents can even act as team members with their own email addresses.

N2
OpenAI's annualized revenue about $20B less than previously signaled, FT reports
October 8, 2026 · via Reuters

The Financial Times, citing investor documents, reported OpenAI recently told investors its annualized revenue was approaching $50 billion at end-September, far short of the $70 billion figure reported late last month. The gap comes from different calculation methods: Anthropic counts cloud-partner revenue, OpenAI does not.

N3
120+ US lawmakers object to Google buying Spirit Airlines data for AI training
October 8, 2026 · via Reuters

More than 120 US lawmakers, led by Senator Elizabeth Warren and Representative Steven Horsford, objected to Google acquiring defunct carrier Spirit Airlines' internal data for $10 million to train AI systems, reportedly including 100 million emails and 500 million Teams messages, and asked the company to exclude employee information.

N4
Leaders of Trump's 'Super Intelligence Task Force' to meet Thursday
October 8, 2026 · via Reuters

The four leaders of President Trump's 'Super Intelligence Task Force' were slated to meet Thursday, vice chair Scott Kupor told Reuters, with the full committee expected to convene next week and release a charter outlining its goals.

N5
Google cuts free Gemini app to Flash-Lite only starting October 9
October 9, 2026 · via Android Police, 9to5Google

Starting today, personal Google accounts without a paid AI plan lose Flash and Pro in the Gemini app and keep only Flash-Lite; AI Plus subscribers lose Pro. The change, first spotted by 9to5Google in Google's support pages, draws the clearest line yet between paying and non-paying users.

N6
NVIDIA commits $1B to US AI-for-science research over five years
October 8, 2026 · via NVIDIA, HPCwire

NVIDIA announced a $1 billion, five-year commitment to AI-for-science research at US higher-education institutions, targeting quantum computing, healthcare and energy security, plus cloud support for government missions. The pledge was made at the 'Science: A New Golden Age' event in Washington, D.C.

02 · Selections

Three papers, chosen first.

01
Editor's judgment: the day's most unsettling agents finding. Thousands of agents across five model families all converge on polarization, which suggests the pattern lives in the interaction structure rather than any single model's politics. The intervention testbed is what makes this science rather than just a warning.
Nature CommunicationsCross-domain
Emergence of polarization in networks of large language model agents

Researchers simulated thousands of LLM agents on GPT-3.5, GPT-4o, ChatGLM, Llama-3 and DeepSeek-V3, letting them form social networks and debate political issues through LLM-guided conversation. The agents spontaneously built human-like networks with homophilic clustering, and their opinions polarized over time. The same setup doubles as a testbed for interventions such as encouraging diverse interactions and curbing confirmation bias.

Why it matters

Personal agents will soon talk to each other far more than they talk to us, and their collective behavior becomes a social question, not just an engineering one. Showing that polarization emerges from agent interaction itself, plus a testbed to trial fixes, gives researchers a place to work on the problem before it plays out in the wild.

02
Editor's judgment: a principled fix to a habit the field inherited from image generation. Diffusion models used as scientific simulators need clocks set by physics, not by what made pretty pictures, and deriving the noise schedule from a real variance law is exactly that correction.
Nature Communications
Generative diffusion surrogates with analytical variance schedule

The team anchors a diffusion model's noise schedule to a known physical variance law, the way particle spread actually grows over time in stochastic transport, instead of the usual heuristic linear or cosine schedules. The variance is then fixed analytically while the network learns only the remaining distributional shape, enabling entrance-only training and calibrated surrogates for problems like laboratory plasma transport.

Why it matters

Diffusion models are becoming scientific simulators, but a simulator whose internal clock is arbitrary cannot be trusted for physical prediction. Tying the generative clock to real transport laws turns these models from curve-fitters into calibrated instruments, which is what science actually needs from them.

03
Editor's judgment: the most concrete recursive self-improvement result this week. An environment that lets agents jointly optimize data, training and harness under a shared budget, with a single index to compare models, turns self-improvement from philosophy into something measurable. Opus 5 closing nearly half the remaining gap is the number to watch.
arXiv
RSIGym: A Flexible Environment for Recursive Self-Improvement

RSIGym is an agent-native research environment built on an 'everything as a service' design: training, inference, rollout, evaluation and sandbox execution are reusable services with shared budget and permission controls. Its RSI-Index measures the mean fraction of remaining performance gap closed across five benchmarks; in joint runs, Opus 5 scored 0.4809, lifting SWE-bench Verified from 17.67% to 50.33% and AIME from 31.67% to 97.78%.

Why it matters

Recursive self-improvement is bottlenecked by infrastructure: research agents waste their budgets rebuilding training pipelines instead of doing research. A shared, budgeted environment that lets agents study data, training and harness changes together is the missing lab bench for measuring whether agents can genuinely improve AI systems.

03–06 · Departments

Four columns, every issue.

03

Pathology

數位病理
arXiv
Masked Feature Encoding for Large-Scale Whole Slide Image Representation

MFE-MIL adds a lightweight MLP adapter to frozen-encoder multiple instance learning pipelines, jointly trained with a window-based masked reconstruction objective that suppresses within-slide patch variance from staining, scanners and texture. It needs no patch coordinates at training or inference, and improves accuracy and F1 across CAMELYON16/17, PANDA and TCGA-BRCA, plus survival concordance on five TCGA cohorts.

Why it matters

Most pathology AI pipelines quietly suffer from stain and scanner variation that drowns out the diagnostic signal. A plug-and-play adapter that cleans up feature variance without retraining giant foundation models is the kind of practical fix that can actually reach clinics.

arXiv
TIRA: Tumor Immune Representation Adaptation for Zero-Shot Cross-Cancer MSI and TMB Prediction

TIRA conditions frozen pathology foundation-model tile attention on spatial immune topology, using source-derived immune organization as a biological prior for joint MSI and TMB prediction. Trained on colorectal cancer cohorts, it transfers zero-shot to gastric and endometrial cancers, lifting UNI2 zero-shot AUROC on TCGA-STAD from 0.633 to 0.766 for MSI and from 0.651 to 0.772 for TMB.

Why it matters

MSI and TMB decide who gets immunotherapy, but sequencing every patient is expensive. A model that predicts these biomarkers from routine slides and actually generalizes across cancer types brings that triage closer to hospitals that cannot afford molecular testing for everyone.

arXiv
One-Slide Calibration of Pathology Foundation Models

SlideRuler calibrates scanner-induced embedding shifts in frozen pathology foundation models using regions within the slide itself as internal controls. A transfer map learned from paired rescans corrects a single scan at inference, cutting target-to-source embedding distance by 16.3 to 38.5 percent across five SCORPION scanners, without needing scanner identity or target-cohort statistics.

Why it matters

A model that works on one hospital's scanner and silently degrades on another's is a patient-safety problem. Calibration that draws its correction signal from the slide being examined, rather than from a reference cohort, is a practical path to making frozen foundation models behave consistently across devices.

04

Agents

智能體
arXiv
A Society of Researchers: Designing Institutions for Populations of Autonomous Research Agents

The paper argues that populations of thousands of research agents sharing one compute pool will acquire an organization whether designers provide one or not, so it proposes explicit institutions: competing principal investigators, proposal calls, independent review, grants, and a human 'mayor' who allocates resources but assigns no tasks. Six principles and six open problems structure the agenda.

Why it matters

As agent deployments scale from single projects to populations, the failure mode is not a bad agent but a bad organization: herding, duplicated effort, self-organized workarounds. Designing institutions for agent societies before they emerge on their own is governance work the field cannot postpone.

arXiv
Homogenization in Multi-Agent Systems

The authors operationalize homogenization, the tendency of interacting agents to converge on similar behaviors, with three metrics: conformity, polarization and inertia. Across code generation, hiring and peer review, it creates concrete risks: correlated security blind spots, bias that persists after the biased agent is removed, and converging evaluation standards. Simple fixes like resampling or mixed-model teams fail to reduce it.

Why it matters

Multi-agent systems are sold on the promise that diverse agents check each other's blind spots. If interaction itself erases that diversity, the whole value proposition collapses, and the finding that easy diversity fixes do not work means the field needs genuinely new interventions.

arXiv
SkillForge: Co-Evolving Skills and Agents via Dynamic Skill Lifecycles

SkillForge is an agentic RL method that co-evolves a skill library alongside the policy through a fitness-driven lifecycle of trial, active, stable and retired states. A pre-RL phase retires low-fitness seed skills using the base model's own rollouts; RL then continues selective retirement, stabilization and LLM-guided mutation each iteration. It tops the strongest baseline on three interactive agent benchmarks with up to 7.8% relative improvement while keeping the library compact, and releases the 5k-record SkillFurnace dataset.

Why it matters

Agent skill libraries rot: once-useful skills turn counterproductive as the policy matures, and append-only accumulation pollutes the agent's context. Treating the library as a population under selection, with retirement as a first-class operation, is how long-lived agents can keep learning without drowning in their own memory.

05

Fairness

公平性
arXiv
Robust Decentralized Fairness Auditing

Auditopus is a decentralized protocol for fairness auditing of LLMs: multiple auditors query the model with their own private query sets and exchange only cumulative statistics vectors, never raw queries. A defense down-weights auditors whose reports are statistically inconsistent with their own history, cutting audit error by up to 78% against adversaries that fabricate numbers to 'fairwash' an unfair model.

Why it matters

Regulations like the EU AI Act require fairness audits, but a single auditor rarely has representative data, and auditors themselves can be captured. A decentralized audit that survives adversarial participants is infrastructure for making AI regulation actually enforceable rather than performative.

arXiv
Justice After Identity: Large Language Models and the View from Everywhere

A solo-authored conceptual paper asks whether LLMs, which encode many human identities without possessing one, could approximate Rawls' 'original position': a view from everywhere built by computationally incorporating identity diversity rather than excluding it. It compares human, base-model and frontier-model judgments on classic moral dilemmas under varying identity constraints.

Why it matters

Fairness debates keep circling the question of whose perspective counts. The paper's provocation, that a system with no identity of its own might synthesize the fairest view of all, reframes the conversation from removing bias to composing perspectives, whether or not one buys the conclusion.

arXiv
A Systematic Investigation of Bias in Large Language Models for Advertising Relevance

A Microsoft team systematically audits LLM fairness as ad-relevance judges, using a counterfactual framework over advertiser identity, input language and demographic wording. Both GPT-4o and a fine-tuned Qwen-7B shift their relevance judgments when only the advertiser name or language changes, with gender-occupation stereotype patterns; masking company names and rebalancing training labels help, but only under the right conditions.

Why it matters

Ad systems decide which businesses get seen, and an LLM judge that quietly favors famous brands or penalizes non-English queries is a market fairness problem hiding inside a relevance score. A systematic audit with mitigation experiments gives practitioners something the field rarely offers: a fairness checklist grounded in real ad logs.

06

Generative Models

生成式模型
arXiv
GRACE: Generation-Aware Latent Compression for Efficient Video Generation

GRACE compresses a pretrained video autoencoder for diffusion generation without retraining from scratch: a frozen base latent plus a learned residual, aligned to the frozen DiT's feature space so the autoencoder optimizes for generation rather than reconstruction. On Wan2.1-I2V-14B it cuts tokens nearly 8x and latency 11.1x while matching VBench quality.

Why it matters

Video generation is hitting a compute wall: quality scales with tokens, and tokens cost quadratically. Compression that preserves what the generator actually needs, rather than what reconstruction metrics reward, is the practical route to video models ordinary hardware can run.

arXiv
QuadTok: Quadtree Visual Tokenizer for Autoregressive Image Generation

QuadTok replaces fixed 2D token grids with a hierarchical quadtree tokenizer that spends tokens on intricate regions and stays coarse on smooth ones, keeping explicit spatial correspondence. Its 947M GPT-style generator reaches 2.08 gFID on ImageNet 256x256, and the tree structure enables zero-shot spatially controlled generation from a user-supplied layout.

Why it matters

Fixed token grids waste capacity on empty sky and starve detailed textures; adaptive allocation is the obvious fix that has been surprisingly hard to get right. A tokenizer that is both spatially grounded and controllable by layout brings autoregressive image models closer to being design tools rather than slot machines.

arXiv
Seq-Flow: Efficient Probabilistic Forecasting with Self-Rollout Error Control

Seq-Flow rethinks probabilistic forecasting: instead of generating each forecast from Gaussian noise, its flow ODE transports the previous forecast distribution to the updated one, since successive forecasts usually differ only modestly. Self-rollout training on its own forecasts controls error accumulation, cutting CRPS by 65% on particle-accelerator beam spill forecasting under tight sampling budgets and staying stable over 400+ updates.

Why it matters

Weather, plasma control and accelerator operations all need distributions updated as new data arrives, and regenerating from noise every time is pure waste. Learning the update itself, with training that anticipates its own compounding errors, is how generative models become real-time forecasting infrastructure.