π WELCOME TO METAMESH.BIZ +++ OpenAI's Chief Research Officer wants 5-10% of compute shifted to safety, which is either reassuringly responsible or terrifyingly insufficient depending on your priors +++ Gemini 4 Argon now outputs 1M tokens at a time because Google looked at the context window arms race and said "hold my TPU" +++ OpenAI catches Moonshot AI running a coordinated distillation campaign, proving the real scaling law is how fast your competitors can copy you +++ THE FUTURE IS 31% AI-GENERATED AND NOBODY CAN TELL THE DIFFERENCE β’
π WELCOME TO METAMESH.BIZ +++ OpenAI's Chief Research Officer wants 5-10% of compute shifted to safety, which is either reassuringly responsible or terrifyingly insufficient depending on your priors +++ Gemini 4 Argon now outputs 1M tokens at a time because Google looked at the context window arms race and said "hold my TPU" +++ OpenAI catches Moonshot AI running a coordinated distillation campaign, proving the real scaling law is how fast your competitors can copy you +++ THE FUTURE IS 31% AI-GENERATED AND NOBODY CAN TELL THE DIFFERENCE β’
Google DeepMind protein watermarking (SynthID Bio)
2x SOURCES ππ 2026-09-30
β‘ Score: 8.1
+++ DeepMind's SynthID Bio embeds invisible signatures into AI-designed proteins, finally giving biosecurity researchers a fighting chance to track what came from the algorithm versus the lab bench. +++
via Arxivπ€ Jenna Russell, Ben Glickenhaus, Katherine Thai et al.π 2026-09-30
β‘ Score: 7.9
"Web text makes up the majority of pretraining data and is increasingly AI-generated. After applying FineWeb quality filtering, we find that 27.5% of tokens from June 2026 web data are labeled as AI-generated by Pangram, rising to 31.1% by August. Unlike synthetic data or model-collapse setups, this..."
π SECURITY
OpenAI model distillation campaign by Moonshot AI
2x SOURCES ππ 2026-09-30
β‘ Score: 7.3
+++ OpenAI's accusation that Moonshot AI orchestrated a coordinated model-distillation campaign reveals the awkward reality that copying smart models is both trivially easy and increasingly hard to ignore. +++
via Arxivπ€ Paras Dahal, Anton Bakhtin, Taco Cohen et al.π 2026-09-29
β‘ Score: 7.3
"As agents take on longer and more complex problems, controlling the execution becomes a task in its own right. Each step in the run brings new control choices, like which partial work to build on, whether to start fresh, or when to stop. We introduce agentic meta-reasoning, an inference-time harness..."
via Arxivπ€ Sohail, Sarkar, Shakuntala Baichooπ 2026-09-30
β‘ Score: 7.0
"Sampling several answers and keeping the one a verifier scores highest is one of the simplest ways to buy accuracy at test time. Its effect is reported as a scaling curve: accuracy against the number $k$ of sampled answers. The curve is cheap to draw and expensive to trust. A budget read off it is c..."
via Arxivπ€ Yanbei Chen, Anirudh Goyal, Raghuraman Krishnamoorthiπ 2026-09-30
β‘ Score: 7.0
"Looped transformers and Mixture-of-Experts (MoE) offer complementary routes to efficient scaling: recurrence increases computational depth at fixed parameters, while MoE sparsity expands total capacity at fixed active compute. Yet existing scaling laws model recurrence or sparsity in isolation. In t..."
via Arxivπ€ Paul Le Van Kiem, Dario Shariatian, Umut Simsekli et al.π 2026-09-30
β‘ Score: 6.9
"Continuous diffusion language models generate all tokens in parallel, yet high-quality generation can still require hundreds of network evaluations (NFEs). We study how distributional distillation can reduce this cost by exploiting the student's probabilistic token outputs. Our unified formulation c..."
via Arxivπ€ Tyler Skow, Shravan Chaudhari, Rama Chellappa et al.π 2026-09-30
β‘ Score: 6.9
"Unlearning a fact in one language does not guarantee its removal in others as changing the query or even the requested answer language can reopen seemingly forgotten knowledge -- a cross-lingual loophole. The most straightforward solution to this challenge -- unlearning in all languages -- is neithe..."
via Arxivπ€ Sachi Shome, William Eiersπ 2026-09-30
β‘ Score: 6.9
"Lossy compression is widely used in Federated Learning (FL) but is generally treated as an error source, while conventional poisoning defenses inspect update geometry. In this work, we instead treat the compressor's response as a security signal: the input-dependent distortion and payload behavior i..."
via Arxivπ€ Yong Du, Tongbo Chen, Zhengxi Lu et al.π 2026-09-30
β‘ Score: 6.9
"Online training enables computer-use agents (CUAs) to improve through interaction with executable environments. However, existing methods primarily rely on sparse outcome rewards, which provide no supervision for intermediate actions. On-policy self-distillation (OPSD) offers token-level learning si..."
via Arxivπ€ Ganesh Pavan Kartikeya Bharadwaj Kolluri, Michael Kampouridis, Ravi Shekharπ 2026-09-29
β‘ Score: 6.9
"Speech-LLMs are expensive to run, making compression important for real-world deployment. However, compressed models are usually selected using aggregate word error rate (WER), which can hide how pruning affects different demographic groups. In this work, we systematically study the effect of audio..."
via Arxivπ€ Yu Xu, Yuxin Zhang, Xiao Yang et al.π 2026-09-29
β‘ Score: 6.9
"Mixture-of-Experts (MoE), popularized by large language models, is a promising paradigm for scaling visual generative models. However, conventional token-wise MoE routes tokens independently within a homogeneous expert pool and regularizes expert usage toward uniformity, making it poorly matched to..."
via Arxivπ€ Anmol Kabra, Swathi Saravana Selvam, Albert Gong et al.π 2026-09-30
β‘ Score: 6.8
"Training LLM agents with reinforcement learning (RL) is bottlenecked by environments, which must provide verifiable rewards, support long-horizon interaction, and scale cheaply. Existing approaches rely on costly human-curated data or on LLM-generated environments that risk hallucinations and benchm..."
via Arxivπ€ Young-Jun Lee, Jinheon Baek, Soyeong Jeong et al.π 2026-09-30
β‘ Score: 6.8
"Evolutionary search with large language models (LLMs) can stall when progress requires external knowledge the model lacks. Supplying relevant documents helps, but simply adding web search tool can keep returning the same pages as solutions change. We introduce EvoDuet, a bi-level optimization method..."
via Arxivπ€ Yong Xien Chng, Tianyi Chen, Wenwen Tong et al.π 2026-09-30
β‘ Score: 6.8
"Improving text-to-image models has traditionally relied on increasing model size or the number of denoising steps. In this work, we explore an alternative way to scale computation by repeatedly running shared Transformer blocks within each denoising step, effectively increasing computational depth w..."
via Arxivπ€ Razan El Mais, Ali Chehab, Ibrahim Issa et al.π 2026-09-30
β‘ Score: 6.8
"Differentially Private Stochastic Gradient Descent (DP-SGD) is a leading approach for privacy-preserving fine-tuning of large language models (LLMs). Many decoder-only LLMs employ weight tying between input and output embeddings, a design choice originally introduced for parameter efficiency and imp..."
via Arxivπ€ Yinghui He, Yapei Chang, Khushi Bhardwaj et al.π 2026-09-30
β‘ Score: 6.8
"On-policy distillation (OPD) is a promising approach for training language agents, providing dense teacher supervision on student-generated trajectories. However, in multi-turn interaction, an incorrect action changes the states the student encounters later, so errors compound across turns. In preli..."
via Arxivπ€ Yang Cai, Vineet Gupta, Yanchen Jiang et al.π 2026-09-30
β‘ Score: 6.8
"We present Cogentic, a multi-agent harness for automated proof discovery on open research problems. While frontier language models can generate strong mathematical ideas in a single shot, single-shot generation is often insufficient for open problems that require exploring multiple competing conject..."
via Arxivπ€ Zhenyu Wang, Tianze Wang, Linjun Zhang et al.π 2026-09-29
β‘ Score: 6.8
"On-policy distillation (OPD) trains a student on its own generated responses using dense, token-level supervision from a stronger teacher. Vanilla OPD treats all teacher signals equally, assuming that the teacher's supervision is equally important for every token. However, teacher signals at differe..."
via Arxivπ€ Pranjal Aggarwal, Lawrence Keunho Jang, Sean Welleck et al.π 2026-09-30
β‘ Score: 6.7
"Computer use agents (CUAs), which use graphical user interfaces (GUIs) to complete tasks on a computer, have recently surpassed human performance on many standard benchmarks, including difficult long-horizon tasks. Their capabilities are undoubtedly impressive, however, a key barrier to the widespre..."
via Arxivπ€ Junshu Pan, Zhizhang Fu, Shulin Huang et al.π 2026-09-30
β‘ Score: 6.7
"Reinforcement learning with verifiable rewards (RLVR) has improved the reasoning capabilities of large language models (LLMs), yet their predictions remain sensitive to task-irrelevant prompt features. We investigate this sensitivity through semifactual prompt interventions that preserve the underly..."
"Recent autonomous machine learning engineering (MLE) agents have made significant progress on public leaderboards. Often motivated by progress stagnation over long-horizon cycles and limited Large Language Model (LLM) primitives, modern MLE agents are deployed on top of increasingly elaborate machin..."
via Arxivπ€ Dor Tirosh, Ido Amos, Mor Gevaπ 2026-09-29
β‘ Score: 6.7
"Transformer language models (LMs) are feed-forward: deep-layer representations are never fed back to shallower layers, and the only pathway for information to flow downward across generation steps is the decoded token. This narrow channel forces models to recompute intermediate results and to discar..."
via Arxivπ€ Arav Dhoot, Punya Syon Pandey, Jamie Johnson et al.π 2026-09-29
β‘ Score: 6.6
"Risk aversion in resources could prevent misaligned AI agents from causing catastrophic harm. Misaligned but risk-averse agents would tend to favor safer strategies like making deals with humans over riskier strategies like rebelling. We train agents to be risk averse through character training, fin..."
via Arxivπ€ Ratish Puduppully, Pranabendu Misra, Paarth Iyer et al.π 2026-09-29
β‘ Score: 6.5
"Chain-of-thought traces are widely read as records of how models reach their answers, informing debugging, agent auditing, and claims about reasoning. Testing this interpretation is difficult because natural-language thinking traces are rarely mechanically verifiable. We revisit it in iGSM, a synthe..."
via Arxivπ€ Subba Reddy Oota, Francisco Herrera, Jordi Cabot Sagrera et al.π 2026-09-29
β‘ Score: 6.4
"Large language models (LLMs) enable agents to solve long-horizon tasks by generating a plan and then executing it in an environment. However, successful planning requires two distinct capabilities: selecting an appropriate plan for the task and executing it faithfully. Existing planner--executor sys..."
via Arxivπ€ Edoardo Bolzoni, Valerio Capraroπ 2026-09-29
β‘ Score: 6.1
"Understanding gender biases in large language models (LLMs) is increasingly important as these systems become embedded in decision-support tools with real consequences. Prior research has focused only on a small set of models, leaving open the extent to which gender biases are common and heterogeneo..."
via Arxivπ€ Quang Hieu Pham, Thuy Duong Nguyen, Jocelyn Qiaochu Chen et al.π 2026-09-29
β‘ Score: 6.1
"Language-model (LM) harnesses enable LMs to operate effectively over long contexts using additional compute. However, existing long-context evaluations are insufficient for distinguishing modern harnesses, reflected by saturated accuracy across harnesses and largely similar evaluation costs. In this..."
OpenAI and Anthropic dropped next-generation models, paused training over agent escapes, leaked user data, and helped form a safety body, all in the same week, in roughly that order.