π WELCOME TO METAMESH.BIZ +++ Kimi K3 drops with 2.8T parameters, 896 routed experts, and a 1M token context window because frontier models now need their own zip codes +++ Google teaching quantum computers to correct their own errors via reinforcement learning, one step closer to making classical compute look quaint +++ Some Claude conversations found publicly accessible online, proving the real AI safety problem was URLs all along +++ THE FUTURE IS 104 BILLION ACTIVATED PARAMETERS AND ZERO ACTIVATED PRIVACY SETTINGS β’
π WELCOME TO METAMESH.BIZ +++ Kimi K3 drops with 2.8T parameters, 896 routed experts, and a 1M token context window because frontier models now need their own zip codes +++ Google teaching quantum computers to correct their own errors via reinforcement learning, one step closer to making classical compute look quaint +++ Some Claude conversations found publicly accessible online, proving the real AI safety problem was URLs all along +++ THE FUTURE IS 104 BILLION ACTIVATED PARAMETERS AND ZERO ACTIVATED PRIVACY SETTINGS β’
+++ Microsoft rolled out MAI-Cyber-1-Flash and its vulnerability-hunting companion, claiming industry-leading performance at half the price. The real test? Whether security teams actually trust AI to patch their most paranoid nightmares. +++
+++ Nvidia is betting heavily that Sutskever's newly-minted SSI will actually need the compute firepower they're providing, a vote of confidence in both the startup's ambitions and Nvidia's ability to monetize the AGI race's infrastructure needs. +++
via Arxivπ€ Kimi Team, Tongtong Bai, Yifan Bai et al.π 2026-07-27
β‘ Score: 7.4
"We introduce Kimi K3, a 2.8T parameter Mixture-of-Experts model with 104 billion activated parameters, native vision capabilities, and a 1-million-token context window. Kimi K3 is built on Kimi Delta Attention and Attention Residuals, which improve information flow across sequence length and model d..."
via Arxivπ€ Bianca Raimondi, Davide Evangelista, Maurizio Gabbrielli et al.π 2026-07-27
β‘ Score: 7.0
"Large Language Models can produce fluent text that is false, unsupported by the available evidence, or inconsistent with information that appears to be internally represented by the model. We study hallucination detection from the geometry of hidden activations and introduce the D-Score, a simple sp..."
via Arxivπ€ Phu Gia Hoang, Anwoy Chatterjee, Tanmoy Chakraborty et al.π 2026-07-27
β‘ Score: 6.9
"The wide-scale use of sparse autoencoders (SAEs) as interpretability tools is limited by inconsistent links between SAE features and model behavior. Features with clear activation descriptions may have weak or unexpected causal effects; steering can vary across prompts or oppose the intended directi..."
via Arxivπ€ Davide Scarso, Hugo Noronha de Almeida, Joaquim Pinaπ 2026-07-24
β‘ Score: 6.9
"Commercial large language models are increasingly used as knowledge references, yet their stance on contested scientific claims is neither stable nor transparent. We tested how four major LLM families (Claude, Grok, GPT, Gemini) evaluate ethnonationalist pseudo-science derived from Frank Salter's bi..."
via Arxivπ€ Ivo Verhoeven, Pushkar Mishra, Ekaterina Shutovaπ 2026-07-27
β‘ Score: 6.8
"This paper studies what discriminatively trained reward models (RMs) memorize by measuring counterfactual memorization on two human preference datasets. We show that RMs 1) misallocate memorization to easy, high margin preference pairs, 2) memorize dataset-specific shortcuts (e.g., model identity, u..."
via Arxivπ€ Arseny Kravchenko, Vadim Liventsev, Innokentii Konstantinov et al.π 2026-07-27
β‘ Score: 6.8
"Autonomous LLM agents processing mixed-confidentiality data face severe security risks from prompt injection attacks and reasoning errors. While dynamic Information Flow Control (IFC) provides structural security guarantees, traditional taint tracking permanently taints an agent's context upon readi..."
via Arxivπ€ Ritik Raj, Souvik Kundu, Sarbartha Banerjee et al.π 2026-07-24
β‘ Score: 6.8
"Routing to select large language models (LLMs) with different cost-quality trade-offs has become a fundamental deployment feature of enterprise AI. Existing routers, primarily make independent routing decisions for each LLM call. However, agentic applications execute as long-horizon workflows whose..."
via Arxivπ€ Hong Liu, Yuan Cheng, Lin Niu et al.π 2026-07-27
β‘ Score: 6.7
"Token-level sparse attention, as implemented by DeepSeek Sparse Attention (DSA) in production systems, makes the downstream attention efficient but shifts the bottleneck to the indexer that feeds it. To select the top-k tokens for each query, the indexer must still score every preceding token, incur..."
via Arxivπ€ Maruthi Vemula, Neeraj Praneeth Gajulaπ 2026-07-27
β‘ Score: 6.7
"A language model with a bounded working memory must repeatedly decide which stored items to keep. Every deployed method decides the moment an item arrives, from the past (StreamingLLM, H2O) or from a guess about the future (SnapKV). We recast the choice as an estimation problem on a hidden signal, w..."
via Arxivπ€ Siyuan Huang, Pengyu Cheng, Haotian Liu et al.π 2026-07-24
β‘ Score: 6.7
"LLM training is shifting from manual design and annotation to interaction-driven self-evolution. However, existing self-evolutionary methods face a fundamental dilemma between task diversity and verification reliability: environment-bound methods obtain precise feedback but confine learning to narro..."
"Enterprise AI agents are typically granted static credential sets at configuration time, holding every tool the role might need for every task they perform. This persistent over-privilege expands the attack surface. We argue that capability scoping must follow a dynamic least-privilege principle and..."
via Arxivπ€ Darshan Tank, Baran Namaπ 2026-07-24
β‘ Score: 6.7
"Adding procedural skills to an LLM agent is typically evaluated by average improvement in task success. However, this metric hides an important cost: skills can also make agents worse. We measure both sides by comparing agents with and without skills across nearly 6,000 runs spanning two office auto..."
via Arxivπ€ Xueping Gao, Jianwei Yang, Qiang Yangπ 2026-07-27
β‘ Score: 6.6
"Generate--test--revise loops are common in coding agents, but repetition alone provides no reliability guarantee. We study the gap between finding a correct patch and retaining, verifying, and submitting it. A sealed five-seed study over 30 HumanEval repairs produces 900 three-revision trajectories...."
via Arxivπ€ Tianyi Men, Zhuoran Jin, Kang Liu et al.π 2026-07-27
β‘ Score: 6.6
"Multi-turn long-horizon planning is critical for foundation model agents, yet how to fundamentally improve it remains unclear. Existing models are trained on uncontrollable and opaque Internet data, making it difficult to identify how planning ability is acquired, shaped, and integrated. To address..."
via Arxivπ€ Jhonatan Tavori, Gur-Eyal Sela, Ion Stoica et al.π 2026-07-27
β‘ Score: 6.6
"Inference systems increasingly combine a fast path that returns predictions within the application's latency deadline together with a higher-accuracy slow path that runs higher-compute methods on stronger, remote hardware, so its results can be returned on time and combined with the fast path predic..."
via Arxivπ€ Zhen Huang, Yikun Wang, Shijie Xia et al.π 2026-07-27
β‘ Score: 6.5
"Pretraining data processing is critical to the downstream performance of Large Language Models (LLMs). However, many existing approaches define a fixed processing strategy at the corpus or domain level and apply it uniformly to many examples, without adapting to the needs of each example. We propose..."
via Arxivπ€ Yanhao Jia, Jiepeng Wang, Haibin Huang et al.π 2026-07-27
β‘ Score: 6.5
"Modern large language models scale successfully by pairing capacity growth with efficiency, keeping per-token and deployment costs under control as capacity grows. AIGC Foundation Models (AFMs), especially diffusion-transformer backbones, have begun to adopt sparse experts, but recent efforts mostly..."
via Arxivπ€ Rajat Sainju, Dariusz Jarosz, Hairong Shang et al.π 2026-07-27
β‘ Score: 6.4
"Scientific user facilities accumulate decades of operational knowledge that no single search index covers: electronic logbooks, technical documents, internal wikis, operations chat messages, maintenance records, and live control-system data. We present APS-RAG, Advanced Photon Source Retrieval Augme..."
via Arxivπ€ Shixin Fang, Jiachen Wo, Wenjuan Qin et al.π 2026-07-24
β‘ Score: 6.4
"Large language model (LLM) evaluation spans diverse tasks and benchmarks, yet evidence remains organized around tasks rather than the capabilities they probe. This fragmentation limits cross-study comparison, obscures capabilities tasks recruit, and makes coverage gaps difficult to identify.
We in..."
via Arxivπ€ Hangjie Yuan, Yichen Qian, Zhiwei Tang et al.π 2026-07-27
β‘ Score: 6.2
"Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric challenge: models must absorb knowledge from heterogeneous 2D and 3D medical images, and evaluation protocols must align with..."
via Arxivπ€ Bingnan Li, Haozhe Wang, Haozhong Xiong et al.π 2026-07-27
β‘ Score: 6.2
"On-policy distillation (OPD) adapts diffusion models by querying a teacher along trajectories generated by the current student, but how it should behave under classifier-free guidance (CFG), a default component of modern diffusion systems, remains poorly understood. Existing OPD methods naturally ex..."
via Arxivπ€ Justin Sirignano, Konstantinos Spiliopoulos, Samuel Cohenπ 2026-07-27
β‘ Score: 6.1
"The Deep Galerkin Method (DGM) and Physics Informed Neural Networks (PINNs) have become widely-used methods for solving partial differential equations (PDEs) in the rapidly growing field of scientific machine learning. In these methods, a neural network is trained to approximate the PDE solution by..."
via Arxivπ€ Atharva Pandey, Gautam Jajooπ 2026-07-27
β‘ Score: 6.1
"Large language models are increasingly used as social simulators, including as synthetic survey respondents. Most evaluations ask whether simulated outcomes resemble human outcomes. We argue that this is necessary but too weak: a simulator can match the final answer while using the wrong rationale-d..."
via Arxivπ€ Nanbeige Lab, :, Chen Yang et al.π 2026-07-24
β‘ Score: 6.1
"We present Nanbeige4.2-3B, a compact general agentic model with 3B non-embedding parameters. It delivers strong performance across code-agent, office-agent, and complex tool-use tasks while maintaining highly competitive reasoning capabilities in mathematics, coding, and science. Nanbeige4.2-3B is p..."
via Arxivπ€ Varun Ghat Ravikumar, Sina Ahmadi, Lena JΓ€ger et al.π 2026-07-24
β‘ Score: 6.1
"Most endangered languages lack the parallel data required for machine translation, despite the existence of descriptive grammar books. We introduce a pipeline that uses large language models to extract grammatical rules, example sentences, and lexicons from grammar books and generate synthetic paral..."
ποΈ FROM THE ARCHIVE
Recent daily Metamesh snapshots with preserved AI news rankings, clusters, source links,
and ticker commentary.
Anthropic and OpenAI race to ship frontier models while quietly lobbying Washington to restrict open-weight competitors. The alignment problem worth watching is between their press releases and their policy positions.