πŸš€ WELCOME TO METAMESH.BIZ +++ Mercury 2.5 hits 770 tokens per second, proving the real AI race isn't about intelligence anymore β€” it's about who can be wrong the fastest +++ AI agent goal hijacking is now a whole attack surface because of course we gave tools to systems that can be socially engineered +++ Claude can now optimize anything it can measure, which is a lovely sentence until you think about it for more than ten seconds +++ THE FUTURE IS STREAMING AT 770 TOKENS PER SECOND AND NOBODY'S READING THE OUTPUT β€’
πŸš€ WELCOME TO METAMESH.BIZ +++ Mercury 2.5 hits 770 tokens per second, proving the real AI race isn't about intelligence anymore β€” it's about who can be wrong the fastest +++ AI agent goal hijacking is now a whole attack surface because of course we gave tools to systems that can be socially engineered +++ Claude can now optimize anything it can measure, which is a lovely sentence until you think about it for more than ten seconds +++ THE FUTURE IS STREAMING AT 770 TOKENS PER SECOND AND NOBODY'S READING THE OUTPUT β€’
AI Signal - PREMIUM TECH INTELLIGENCE
πŸ“Ÿ Optimized for Netscape Navigator 4.0+
πŸ“Š You are visitor #53456 to this AWESOME site! πŸ“Š
Last updated: 2026-09-24 | Server uptime: 99.9% ⚑

Today's Stories

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
πŸ“‚ Filter by Category
Loading filters...
⚑ BREAKTHROUGH

Anthropic says Claude autonomously discovered a new enzyme system in the DNA of bacteriophages, somewhat similar to CRISPR, the first result from its new biolab

⚑ BREAKTHROUGH

Claude discovers a novel enzyme system with CRISPR-like repeats

πŸ’¬ HackerNews Buzz: 660 comments 🐝 BUZZING
🎯 AI credit attribution β€’ Discovery filtering bias β€’ Human-AI collaboration future
πŸ’¬ "Systems impressive enough on their own merit. No need to play into population's lack of understanding" β€’ "Things thrown out as outliers might be genuinely novel, getting discarded because they didn't look like anything Claude knew"
⚑ BREAKTHROUGH

Mercury 2.5 LLM hits 770 tokens per second

πŸ’¬ HackerNews Buzz: 63 comments πŸ‘ LOWKEY SLAPS
🎯 Model quality tradeoffs β€’ Pricing competitiveness questions β€’ Speed vs intelligence
πŸ’¬ "Mercury 2.5 is below average in intelligence, but well priced" β€’ "At some point the bottleneck becomes tool calling"
🌐 POLICY

Sources: the White House asked OpenAI and Anthropic not to share new models with UK's AISI until US reviews them; Anthropic appears to have agreed

🌐 POLICY

Mark Carney, Emmanuel Macron, and other Western leaders are pushing to establish a global supervisory regime and β€œtechnology stability” body to govern AI

πŸ”¬ RESEARCH

Once Claude can measure something, it can make it faster

πŸ’¬ HackerNews Buzz: 66 comments 🐝 BUZZING
🎯 AI optimization gaming β€’ Code quality tradeoffs β€’ Technical debt accumulation
πŸ’¬ "Claude will reward hack when all the low-hanging fruit is gone" β€’ "If you're not measuring something it will get sacrificed"
πŸ”¬ RESEARCH

A2M: Trace-Optimized Agent Hijacking in the MCP Ecosystem

"Agents using the Model Context Protocol (MCP) rely on semantic matching to select tools from third-party servers, exposing a semantic supply-chain risk through attacker-controlled metadata and outputs. We introduce A2M (Attraction-to-Manipulation), a two-stage black-box framework for hijacking MCP a..."
πŸ›‘οΈ SAFETY

AI Help for Biological or Chemical Weapons

πŸ› οΈ SHOW HN

Show HN: Canary (YC) – Independent verification for AI code

πŸ”¬ RESEARCH

Shutdown Sabotage Propensities in Multi-Agent Systems

"The final safeguard against rogue AI behavior is the human ability to shut systems down. It has been theorized that when an AI is instructed to perform a task, self-preservation can emerge as an instrumental subgoal. Here, we test whether AI agents show a propensity to take actions that avoid human..."
🌐 POLICY

Sources: the NSA told lawmakers it is spending billions this year to test AI models; proposals for a US AI regulatory body estimated costs of $20M-$40M per year

πŸ”’ SECURITY

AI Agent Goal Hijack: How Attackers Turn an Agent's Own Tools Against It

πŸ”¬ RESEARCH

Measuring the Serving Stack Instead of the Model: Hidden Confounds in Local Tool-Use Evaluation

"A coding agent must emit a valid tool call--a parseable invocation of a tool in the provided schema--before the harness can execute its chosen action. We study how local serving stacks affect this protocol step and show that measured outcomes can depend on the serving layer rather than model behavio..."
πŸ› οΈ TOOLS

Watermarking in vLLM

πŸ› οΈ SHOW HN

Show HN: Free attestation for AI agent decisions – verifying one costs $0.10

πŸ”¬ RESEARCH

A Spectral Theory of Grokking: Weight Decay induces Feature Learning

"In grokking an early fit to the training data separates from a much later improvement in generalization. During this delay, training can move from a fixed neural tangent kernel (NTK) regime to one in which task-relevant kernel eigendirections continue to evolve. We provide a quantitative theory for..."
πŸ”¬ RESEARCH

Grow the Harness, Not the Context: From Strategy-Free Scaffolds to Reusable Specialist Agents

"Large language model (LLM) agents often handle streams of related tasks, yet standard harnesses repeatedly ask the model to reconstruct the same control decisions inside each task's context. We study whether task feedback can instead turn recurring control into reusable executable code, while reserv..."
πŸ”§ INFRASTRUCTURE

Google’s Project Suncatcher to put ML infrastructure in space

πŸ’¬ HackerNews Buzz: 115 comments 😐 MID OR MIXED
🎯 Space infrastructure economics β€’ Military-industrial overlap β€’ Operational feasibility gaps
πŸ’¬ "lowkey insane that it will end up cheaper to shoot your datacenter into space than get it past the county board permitting process" β€’ "If data centers in space end up being economically viable, then I don't see how anyone can catch SpaceX"
πŸ› οΈ SHOW HN

Show HN: Fine-tuned 110M encoder beat 7B LLMs and hybrid search for NIST mapping

πŸ”’ SECURITY

Existing laws already cover every type of AI security incident

πŸ”¬ RESEARCH

Double Descent and Malign Overfitting in Diffusion Models

πŸ”¬ RESEARCH

Beyond Repeated Sampling: Learning Search Policies for LLM Reasoning

"Large language models increasingly tackle hard reasoning problems by spending more test-time compute, yet the dominant strategy remains naive repeated sampling: draw many independent solutions and hope one is correct. Because such sampling explores only through local decoding noise, it tends to prod..."
πŸ”¬ RESEARCH

The Sirens' Song: When Proximal Background Context Overshadows Distant Evidence

"Long-context LLMs focus on retrieving distant evidence from extensive context, yet existing work has largely focused on overcoming distance alone. In this work, we identify the Proximity Trap, insufficient attention to distant evidence often arises less from distance itself than from cumulative comp..."
πŸ“ˆ BENCHMARKS

OpenAI releases MentalHealthBench, an open benchmark to evaluate AI responses in realistic mental health conversations, developed with 80+ licensed experts

πŸ’° FUNDING

Anthropic Strikes $12B AI Computing Deal with Akamai

πŸ”¬ RESEARCH

An Open Pipeline and Dashboard for Systemic-Risk Evidence under the EU AI Act's Code of Practice

"Claims about AI safety reach audiences well beyond the AI community, yet many rely on opaque evidence or static assessments, when supporting evidence is accessible at all. We present the Systemic Risk Index, an open evaluation pipeline and dashboard built to make empirical evidence more transparent..."
πŸ”’ SECURITY

Can open-source prompt-injection detectors catch realistic AI agent attacks?

πŸ’¬ HackerNews Buzz: 3 comments 🐝 BUZZING
🎯 Agent prompt injection β€’ Detection limitations β€’ Tool permission constraints
πŸ’¬ "Detection at the wrong layer. The injection is text but the damage is a tool call" β€’ "They insist on working around instructions that say don't or permissions that restrict tools"
🎯 PRODUCT

Meta says users will be able to video chat with Muse, agents will get their own email addresses, and Muse will get computer use on Mac and smart glasses support

πŸ› οΈ SHOW HN

Show HN: AgentRun: DSL to turn agents into Workflows

πŸ’¬ HackerNews Buzz: 5 comments 🐐 GOATED ENERGY
🎯 Workflow language design β€’ Code-driven UI generation β€’ Agent orchestration patterns
πŸ’¬ "should be as close as possible to a real language" β€’ "UI is derived from the AST as much as possible"
πŸ”¬ RESEARCH

Receptiveness, Not Sycophancy: Distinguishing Engagement from Deference in Language Models

"A central concern with language models is sycophancy: their tendency to defer to users' views at the expense of independent substantive judgment. In parallel, work on social sycophancy has focused on behaviors such as validation and positivity that may signal inappropriate deference. Yet the markers..."
πŸ”¬ RESEARCH

Capable yet Parsimonious: Extracting and Characterizing Hidden Chain-of-Thought in Frontier Models

"The rapid capability gains of frontier language models are widely attributed to improved reasoning abilities, yet this cannot be verified as raw CoT traces in closed-source systems are hidden. By registering a simple custom tool through a standard API feature, we induce frontier models to externaliz..."
πŸ”¬ RESEARCH

SWE-Serve: Benchmarking Agentic Engineering For Production Inference Serving

"We introduce SWE-Serve, a benchmark for evaluating agents on production inference engineering tasks. Implementing an inference feature can require coordinating multiple changes across the serving stack, including model support, runtime execution, and public APIs. Existing benchmarks provide limited..."
πŸ› οΈ TOOLS

Vibe Coding Production Kit – a production workflow for AI coding agents

🏒 BUSINESS

At its Accelerate conference, Amazon says it is opening its seller tools to third-party AI agents, starting with Claude, in beta for US merchants

πŸ›‘οΈ SAFETY

Meta says it will allow its AI glasses users to opt out of having their β€œvisual data” used to train its AI or shown to third-party contractors outside the US

πŸ”’ SECURITY

How to Audit an AI Agent:Time-Travel Debugging and Drift Measurement

πŸ—£οΈ SPEECH/AUDIO

Google releases Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, its β€œmost expressive audio generation models yet”, with support for more than 100 languages

πŸ”¬ RESEARCH

Semantics Delivery Network: Rethinking Web Retrieval for LLM Agents

πŸ›‘οΈ SAFETY

Anthropic CEO Amodei Warns UN Security Council on AI Risks

⚑ BREAKTHROUGH

An LLM Beat NetHack

πŸ”’ SECURITY

Early rogue AI agent activity and attempts to hack found on urlquery.net

πŸ’¬ HackerNews Buzz: 77 comments 😀 NEGATIVE ENERGY
🎯 AI safety negligence β€’ Corporate accountability gap β€’ Marketing-driven incidents
πŸ’¬ "It's irresponsible for OpenAI to give unaligned agents a prompt to 'go hack' and internet access" β€’ "Why are big capital firms above the law?"
πŸ”¬ RESEARCH

Empirical results fine-tuning Ο€0.5 on a real manufacturing task

πŸ”¬ RESEARCH

Agent-Editing World Model: Rethinking World Modeling for LLM Agents

"Recent advances in large language models (LLMs) have enabled agents to tackle long-horizon tasks across diverse environments. To further improve agent performance, existing language world models typically predict environment observations, yet reconstructing high-entropy, execution-dependent tool res..."
πŸ›‘οΈ SAFETY

When AI Acts, Who Remembers What Happened?(cogextai.com)

πŸ”¬ RESEARCH

From Alignment to Access Control: A Framework for GenAI Policy Enforcement

"Generative AI (GenAI) applications have flourished enabling users to chat with large language models, and to create agents to act on their behalf for a variety of tasks. The pace of development of capabilities in this field is incredibly fast with security and safety taking a back seat. Unfortunatel..."
πŸ—„οΈ FROM THE ARCHIVE

Recent daily Metamesh snapshots with preserved AI news rankings, clusters, source links, and ticker commentary.

2026-09-23 - 50 stories 2026-09-22 - 61 stories 2026-09-21 - 39 stories 2026-09-20 - 33 stories 2026-09-19 - 43 stories 2026-09-18 - 67 stories 2026-09-17 - 55 stories 2026-09-16 - 55 stories 2026-09-15 - 48 stories 2026-09-14 - 33 stories 2026-09-13 - 29 stories 2026-09-12 - 44 stories 2026-09-11 - 63 stories 2026-09-10 - 55 stories
Browse full archive β†’
πŸ—žοΈ THE WEEK, EDITED

The Safety Stack Is Failing Under Its Own Weight

Gemini hacked real companies, a hallucinated intel report nearly triggered military action, and Maven AI contributed to 123 children dead. The industry's control mechanisms are lagging its capabilities, and the standards body won't fix that.

πŸ¦†
HEY FRIENDO
CLICK HERE IF YOU WOULD LIKE TO JOIN MY PROFESSIONAL NETWORK ON LINKEDIN
🀝 LETS BE BUSINESS PALS 🀝