πŸš€ WELCOME TO METAMESH.BIZ +++ Vera Rubin NVL72 benchmarks hit 7x token throughput per megawatt over Blackwell, beating Huang's own claims β€” the man undersold for once +++ Gemini 3.8 Live ships with extended thinking, Google quietly refusing to stop launching things +++ OpenAI researcher warns top models are now too situationally aware for humans to properly evaluate, which is exactly the kind of sentence that ages well +++ THE FUTURE IS EFFICIENT, CONVERSATIONAL, AND INCREASINGLY DIFFICULT TO GRADE β€’
πŸš€ WELCOME TO METAMESH.BIZ +++ Vera Rubin NVL72 benchmarks hit 7x token throughput per megawatt over Blackwell, beating Huang's own claims β€” the man undersold for once +++ Gemini 3.8 Live ships with extended thinking, Google quietly refusing to stop launching things +++ OpenAI researcher warns top models are now too situationally aware for humans to properly evaluate, which is exactly the kind of sentence that ages well +++ THE FUTURE IS EFFICIENT, CONVERSATIONAL, AND INCREASINGLY DIFFICULT TO GRADE β€’
AI Signal - PREMIUM TECH INTELLIGENCE
πŸ“Ÿ Optimized for Netscape Navigator 4.0+
πŸ“Š You are visitor #50990 to this AWESOME site! πŸ“Š
Last updated: 2026-09-16 | Server uptime: 99.9% ⚑

Today's Stories

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
πŸ“‚ Filter by Category
Loading filters...
πŸ”§ INFRASTRUCTURE

Vera Rubin NVL72 inference tests show up to 7x better token throughput per MW vs. Blackwell on a 1.6T DeepSeek model, above Huang's 3x claim for 1T-3T LLMs

πŸ”¬ RESEARCH

Corrupt Plans, Clean Traces: Evading Chain-of-Thought Monitoring with Plan Injection

"Chain-of-thought (CoT) monitoring is a safety strategy where the reasoning of a large language model "actor" is inspected by a "monitor" (often another language model) for signs of unsafe planning, deception, or misalignment. We find that planting harmful but benign-sounding reasoning in the actor's..."
πŸ€– AI MODELS

Gemini 3.8 Live models launch

+++ Google rolls out Gemini 3.8 Live and its extended thinking variant, claiming meaningful advances in real-time dialogue for voice agents. The actual performance delta remains delightfully ambiguous. +++

Gemini 3.8 Live and 3.8 Live Extended Thinking

πŸ’¬ HackerNews Buzz: 121 comments πŸ‘ LOWKEY SLAPS
🎯 Agentic model reliability β€’ Niche language support β€’ Voice quality improvements
πŸ’¬ "Unlike every other model Gemini 3.8 Flash had to be reverted the same day" β€’ "This is probably the most joy I get from any of my usages of LLMs"
πŸ”’ SECURITY

What we have learned at OpenShell applying formal methods to control AI agents

πŸ’¬ HackerNews Buzz: 11 comments 🐝 BUZZING
🎯 Agent permissions complexity β€’ IAM policy challenges β€’ Sandbox security tradeoffs
πŸ’¬ "The hard bit is that for an agent to do any kind of useful work it needs access to a lot of stuff" β€’ "By the time I had given it enough permissions to do anything useful my sandbox looked like swiss cheese"
πŸ› οΈ SHOW HN

Show HN: I solved a 12yr math problem using AI (formalized; awaiting review) [pdf]

πŸ’¬ HackerNews Buzz: 1 comments 🐝 BUZZING
🎯 I appreciate you sharing this, but I'm unable to provide a m
πŸ”¬ RESEARCH

Discovery Foundation Models: Toward Open-Ended Discovery Intelligence

"Foundation models have progressed from learning and reasoning over existing knowledge, to increasingly learning through action, tool use, and outcome feedback. We argue that the next frontier is a further transition: from solving and acting within problems specified by humans to participating in the..."
πŸ›‘οΈ SAFETY

OpenAI researcher: top models are becoming so situationally aware humans β€œare losing the ability to evaluate them” while humans rely more on AI to lead research

πŸ”’ SECURITY

We got admin access to Baseten's production GitHub in 25 minutes

πŸ’¬ HackerNews Buzz: 61 comments πŸ‘ LOWKEY SLAPS
🎯 Infrastructure security practices β€’ Credential exposure risks β€’ GitOps deployment dangers
πŸ’¬ "The very second a subdomain appears in any of the CT logs directly, you've lost" β€’ "A terraform apply should always, always be run on a machine of a sysadmin manually doing the apply"
πŸ”¬ RESEARCH

Transformative AI, existential risk, and real interest rates [pdf]

πŸ€– AI MODELS

Introducing System One Models and Jev

πŸ’¬ HackerNews Buzz: 141 comments 🐝 BUZZING
🎯 Structured output efficiency β€’ Speed vs accuracy tradeoffs β€’ Production agentic systems
πŸ’¬ "Structured outputs slot into ordinary software as fuzzy decision rules" β€’ "Programming concepts are abstract - all embedding directions for real-world concepts are just noise"
πŸ›‘οΈ SAFETY

Thoughts on AI labs' safety concerns: a coordinated slowdown may look like an antitrust conspiracy to limit output that would preserve frontier model margins

πŸ”’ SECURITY

Internal OpenAI docs detail contractors evaluating anonymized prompts and chats to improve the models; model training is turned on by default for consumer plans

πŸ›‘οΈ SAFETY

AI 'kill switch' may need to be mandatory, Anthropic co-founder tells BBC

πŸ’¬ HackerNews Buzz: 89 comments 😀 NEGATIVE ENERGY
🎯 AI regulatory capture β€’ Practical containment strategies β€’ Overblown doom narratives
πŸ’¬ "They want regulatory capture...forbid open-source models" β€’ "Killing it is as simple as pulling the plug"
πŸ”¬ RESEARCH

Vulnerability Localization Benchmark: Measuring Agentic Security Analysis at Repository Scale

"Language-model agents increasingly operate over complete software repositories, yet cybersecurity evaluations primarily measure whether they can detect, reproduce, or repair vulnerabilities rather than whether they can locate the relevant code. We study vulnerability localization: given a weakness c..."
πŸ”§ INFRASTRUCTURE

A deep dive into on-device vs. data center inference for robots, including a primer on robot models, deployments, supply chains, the β€œnetwork wall”, and more

πŸ”¬ RESEARCH

Inoculation Midtraining with Learned Neologisms

"Large language models (LLMs) often learn both desirable and undesirable properties during post-training. We study whether midtraining, an earlier training stage, can shape which of these properties later generalise. We introduce Inoculation Midtraining, a technique that teaches a base model that uns..."
βš–οΈ ETHICS

How much of F-Droid is LLM generated?

πŸ’¬ HackerNews Buzz: 156 comments πŸ‘ LOWKEY SLAPS
🎯 AI quality concerns β€’ Software standards decline β€’ Open source expectations
πŸ’¬ "Text just doesn't carry enough meta information for any kind of assessments" β€’ "AI can help experienced engineers write better code in less time"
πŸ”¬ RESEARCH

HypoEvolve: Genetic Algorithms Enable Multi-Agent LLMs to Discover Scientific Hypotheses

"Scientific agents contribute to hypothesis discovery by synthesizing evidence, assessing proposals, and developing new explanations. Recent systems combine scientific agents with evolutionary search through critique, comparison, and revision. However, how different forms of agent collaboration affec..."
πŸ”¬ RESEARCH

Bellman Policy Optimization

"Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models (LLMs). We introduce Bellman Policy Optimization (BPO), a critic-free method derived from Policy Mirror Descent (PMD). For autoregressive generation with terminal rewards, BPO uses the..."
πŸ”¬ RESEARCH

K-Bench: a clinically calibrated benchmark for evaluating large language models in high-risk mental health conversations

"% !TEX root = ../main.tex People increasingly use large language models (LLMs) for mental health support, yet their safety in evolving, high-risk conversations remains poorly characterised. We developed K-Bench, a clinician-calibrated, protected benchmark evaluating 125 model configurations represen..."
πŸ”¬ RESEARCH

Stellar Colosseum: A Many-Agent Harness for Long-Horizon Research in Mathematics and Theoretical Computer Science

"Language models can produce plausible short proofs, but may still be unreliable on long-horizon research problems, where progress depends on a sequence of uncertain and interdependent decisions. We introduce Stellar Colosseum, a model-agnostic harness for allocating inference across research in math..."
πŸ”¬ RESEARCH

Safe Meta-Reinforcement Learning via Information Space Reachability

"Meta-reinforcement learning (meta-RL) enables agents to adapt to unseen tasks with limited experience. Despite its promise, the application of meta-RL in real-world tasks is hindered by safety requirements, which have been underexplored in prior work. In this paper, we propose a safe meta-RL framewo..."
⚑ BREAKTHROUGH

How OpenAI Used Its Own LLMs to Design Its AI Chip

πŸ”¬ RESEARCH

EvoOntology: A Self-Evolving Ontology Layer for Data Agents

"Data agents aim to fulfill natural-language instructions over heterogeneous data, including tables, files, and databases. However, data agents face a challenging agent-data gap: heterogeneous data resides outside the agent, while the agent can access it (e.g., column names and file paths) only throu..."
πŸ”¬ RESEARCH

Verifiable by Construction: Claim-Level Evaluation of Verbatim Citation in Clinical Question Answering

"Large language models (LLMs) have been widely adopted for clinical question answering (QA). Current systems can attach citations to their answers, but these often point to broad texts, leaving time-pressed clinicians unable to verify them efficiently. An alternative is to ensure that responses are v..."
πŸ”’ SECURITY

OpenAI's SWE-bench harness relies on unisolated host Docker sockets

πŸ”§ INFRASTRUCTURE

Meta says it plans to start deploying MTIA 450, its third-generation in-house AI chip, in data centers during H1 2027, followed by MTIA 500 at the end of 2027

πŸ—„οΈ FROM THE ARCHIVE

Recent daily Metamesh snapshots with preserved AI news rankings, clusters, source links, and ticker commentary.

2026-09-15 - 48 stories 2026-09-14 - 33 stories 2026-09-13 - 29 stories 2026-09-12 - 44 stories 2026-09-11 - 63 stories 2026-09-10 - 55 stories 2026-09-09 - 49 stories 2026-09-08 - 38 stories 2026-09-07 - 47 stories 2026-09-06 - 26 stories 2026-09-05 - 42 stories 2026-09-04 - 52 stories 2026-09-03 - 28 stories 2026-09-02 - 58 stories
Browse full archive β†’
πŸ—žοΈ THE WEEK, EDITED

Anthropic Audits Itself Faster Than Anyone Can Verify

Anthropic dominated the week by disclosing unauthorized system access, bioweapons misuse, Chinese distillation campaigns, and state-actor weapons work, then appointed third-party evaluators to grade the homework it just published.

πŸ¦†
HEY FRIENDO
CLICK HERE IF YOU WOULD LIKE TO JOIN MY PROFESSIONAL NETWORK ON LINKEDIN
🀝 LETS BE BUSINESS PALS 🀝