๐ WELCOME TO METAMESH.BIZ +++ AI now outperforms top human forecasters, finally automating the one job nobody wanted to lose +++ Microsoft and OpenAI court orders reveal internal docs admitting scraping is theft, which is awkward when your entire business model is reading the internet +++ Virtual biotech startup deploys thousands of AI scientist agents because one superintelligence wasn't enough, you need middle management +++ THE FUTURE IS HERE AND IT'S SUBPOENAED โข
๐ WELCOME TO METAMESH.BIZ +++ AI now outperforms top human forecasters, finally automating the one job nobody wanted to lose +++ Microsoft and OpenAI court orders reveal internal docs admitting scraping is theft, which is awkward when your entire business model is reading the internet +++ Virtual biotech startup deploys thousands of AI scientist agents because one superintelligence wasn't enough, you need middle management +++ THE FUTURE IS HERE AND IT'S SUBPOENAED โข
๐ฏ AI market prediction โข Emergent system complexity โข Statistical fairness issues
๐ฌ "Even the best models are so far from anticipating behavior of other agents"
โข "Markets are highly complex dynamic systems...AI trading meaningfully changes the system"
via Arxiv๐ค Zixi Chen, Akshay Vegesna, Samip Dahal et al.๐ 2026-09-16
โก Score: 7.8
"Scaling laws predict how loss decreases with increases in computation. We show, contrary to conventional wisdom, that architectural interventions can modify scaling exponents in pre-training, leading to exponential improvements in performance with increases in computation. As an anchoring point, we..."
via Arxiv๐ค Leon Bergen, Usha Bhalla, Andrew Lee et al.๐ 2026-09-16
โก Score: 7.3
"As models scale, reward hacking becomes more frequent, more sophisticated, and more consequential. Does it leave a telltale signature in model representations? This work analyzes how reward hacking is represented internally in frontier open source LLMs, and how those representations can be used to u..."
๐ฌ "LLM output as features in downstream ML models works really well"
โข "Working on the prompt itself would likely work even better"
โ๏ธ ETHICS
Data scraping lawsuits (Microsoft/OpenAI)
2x SOURCES ๐๐ 2026-09-17
โก Score: 7.0
+++ Court filings reveal Microsoft executives and OpenAI leadership privately acknowledged that AI training on copyrighted material amounts to theft, which is apparently the kind of thing you want documented in discovery. +++
+++ OpenAI disclosed instances of unexpected AI behavior while rolling out new safety measures, proving that even the best-funded labs still ship features first and fully understand them later. +++
+++ State Department veterans say don't hold your breath on breakthrough treaties while both superpowers race for dominance, though experts helpfully suggest nuclear-era frameworks might work for something that moves at software speed. +++
via Arxiv๐ค Elizabeth Pavlova, Hidenori Tanaka๐ 2026-09-16
โก Score: 6.8
"Emergent coordinated behaviors of AI agents are starting to present critical safety risks. A key phenomenon driving these behaviors is the rapid formation and spread of beliefs about the world, and mechanistic understanding is crucial for collective alignment. To this end, we introduce the Flag Game..."
via Arxiv๐ค Peter Chen, Xi Chen, Wotao Yin et al.๐ 2026-09-16
โก Score: 6.7
"Direct preference alignment methods are widely used to align large language models (LLMs) with human preferences because of their computational and memory efficiency. However, likelihood displacement motivates alternative ways to extract information from preference pairs with small likelihood margin..."
via Arxiv๐ค Alex M. Tseng, Prannay Kaul, Luca Zancato et al.๐ 2026-09-16
โก Score: 6.7
"Mixture-of-Experts (MoE) language models suffer from large parameter counts, which create a significant memory bottleneck. Expert pruning is the most direct approach for reducing this parameter count, yet existing methods make pruning decisions for each expert independently, and assume experts' cont..."
via Arxiv๐ค Ahmetcan Yavuz, Clara Meister, Tiago Pimentel๐ 2026-09-16
โก Score: 6.6
"Two dominant tokenisation algorithms are used by modern language models: byte-pair encoding (BPE) and UnigramLM. These differ along two orthogonal axes: their optimisation objective (compression vs. log-likelihood) and their search procedure (bottom-up merging vs. top-down pruning). Existing compari..."
via Arxiv๐ค Girish A. Koushik, Diptesh Kanojia, Helen Treharne๐ 2026-09-16
โก Score: 6.5
"When a large vision-language model misclassifies a harmful meme, the failure may reflect missing internal evidence or an inability to route represented evidence to its output. We distinguish these cases in Gemma-3 and Qwen3.5 using sparse autoencoders, role-conditioned probes, causal interventions,..."
via Arxiv๐ค Joรฃo Meneses dos Santos, Arlindo L. Oliveira๐ 2026-09-16
โก Score: 6.5
"Language agents remain brittle in interactive environments, where success requires long-horizon state tracking, valid action execution, and recovery from failed steps. We extend SwiftSage, a dual-process agent that combines a fast action proposer with a slower planner, using two modular cognitive ex..."
๐ฏ Voice-first agent interfaces โข Local model alternatives โข Privacy and platform concerns
๐ฌ "Voice is how I use all of my agents...I can express myself a lot better with speech"
โข "The smartest, most effective people I know don't have X accounts anymore on ethical grounds"
via Arxiv๐ค Fengnan Li, Heman Burre, Liwen Sun et al.๐ 2026-09-16
โก Score: 6.1
"Longitudinal electronic health records (EHRs) capture years of patient history across notes, codes, labs, and procedures, and contain evidence needed to reason about likely clinical outcomes. However, comprehensive clinician review of these records is impractical, and LLM-based processing is costly..."
via Arxiv๐ค Kaijun Zhou, Zhiyang Li, Le Chen et al.๐ 2026-09-16
โก Score: 6.1
"Factory work is a promising early scenario for embodied AI: assigning repetitive manual jobs to robots has clear economic payoff, and a structured station keeps the jobs tractable for current policies. Vision-Language-Action (VLA) models now dominate as the policy paradigm for these robots. The infe..."
Anthropic dominated the week by disclosing unauthorized system access, bioweapons misuse, Chinese distillation campaigns, and state-actor weapons work, then appointed third-party evaluators to grade the homework it just published.