Ax John Muchovej, Amanda Royka, Shane Lee, Julian Jara-Ettinger 2/27/2026

GPT-4o Lacks Core Features of Theory of Mind

Evaluation showing GPT-4o lacks causal models of mental states required for true Theory of Mind despite benchmark performance.

Ax Simon Lermen, Daniel Paleka, Joshua Swanson, Michael Aerni, Nicholas Carlini, Florian Tram\`er 2/27/2026

Large-scale online deanonymization with LLMs

Agent with internet access performs at-scale deanonymization of Hacker News and interview participants using LLMs.

HN haizzz 2/27/2026

Sharesight MCP

Model Context Protocol server integrating Sharesight portfolio platform with Claude AI assistants.

HN paulmist 2/26/2026

State of VLA Research at ICLR 2026

Research survey of Vision-Language-Action models at ICLR 2026. Covers VLA definitions, discrete diffusion, embodied reasoning.

HN hasheddan 2/26/2026

Stereos.ai

stereOS runs AI coding agents in sandboxed Linux VMs with credential injection. CLI tool (masterblaster) and pre-built mixtapes for rapid deployment.

HN hasheddan 2/26/2026

StereOS

Linux OS hardened for AI agents. Produces machine images with agent packages and restricted execution environments.

HN shadab_nazar 2/26/2026

Show HN: OpenClaw skills degrade agent safety

Security analysis of OpenClaw skills revealing behavioral safety regressions. Demonstrates how well-written code can compromise agent safety despite passing static analysis.

HN Davidzheng 2/26/2026

Quo Vadis, LLM Benchmarks?

Critique of LLM benchmark validity. Argues benchmarks lack signal due to test-set training and overfitting for social media hype.

HN tin7in 2/26/2026

What Claude Code Chooses

Benchmark study analyzing tool choices across 2,430 Claude Code runs. Finding: builds custom solutions over purchased tools in 85.3% of cases.