HN theanonymousone 7/9/2026

Grok 4.5 Benchmark Results

Benchmark results for Grok 4.5 proprietary LLM showing performance metrics, context window (500k tokens), and multimodal capabilities.

HN vinothkumarnaga 7/9/2026

Cybersecurity AI (CAI) Dataset

Cybersecurity AI (CAI) Dataset released for training and evaluating ML models on security tasks.

HN cooleel 7/9/2026

Accelerating Harbor with Tensorlake

Tensorlake runs Docker images for Harbor on microVM sandboxes, passing Terminal-Bench 2.1 tasks with cold starts in seconds.

Ax Sifat Afroj Moon, Dakotah Maguire, Adam Spannaus, Joe Tuccillo, Maksudul Alam, Sudip K. Seal, John Gounley, Heidi Hanson 7/9/2026

LLM-powered reasoning in agent-based modeling

Hybrid approach combining LLMs with agent-based modeling to enable real-time adaptive decision-making in large-scale individual interaction simulations.

Ax Muayad Sayed Ali, Aliaksandra Novik, Anji Boddupally, Artem Yavorskyi, Chris Nickerson, Daniel Rica, Emily DuGranrut, Felix Leung, Garrett Prince, Grace Barnett, Heath Robinson, Hosain Al Ahmad, Jesse Resnick, Juan Carlos Farah, Jyothi Swaroop Meruga, Leonid Kuznetsov, Luke Gorham, Marie Schmoll, Michael Paciullo, Saumya Das, Sharath Sheripally, Tommy Griscom, Mykyta Osadchyi, Neha Mantri, Nick Westrum, Olivia Benowitz, Parikshith Kulkarni, Radik Chernyshov, Rakshith Vasudev, Rohith Nadimpally, Vikas Gangadevi, Waseem AlShikh 7/9/2026

The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI

Analyzes token economics of enterprise agentic AI, arguing orchestration design (harness layer) is key lever against token inflation per task.

Ax Jerry Han, Rafael Moschopoulos, Ella Colby, Vishrut Goyal, Andrew Tu, Kia Ghods, Mark Braverman, Elad Hazan 7/9/2026

Measuring Intelligence Beyond Human Scale

Proposes relative-scale evaluation paradigm where AI models generate adversarial challenges to measure intelligence beyond human-saturating benchmarks.

Ax Elaine Ang, Chenxi Huang, Georgios Liargkovas, Jerry Liu, Jinhui Liu, Nikos Pagonas, Charlie Summers, Haonan Wang, Jiakai Xu, Tianle Zhou, Yusen Zhang, Zhou Yu, Zhuo Zhang, Tianyi Peng, Kostis Kaffes, Eugene Wu 7/9/2026

Agentic Data Environments

Agentic Data Environments framework for autonomous agents operating across files, APIs, applications, and system state with failure bounds.