HN ZLStas 2/20/2026

Using classic dev books to guide AI agents?

Open-source project structuring software engineering principles into reusable skills for AI agents on code review and system design tasks. Explores practical agent capability abstraction.

HN beowa 2/20/2026

Show HN: Legal RAG Bench

Legal RAG Bench benchmark evaluates hallucinations, retrieval failures, and reasoning in legal RAG systems. Findings show embedding models critical, not generative models.

HN sidnarsipur 2/20/2026

The path to ubiquitous AI (17k tokens/sec)

Analysis of barriers to AI adoption: latency and cost. Discusses performance improvements (17k tokens/sec) enabling better human-AI collaboration for coding.

HN matt_d 2/20/2026

Proof Assistants in the Age of AI

Formal proof assistants increasingly matter as AI generates verified mathematics; collaboration environment for humans and AI.

HN kukla3 2/20/2026

Agentic AI and the Mythical Agent Month

Position paper examining whether AI agents can overcome Brooks' Law through scalable agency, exploring theoretical advantages of instantaneous context loading.

Ax Mubashara Akhtar, Anka Reuel, Prajna Soni, Sanchit Ahuja, Pawan Sasanka Ammanamanchi, Ruchit Rawal, Vil\'em Zouhar, Srishti Yadav, Chenxi Whitehouse, Dayeon Ki, Jennifer Mickel, Leshem Choshen, Marek \v{S}uppa, Jan Batzner, Jenny Chim, Jeba Sania, Yanan Long, Hossein A. Rahmani, Christina Knight, Yiyang Nan, Jyoutir Raj, Yu Fan, Shubham Singh, Subramanyam Sahoo, Eliya Habba, Usman Gohar, Siddhesh Pawar, Robert Scholz, Arjun Subramonian, Jingwei Ni, Mykel Kochenderfer, Sanmi Koyejo, Mrinmaya Sachan, Stella Biderman, Zeerak Talat, Avijit Ghosh, Irene Solaiman 2/20/2026

When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

Systematic analysis of benchmark saturation across 60 LLM benchmarks, showing many quickly lose ability to differentiate best-performing models.