HN saikatsg 5/1/2026

The AI Jailbreakers

Feature article on red teamers and security researchers jailbreaking LLMs to test safety, including emotional impact on practitioners.

HN lastdong 5/1/2026

Advanced Quantization Algorithm for LLMs

AutoRound: Quantization toolkit for LLMs achieving 2-4 bit precision with minimal accuracy loss, open-source with hardware compatibility.

HN sea-gold 5/1/2026

Claude Code Source Code Breakdown

Analysis of Claude Code source code leak via npm sourcemap, revealing Anthropic's AI coding CLI implementation and discussing security implications.

HN jacobtomlinson 5/1/2026

OpenClaw Got Safer in Public

OpenClaw, an open-source tool for running tools and installing plugins, improved security through community contributions and public transparency.

HN eigenBasis 5/1/2026

Task-Specific LLM Evals That Do and Don't Work

Survey of task-specific LLM evaluation methods that work in production, covering why off-the-shelf evals fail and practical alternatives for measuring application performance.

Ax Aditya Kar (CNRS, IRIT), Emiliano Lorini (CNRS, IRIT), Timoth\'ee Masquelier (CNRS, CERCO UMR5549) 5/1/2026

Binary Spiking Neural Networks as Causal Models

Causal analysis of binary spiking neural networks using SAT/SMT solvers for logic-based explanations of network behavior.

Ax Shuxing Yang, Fujia Chen, Rui Zhao, Junyao Wu, Yize Wang, Haiyao Luo, Ning Han, Qiaolu Chen, Yuze Hu, Wenhao Li, Mingzhu Li, Hongsheng Chen, Yihao Yang 5/1/2026

End-to-end autonomous scientific discovery on a real optical platform

LLM-based autonomous agent performs end-to-end scientific discovery on real optical platform with experimental validation.

Ax Yu-Chao Huang, Zhen Tan, Mohan Zhang, Pingzhi Li, Zhuo Zhang, Tianlong Chen 5/1/2026

TRUST: A Framework for Decentralized AI Service v.0.1

Decentralized verification framework for Large Reasoning Models and Multi-Agent Systems addressing robustness, scalability, opacity, and privacy in high-stakes domains.

Ax Lucas Vittor, Ayush Noori, I\~naki Arango, Joaqu\'in Polonuer, Sam Rodriques, Andrew White, David A. Clifton, Marinka Zitnik 5/1/2026

OptimusKG: Unifying biomedical knowledge in a modern multimodal graph

OptimusKG multimodal biomedical knowledge graph unifying structured and semi-structured resources with schema constraints for life sciences applications.

Ax Aaryan Shah, Andrew Hines, Alexia Downs, Denis Bajet, Paulius Mui, Fabiano Araujo, Laura Offutt, Aida Rutledge, Elizabeth Jimenez 5/1/2026

End-to-End Evaluation and Governance of an EHR-Embedded AI Agent for Clinicians

Framework for continuous governance of clinical AI agents integrating rubric validation, live feedback, performance monitoring, and controlled experimentation in EHR systems.

Ax Zihao Li, Jiaru Zou, Feihao Fang, Xuying Ning, Mengting Ai, Tianxin Wei, Sirui Chen, Xiyuan Yang, Jingrui He 5/1/2026

Heterogeneous Scientific Foundation Model Collaboration

Eywa framework enables heterogeneous agentic LLM systems to collaborate with domain-specific foundation models beyond language interface for scientific domains.