LB arkaung.github.io via yelianung 4/27/2026

TurboQuant: A First-Principles Walkthrough

TurboQuant: quantization method compressing KV caches and embeddings to 2-4 bits using random rotations with provably near-optimal distortion, no training required.

HN RamtinJ95 4/27/2026

State of my AI-assisted development workflows

Long-form technical article documenting author's AI-assisted development workflows and tools over 3 months, comparing setup changes and tool integration patterns.

HN danboarder 4/27/2026

MemPalace – Local-first AI memory

MemPalace is local-first AI memory system storing verbatim conversation history with semantic search retrieval, achieving 96.6% R@5 on LongMemEval without API calls.

Ax Weitao Li, Boran Xiang, Xiaolong Wang, Zhinan Gou, Weizhi Ma, Yang Liu 4/27/2026

UR$^2$: Unify RAG and Reasoning through Reinforcement Learning

UR² unifies retrieval-augmented generation and reasoning via reinforcement learning, enabling LLMs to learn when and how to retrieve and reason across diverse domains.

Ax Zahra Yousefijamarani, Xinglu Wang, Qian Wang, Morgan Lindsay Heisler, Taha Shabani, Niloofar Gholipour, Parham Yassini, Hong Chang, Kan Chen, Qiantao Zhang, Xiaolong Bai, Jiannan Wang, Ying Xiong, Yong Zhang, Zhenan Fan 4/27/2026

HFX: Joint Design of Algorithms and Systems for Multi-SLO Serving and Fast Scaling

HFX system jointly optimizes algorithms and infrastructure for multi-task LLM serving under service-level objectives and dynamic workloads with elastic scaling.

Ax Xingyu Shen, Yingfa Chen, Zhen Leng Thai, Xu Han, Zhiyuan Liu, Maosong Sun 4/27/2026

StateX: Enhancing RNN Recall via Post-training State Expansion

StateX improves RNN recall on long contexts by expanding fixed-size recurrent states post-training, addressing compression limitations in linear attention and state-space models.

Ax Matthew Sharp, Omer Bilgin, Iason Gabriel, Lewis Hammond 4/27/2026

Agentic Inequality

Examines inequality issues arising from unequal access to and capabilities of autonomous AI agents in political and economic systems.

Ax Christoph B\"uhler, Matteo Biagiola, Luca Di Grazia, Guido Salvaneschi 4/27/2026

AgentBound: Securing Execution Boundaries of AI Agents

AgentBound introduces the first access control framework for securing AI agents using Model Context Protocol, preventing unrestricted system access.

Ax Rebonto Haque, Oliver M. Turnbull, Anisha Parsan, Nithin Parsan, John J. Yang, Anna L. Beukenhorst, Charlotte M. Deane 4/27/2026

Mechanistic Interpretability of Antibody Language Models Using SAEs

Mechanistic interpretability study using sparse autoencoders to understand and steer antibody language models through latent feature analysis.

Ax Kaibo Huang, Jin Tan, Yukun Wei, Wanling Li, Zipei Zhang, Hui Tian, Zhongliang Yang, Linna Zhou 4/27/2026

AgentMark: Utility-Preserving Behavioral Watermarking for Agents

AgentMark proposes utility-preserving watermarking techniques for LLM-based agents to protect IP and identify planning behaviors in multi-step task execution.

Ax Deming Chen, Vijay Ganesh, Weikai Li, Yingyan Celine Lin, Yong Liu, Subhasish Mitra, David Z. Pan, Ruchir Puri, Jason Cong, Yizhou Sun 4/27/2026

Report for NSF Workshop on AI for Electronic Design Automation

NSF workshop report on AI applications in Electronic Design Automation covering LLMs, GNNs, RL, and neurosymbolic methods.

Ax Yixiang Fang, Arijit Khan, Tianxing Wu, Da Yan, Shu Wang 4/27/2026

LLM+Graph@VLDB'2025 Workshop Summary

Workshop summary on integrating LLMs with graph data, covering algorithms and systems bridging LLMs, graphs, and machine learning.

Ax Francesco Andrea Causio, Vittorio De Vita, Olivia Riccomi, Michele Ferramola, Federico Felizzi, Alessandro Tosi, Antonio Cristiano, Lorenzo De Mori, Chiara Battipaglia, Melissa Sawaya, Luigi De Angelis, Marcello Di Pumpo, Alessandra Piscitelli, Pietro Eric Risuleo, Alessia Longo, Giulia Vojvodic, Mariapia Vassalli, Bianca Destro Castaniti, Nicol\`o Scarsi, Manuel Del Medico 4/27/2026

EuropeMedQA Study Protocol: A Multilingual, Multimodal Medical Examination Dataset for Language Model Evaluation

EuropeMedQA dataset evaluates LLM performance on multilingual multimodal medical exams from European regulatory sources.

Ax Martin Balko, Jan Greb\'ik, Pavel Hub\'a\v{c}ek, Martin Kouteck\'y, Mat\v{e}j Kripner, V\'aclav Rozho\v{n}, Robert \v{S}\'amal, Adri\'an Z\'ame\v{c}n\'ik 4/27/2026

Bolzano: Case Studies in LLM-Assisted Mathematical Research

Bolzano orchestrates parallel LLM prover and verifier agents with persistent knowledge base to produce novel mathematical and CS results.

Ax Sravanth Kodavanti, Sowmya Vajrala, Srinivas Miriyala, Utsav Tiwari, Uttam Kumar, Utkarsh Kumar Mahawar, Achal Pratap Singh, Arya D, Narendra Mutyala, Vikram Nelvoy Rajendiran, Sharan Kumar Allur, Euntaik Lee, Dohyoung Kim, HyeonSu Lee, Gyusung Cho, JungBae Kim 4/27/2026

Unlocking the Edge deployment and ondevice acceleration of multi-LoRA enabled one-for-all foundational LLM

Framework for on-device LLaMA inference on smartphones with multi-LoRA support, achieving edge deployment on Qualcomm chipsets.

Ax Zhaokun Wang, Jinyu Guo, Jingwen Pu, Hongli Pu, Meng Yang, Xunlei Chen, Jie Ou, Wenyi Li, Guangchun Luo, Wenhong Tian 4/27/2026

CAP: Controllable Alignment Prompting for Unlearning in LLMs

CAP enables selective knowledge unlearning in LLMs through controllable prompting without model weight access for regulatory compliance.

Ax Mohamed Ali Souibgui, Jan Fostier, Rodrigo Abad\'ia-Heredia, Bohdan Denysenko, Christian Marschke, Igor Peric 4/27/2026

LayerBoost: Layer-Aware Attention Reduction for Efficient LLMs

LayerBoost reduces attention complexity in transformers by applying layer-aware modifications instead of uniform replacement across all layers.