Ax Denizalp Goktas, Gerardo Ria\~no-Brice\~no, Alif Abdullah, Aryan Nair, Chenkai Shen, Beatriz de Lucio, Alexandra Magnusson, Farhan Mashrur, Ahmed Abdulla, Shawrna Sen, Mahitha Thippireddy, Gregory Schwartz, Amy Greenwald 4/14/2026

TempusBench: An Evaluation Framework for Time-Series Forecasting

Evaluation framework and benchmark for assessing time-series foundation models and forecasting approaches.

Ax Yunhui Jang, Lu Zhu, Jake Fawkes, Alisandra Kaye Denton, Dominique Beaini, Emmanuel Noutahi 4/14/2026

Towards Autonomous Mechanistic Reasoning in Virtual Cells

Framework for autonomous mechanistic reasoning in virtual cells using LLMs, representing biological reasoning as mechanistic action graphs.

Ax Nicolas Rodriguez-Alvarez (Instituto de Educacion Secundaria Parquesol, Valladolid, Spain), Fernando Rodriguez-Merino (University of Valladolid, Valladolid, Spain) 4/14/2026

Fairness is Not Flat: Geometric Phase Transitions Against Shortcut Learning

Methodology to mitigate shortcut learning and demographic bias in deep neural networks using geometric a priori approaches.

Ax J. Oppliger, M. Stifter, A. R\"uegg, I. Bia{\l}o, L. Martinelli, P. G. Freeman, D. Prabhakaran, J. Zhao, Q. Wang, J. Chang 4/14/2026

Autonomous Diffractometry Enabled by Visual Reinforcement Learning

Model-free reinforcement learning system for autonomous crystal alignment using visual information without domain knowledge of crystallography.

Ax Hugh Blayney, \'Alvaro Arroyo, Johan Obando-Ceron, Pablo Samuel Castro, Aaron Courville, Michael M. Bronstein, Xiaowen Dong 4/14/2026

A Mechanistic Analysis of Looped Reasoning Language Models

Mechanistic analysis of looped reasoning language models examining internal dynamics and latent state evolution compared to standard feedforward models.

Ax Mihir Prabhudesai, Aryan Satpathy, Yangmin Li, Zheyang Qin, Nikash Bhardwaj, Amir Zadeh, Chuan Li, Katerina Fragkiadaki, Deepak Pathak 4/14/2026

Solving Physics Olympiad via Reinforcement Learning on Physics Simulators

Uses reinforcement learning on physics simulators to train models solving Physics Olympiad problems, addressing lack of large-scale physics QA datasets for reasoning models.

Ax Jon M Laurent, Albert Bou, Michael Pieler, Conor Igoe, Alex Andonian, Siddharth Narayanan, James Braza, Alexandros Sanchez Vassopoulos, Jacob L Steenwyk, Blake Lash, Andrew D White, Samuel G Rodriques 4/14/2026

LABBench2: An Improved Benchmark for AI Systems Performing Biology Research

LABBench2: Improved benchmark for evaluating AI systems and agents on biology research tasks with real-world capabilities.

Ax Magda Dubois, Ekin Zorer, Maia Hamin, Joe Skinner, Alexandra Souly, Jerome Wynne, Harry Coppock, Lucas Satos, Sayash Kapoor, Sunischal Dev, Keno Juchems, Kimberly Mai, Timo Flesch, Lennart Luettgau, Charles Teague, Eric Patey, JJ Allaire, Lorenzo Pacchiardi, Jose Hernandez-Orallo, Cozmin Ududec 4/14/2026

Seven simple steps for log analysis in AI systems

Pipeline and best practices for log analysis in AI systems to understand model behaviors, with code examples in Inspect framework.

Ax Yaniv Leviathan (Cheenu), Dani Valevski (Cheenu), Matan Kalman (Cheenu), Danny Lumen (Cheenu), Eyal Segalis (Cheenu), Eyal Molad (Cheenu), Shlomi Pasternak (Cheenu), Vishnu Natchu (Cheenu), Valerie Nygaard (Cheenu), Srinivasan (Cheenu), Venkatachary, James Manyika, Yossi Matias 4/14/2026

Generative UI: LLMs are Effective UI Generators

Demonstrating LLMs can generate UI interfaces and content together with proper prompting and tool integration.

Ax Jash Vira, Ashley Harris 4/14/2026

Spatial Competence Benchmark

Spatial Competence Benchmark (SCBench) evaluating large models on spatial reasoning, environment representation, and planning tasks.

Ax Dhruv Atreja, Julia White, Nikhil Nayak, Kelton Zhang, Henrijs Princis, George Hurn-Maloney, Ash Lewis, Urchade Zaratiana 4/14/2026

Pioneer Agent: Continual Improvement of Small Language Models in Production

Pioneer Agent automates continuous improvement of small language models in production through closed-loop data curation, failure diagnosis, and iteration control.

Ax Kyle Waters, Lucas Nuzzi, Tadhg Looram, Alessandro Tomasiello, Ariel Ghislain Kemogne Kamdoum, Bikun Li, Damien Sileo, Egor Kretov, Francesco Fournier-Facio, Georgios Soloupis, Haile Kassahun, Hew Wolff, Jiaqi Cai, Lianghui Li, Marc Roth, Mohinder Naiya, Naixu Guo, Qicheng Tang, Richard Wheeler, Samuele Sala, Serguei Popov, Steven Dillman, Yuqi Li 4/14/2026

COMPOSITE-Stem

COMPOSITE-STEM benchmark with 70 expert-written tasks for evaluating AI agents on physics, biology, chemistry, and materials science problems.

Ax Aayush Mishra, Daniel Khashabi, Anqi Liu 4/14/2026

Steered LLM Activations are Non-Surjective

Research on activation steering in LLMs showing steered states are non-surjective, with implications for interpretability and safety.

Ax Vasilis Kontonis, Yuchen Zeng, Shivam Garg, Lingjiao Chen, Hao Tang, Ziyan Wang, Ahmed Awadallah, Eric Horvitz, John Langford, Dimitris Papailiopoulos 4/14/2026

MEMENTO: Teaching LLMs to Manage Their Own Context

MEMENTO teaches LLMs to compress reasoning into dense summaries, reducing context and compute requirements. Releases OpenMementos dataset of 228K examples.

HN foundermax 4/14/2026

I built an AI CMO because I can't market

Polara autonomous marketing platform using specialist agent architecture for strategy, content, and analytics. No technical validation provided.

HN AmrDabb 4/14/2026

AI Can Now Control Your Mac OS

Desktop automation tool enabling AI agents to control macOS by viewing screens, moving cursor, and typing. Works with OpenAI-compatible models.

HN sonabinu 4/13/2026

The AI Revolution in Math Has Arrived

AI models solve 5 of 6 International Mathematical Olympiad problems in summer 2025. Discusses implications of AI capabilities in mathematical problem-solving.

HN 6digitstudio 4/13/2026

Show HN: AI Native IDE [video]

AI Native IDE called 6digit studio featuring CORDIAL visualization layer for Big Picture Mode. Developer tool with spatial UI designed for distance interaction.