Ax Anirudh Ajith, Amanpreet Singh, Jay DeYoung, Nadav Kunievsky, Austin C. Kozlowski, Oyvind Tafjord, James Evans, Daniel S. Weld, Tom Hope, Doug Downey 2/25/2026

PreScience: A Benchmark for Forecasting Scientific Contributions

PreScience benchmark for forecasting scientific advances using AI. Tests if models trained on historical research can predict future collaborators, impactful directions, and emerging problems.

Ax Haotian Si, Changhua Pei, Xiao He, Zeyan Li, Zhe Xie, Zexin Wang, Jiyao Hu, Zhaoyang Yu, Tieying Zhang, Dan Pei, Jianhui Li, Gaogang Xie 2/25/2026

KairosVL: Orchestrating Time Series and Semantics for Unified Reasoning

KairosVL integrates time series analysis with semantic reasoning using reinforcement learning for complex temporal prediction tasks combining numerical and contextual understanding.

Ax Hongbin Zhong, Fazle Faisal, Luis Fran\c{c}a, Tanakorn Leesatapornwongsa, Adriana Szekeres, Kexin Rong, Suman Nath 2/25/2026

ActionEngine: From Reactive to Programmatic GUI Agents via State Machine Memory

ActionEngine framework enables GUI agents to operate programmatically via state machine memory instead of reactive step-by-step vision-language model calls, reducing costs and improving accuracy.

Ax Bo Zhang, Jinfeng Zhou, Yuxuan Chen, Jianing Yin, Minlie Huang, Hongning Wang 2/25/2026

Grounding LLMs in Scientific Discovery via Embodied Actions

EmbodiedAct framework grounds LLM scientific reasoning in physical simulation with runtime perception to detect anomalies and enable interactive discovery.

Ax Vaidehi Bagaria, Bijo Sebastian, Nirav Patel 2/25/2026

Recursive Belief Vision Language Model

Recursive Belief Vision Language Model (RBVLM) maintains hidden state for long-horizon manipulation under partial observability, reducing action repetition and inference latency.

Ax Julien Dallot, Yuval Emek, Yuval Gil, Maciej Pacut, Stefan Schmid 2/25/2026

Online Algorithms with Unreliable Guidance

Introduces online algorithms with unreliable guidance (OAG) model for ML-augmented online decision making that separates predictive and algorithmic components.

Ax Shitian Zhao, Shaoheng Lin, Ming Li, Haoquan Zhang, Wenshuo Peng, Kaipeng Zhang, Chen Wei 2/25/2026

PyVision-RL: Forging Open Agentic Vision Models via RL

PyVision-RL is an open-source reinforcement learning framework for training multimodal agentic models that prevents interaction collapse and sustains multi-turn reasoning via tool rewards.

Ax Varvara Sazonova, Dmitri Shmelkin, Stanislav Kikot, Vasily Motolygin 2/25/2026

Pipeline for Verifying LLM-Generated Mathematical Solutions

Pipeline for automatic and interactive verification of LLM-generated mathematical solutions as alternative to answer-only evaluation, enabling formal and informal solution generation.

Ax Yaacov Pariente, Vadim Indelman 2/25/2026

POMDPPlanners: Open-Source Package for POMDP Planning

POMDPPlanners is an open-source Python package for evaluating POMDP planning algorithms with benchmarks, hyperparameter optimization, and parallel simulation capabilities.

Ax David Koplow, Tomer Galanti, Tomaso Poggio 2/25/2026

Tool Building as a Path to "Superintelligence"

Benchmark measuring step-success probability for LLM superintelligence via test-time search using GF(2) circuit reconstruction tasks.

Ax Mehdi Acheli, Walid Gaaloul 2/25/2026

Motivation is Something You Need

Novel training paradigm inspired by affective neuroscience using motivation states to activate a larger model intermittently alongside continuous base model training.

Ax Debjit Paul, Daniel Murphy, Milan Gritta, Ronald Cardenas, Victor Prokhorov, Lena Sophia Bolliger, Aysim Toker, Roy Miles, Andreea-Maria Oncescu, Jasivan Alex Sivakumar, Philipp Borchert, Ismail Elezi, Meiru Zhang, Ka Yiu Lee, Guchun Zhang, Jun Wang, Gerasimos Lampouras 2/25/2026

A Benchmark for Deep Information Synthesis

DEEPSYNTH benchmark evaluates LLM-based agents on complex tasks requiring multi-source information synthesis and tool use like web browsing and code execution.

Ax Tony Feng, Junehyuk Jung, Sang-hyun Kim, Carlo Pagano, Sergei Gukov, Chiang-Chiang Tsai, David Woodruff, Adel Javanmard, Aryan Mokhtari, Dawsen Hwang, Yuri Chervonyi, Jonathan N. Lee, Garrett Bingham, Trieu H. Trinh, Vahab Mirrokni, Quoc V. Le, Thang Luong 2/25/2026

Aletheia tackles FirstProof autonomously

Aletheia, an AI agent powered by Gemini 3 Deep Think, autonomously solved 6 out of 10 FirstProof mathematical problems.