TempusBench: An Evaluation Framework for Time-Series Forecasting
Evaluation framework and benchmark for assessing time-series foundation models and forecasting approaches.
Evaluation framework and benchmark for assessing time-series foundation models and forecasting approaches.
Semi-supervised reinforcement learning approach using knowledge-enhanced data synthesis to improve medical reasoning in LLMs.
Python package providing bioacoustic deep learning models and evaluation tools for analyzing passive acoustic monitoring data.
Study of multi-class linear classification in transformers using feature and label permutation equivariance to extract interpretable algorithms.
Analytical formalism decomposing Hessian matrices in neural networks with DAG architectures into inter-layer blocks.
Framework for autonomous mechanistic reasoning in virtual cells using LLMs, representing biological reasoning as mechanistic action graphs.
Methodology to mitigate shortcut learning and demographic bias in deep neural networks using geometric a priori approaches.
Model-free reinforcement learning system for autonomous crystal alignment using visual information without domain knowledge of crystallography.
ClawGUI framework for training, evaluating, and deploying GUI agents that interact with software through visual interfaces, with online RL and evaluation infrastructure.
Mechanistic analysis of looped reasoning language models examining internal dynamics and latent state evolution compared to standard feedforward models.
Uses reinforcement learning on physics simulators to train models solving Physics Olympiad problems, addressing lack of large-scale physics QA datasets for reasoning models.
SHANG++: Accelerated stochastic gradient descent methods robust to multiplicative noise in gradient updates.
LABBench2: Improved benchmark for evaluating AI systems and agents on biology research tasks with real-world capabilities.
VTC: DNN compilation method using virtual tensors to eliminate data movement in neural network workloads including LLMs.
Pipeline and best practices for log analysis in AI systems to understand model behaviors, with code examples in Inspect framework.
Benchmark measuring humanization and anti-detection capabilities of mobile GUI agents against platform countermeasures.
Demonstrating LLMs can generate UI interfaces and content together with proper prompting and tool integration.
Object-Oriented Programmatic World Modeling (OOWM) for embodied reasoning and planning in robotic tasks using LLMs.
MobiFlow: Benchmark for mobile agents using trajectory fusion for real-world GUI task evaluation without system-level APIs.
Architecture for maintaining persistent identity in AI agents through multi-anchor memory to prevent catastrophic forgetting.
Spatial Competence Benchmark (SCBench) evaluating large models on spatial reasoning, environment representation, and planning tasks.
ECHO: Speculative decoding optimization for LLM inference in high-concurrency serving with sparse gating.
Benchmark evaluating LLMs as text-only controllers for exploration and navigation in gridworlds under partial observability.
Method for self-calibrating LLMs at test-time through discriminative distillation to reduce overconfidence without labeled data.
Study of backdoor security vulnerabilities in flow-matching Vision-Language-Action models used for robotics, exploiting vector field dynamics.
NeuroPath system for motor imagery decoding from EEG signals for brain-computer interfaces in prosthetics and rehabilitation.
Lightweight speech activity-based approach for real-time voicemail detection in telephony using tree ensemble classification.
Zero-shot modular pipeline for traffic accident detection, localization, and classification without labeled training data.
ASTRA silicon-photonic accelerator for transformers using stochastic computing to reduce computation and memory demands.
Pioneer Agent automates continuous improvement of small language models in production through closed-loop data curation, failure diagnosis, and iteration control.
COMPOSITE-STEM benchmark with 70 expert-written tasks for evaluating AI agents on physics, biology, chemistry, and materials science problems.
Research on activation steering in LLMs showing steered states are non-surjective, with implications for interpretability and safety.
MEMENTO teaches LLMs to compress reasoning into dense summaries, reducing context and compute requirements. Releases OpenMementos dataset of 228K examples.
Proposes hybrid fine-tuning paradigm for LLMs combining full and parameter-efficient approaches with convergence analysis framework.
Evaluates reproducibility of ColBERT-v2 and ConstBERT retrieval models across different query types, finding architectural limitations on long narrative queries.
Context compression daemon (entroly) for LLM applications reducing token usage by 90% through self-evolving compression. Claims token-negative learning.
MCP server integrating Claude with personal finance data via Model Context Protocol. Enables AI agents to access bank accounts, cards, and investments read-only.
Rust open-source PAM module replacing SSH keys with short-lived OIDC tokens secured via DPoP cryptographic proof. Developer tool for infrastructure.
Integration making Shopee e-commerce products machine-readable for ChatGPT and Perplexity through structured data. Enables AI agent product discovery.
Skywork AI technical report on Matrix-Game 3.0 real-time streaming world model with long-horizon memory for video generation. Published on GitHub/HuggingFace.
OpenAI acquires Hiro Finance, a personal finance startup. Suggests direction toward financial AI agents. Limited technical details; appears to be acquihire.
Polara autonomous marketing platform using specialist agent architecture for strategy, content, and analytics. No technical validation provided.
Desktop automation tool enabling AI agents to control macOS by viewing screens, moving cursor, and typing. Works with OpenAI-compatible models.
Open source browser extension providing AI-powered code reviews for GitLab, GitHub, and Bitbucket with Ollama local model support.
Research agent for automated computer vision dataset curation using retrieval, annotation, and synthesis models composed through multi-agent orchestration.
Persistence layer for AI agent workflows enabling save, resume, and replay across sessions and crashes. Supports JavaScript, Python, and MCP agents.
AI models solve 5 of 6 International Mathematical Olympiad problems in summer 2025. Discusses implications of AI capabilities in mathematical problem-solving.
Open source LLM memory management system using memory palace concept for storing and retrieving agent memories. Addresses memory management challenges in language models.
AI Native IDE called 6digit studio featuring CORDIAL visualization layer for Big Picture Mode. Developer tool with spatial UI designed for distance interaction.
Analysis of AI trading bots and LLM-based investing. Reviews early retail trading attempts showing results indistinguishable from random, discusses limitations.