Polaris: A G\"odel Agent Framework for Small Language Models through Experience-Abstracted Policy Repair
Polaris introduces Gödel agents for small language models enabling recursive self-improvement through policy repair via experience abstraction.
Polaris introduces Gödel agents for small language models enabling recursive self-improvement through policy repair via experience abstraction.
SEDGE framework for generating structured data beyond training distribution with conditions for reliable extrapolation.
Research on training LLMs to explicitly express uncertainty signals within responses during reasoning or at answer time.
UniMamba combines state-space models and attention for efficient multivariate time series forecasting with linear complexity.
Integrates persistent homology and contraction operations into graph neural networks for improved representation learning.
Framework for dynamic mid-generation abstention in chain-of-thought reasoning to reduce wasted compute on incorrect long responses.
Quotient-Space Diffusion Models leverage symmetry in generative tasks for faster 3D molecule structure generation.
LayerBoost proposes layer-aware attention reduction to improve efficient LLM inference by replacing softmax attention selectively across layers.
Research on Self-Preference Bias in LLM-as-a-Judge systems showing LLMs systematically favor their own outputs during evaluation.
Predict-then-Diffuse method for diffusion LLMs enabling adaptive response length under compute budget constraints with parallel generation.
DASE stopping heuristic for LLM ensembles enabling early commitment on consensus and adaptive deliberation with calibrated signals.
Universal Semi-Supervised Learning framework addressing scarce labeled data and unknown unlabeled distributions using structural inference.
PIQL framework integrating privileged information to accelerate learning and improve generalization in tabular foundation models.
Non-monotonic latency behavior in Apple MPS transformer decoding with KV cache interactions, identifying 21x latency spikes.
Multi-task bilevel optimization extending bilevel learning beyond single-task settings with equality constraints for complex ML problems.
RubricRefine improves tool-use agent reliability through training-free pre-execution refinement using rubric-based feedback for code generation.
Full-pipeline FP4 quantization training for large language models on native FP4 hardware, studying MXFP4 in transformer pretraining.
Tree-of-Thought reasoning acceleration for LLMs via speculative exploration, addressing reward dependency bottleneck in complex task solving.
Thompson sampling algorithm for offline-to-online learning addressing distribution shift between offline data and online environments.
Theoretical framework of entropy mechanics in LLM reinforcement learning with verifiable rewards analyzing token-level policy updates.
GEAR: granularity-adaptive advantage reweighting for LLM agents using self-distillation for fine-grained credit assignment.
Supervised fine-tuning analysis for procedural-skill learning across Qwen3.5 model scales (0.8B-4B).
Sparse-to-dense reward principle for LLM post-training combining GRPO sparse rewards with dense token-level distillation.
Continual learning method combining parameter updates and in-context learning for LLMs to adapt without catastrophic forgetting.
WriteSAE: sparse autoencoder decomposing recurrent language model cache writes for interpretability.
ToolMol: evolutionary agentic framework using LLMs with molecular tools for multi-objective drug discovery.
Benchmarking agentic AI for neuroscience data reuse and format standardization across fragmented experimental datasets.
TIDE: graph neural network out-of-distribution detection via information decomposition for robust node classification.
MLGIB: Graph Neural Network method addressing over-squashing in multi-label graphs via information bottleneck.
EMO: progressive training method for Mixture-of-Experts models addressing efficiency paradox in sparse MoE scaling.
Speculative Interaction Agents: framework for real-time LLM agents using asynchronous I/O and speculative tool calling under 1-second latency.
Novel projected gradient methods for nonconvex smooth optimization with improved iteration complexity and auto-conditioned stepsizes.
OMAC framework automatically optimizes multi-agent LLM systems through cost-effectiveness analysis and collaborative patterns.
ActivePusher combines active learning with learned residual dynamics models for nonprehensile robotic manipulation.
Scalable subset selection method for linear mixed models with thousands of candidate predictors.
VER combines multiple vision foundation models through distillation and dynamic routing for flexible robotic learning tasks.
TRIM method uses token-wise attention saliency to identify important samples for efficient LLM instruction tuning with reduced data requirements.
Uses generative models as acquisition functions for batch Bayesian optimization, enabling large-scale optimization of non-continuous and high-dimensional design spaces.
Interactive physical reasoning agent learning human-like causal understanding from game interaction with visual domain gaps.
OPT-ENGINE benchmark evaluating LLM capabilities in optimization modeling across LP to MIP with controlled complexity scaling.
Agent-designed agentic workflows via reinforced canvas editing with graph-level feedback and in-loop error repair for complex multi-step tasks.
Conformal prediction framework for adaptive reasoning in LLMs, controlling risk-accuracy tradeoff when allocating compute budget for inference.
Proxy compression training scheme for language models preserving efficiency of tokenization while enabling raw-byte inference interface.
Top-W: geometry-aware decoding for LLMs using Wasserstein distance over token embeddings to balance diversity and coherence.
Framework extracting distribution maps from LLM next-token probabilities for better statistical text analysis beyond perplexity metrics.
Combines Mamba state-space models with LLMs for dynamic fMRI graph learning in autism diagnosis using multimodal reasoning.
Multi-agent LLM framework for robotic manipulation with closed-loop visual feedback, replacing specialized models with general reasoning.
Scalable reward modeling framework for robotics using trajectory comparisons to handle failed and suboptimal trajectories.
Speculative decoding optimization for LLM inference in high-concurrency serving using elastic sparse gating and dynamic trees.
Scalable framework for training and evaluating agents in claw-style environments with file systems, tools, and persistent workspace.