Does Your Reasoning Model Implicitly Know When to Stop Thinking?
Study on whether Large Reasoning Models know when to stop thinking, addressing redundancy in long chains-of-thought.
Study on whether Large Reasoning Models know when to stop thinking, addressing redundancy in long chains-of-thought.
Training method for Large Reasoning Models using adaptive reflection and length penalties to reduce unnecessary token consumption.
ForesightSafety Bench evaluates frontier risks in autonomous AI with unpredictable and difficult-to-control behaviors.
IntentCUA framework for computer-use agents with intent-aligned planning and multi-agent coordination over long horizons.
Layered execution structures for tool orchestration in agentic systems with reflective error correction mechanisms.
Aletheia AI agent solved 6/10 FirstProof mathematics challenges autonomously using Gemini 3 Deep Think reasoning.
Framework for causal embeddings enabling multiple detailed models to map into sub-systems of coarser causal models.
Examines AI agents with persistent state, tool access, and skills for autonomous execution of social science research pipelines.
ConstraintBench evaluates whether LLMs can directly solve constrained optimization problems without solver access.
Study on LLM vulnerability to jailbreak attacks using classical Chinese prompts to bypass safety constraints.
Dispatcher/Executor principle for multi-task reinforcement learning using abstraction to improve generalization across tasks.
R2GenCSR uses LLMs with visual feature extraction from X-ray images for automated radiology report generation.
Research framework for sparse counterfactual explanations using optimal transport and Shapley values for model interpretability.
Research on robust watermarking techniques for distinguishing generated from real content in generative models.
Research paper on grounding LLMs with real-time financial data for knowledge-aware financial agent applications.
Semantic parallelism technique for efficient MoE LLM inference via model-data co-scheduling reducing communication bottlenecks.
Optimization perspective on reward model quality in RLHF showing accuracy alone doesn't capture effective teacher properties.
Domain decomposition approach for neural operators to improve geometry generalization and transferability in PDE solving.
LLM-empowered hierarchical RIC controller for O-RAN addressing cooperation, computational demands, and domain-specific adaptation.
FineScope framework using SAE-guided data selection for domain-specific LLM pruning and finetuning with maintained performance.
Feature selection method using permutation-invariant embeddings and policy-guided search for complex feature interactions.
Agentic Predictor using multi-view encoders for performance prediction in LLM-based agentic workflows without exhaustive evaluation.
Method bridging target-free and target-based deep reinforcement learning to reduce memory requirements and improve update propagation.
Framework converting generative multimodal LLMs into zero-shot discriminative embedding models without extensive pre-training.
OM2P offline multi-agent reinforcement learning using flow-based generative models with improved sampling efficiency.
Framework for improving LLM context-aided forecasting with diagnostic tools and reduced computational costs for practical deployment.
AC3 reinforcement learning framework for long-horizon robotic manipulation using continuous action chunking with sparse rewards.
LumiMAS framework for real-time monitoring and observability of multi-agent systems with LLMs, addressing system-wide failure detection.
SAT reduction approach for automating input/output logics, a family of deontic logics for reasoning over norms and obligations.
Latent Self-Consistency: method for reliable majority voting in LLM outputs handling both short and long-form reasoning tasks consistently.
Once4All: LLM-synthesized test generator for SMT solver fuzzing using skeleton guidance to uncover bugs in evolving solver versions.
Veritas: pattern-aware deepfake detection system with HydraFake dataset bridging gap between academic benchmarks and industrial deployment.
Draw-In-Mind: multimodal model rebalancing designer and painter roles to improve precision in image editing tasks.
LLaDA diffusion-based large language model applied to automatic speech recognition with deliberation-based post-processing for Whisper transcripts.
E-CIT: plug-and-play ensemble framework for conditional independence testing to reduce computational bottlenecks in constraint-based causal discovery.
Investigation of in-context learning emergence in world models for environmental dynamics prediction beyond static zero-shot performance.
Study of activation function design's role in preventing plasticity loss during continual learning, beyond catastrophic forgetting.
Meta-weighted online sampling approach for aligning LLMs by reducing distribution mismatch between offline preference data and evolving model policy.
MobileLLM-R1: sub-billion parameter language models with chain-of-thought reasoning and open training recipes, challenging assumptions about model size requirements.
DataMind: scalable data-analytic agent system with open-source training recipes for multi-step reasoning over diverse-format, large-scale data files.
BEV-VLM: trajectory planning approach using vision-language models with bird's-eye-view representations from fused camera and LiDAR data.
VoiceBridge: one-step latent bridge model for general speech restoration from diverse distortions with energy-preserving VAE design.
Max-V1: vision-language model framework for autonomous driving that formulates trajectory planning as next waypoint prediction via language.
FLOP: score-based causal discovery algorithm for linear models using fast parent selection and Cholesky-based updates to find optimal causal graphs.
CMT-Benchmark: 50 expert-level condensed matter theory problems for evaluating LLMs on advanced scientific reasoning and code generation.
Permutation-invariant feature selection method using generative models to capture feature interactions while improving robustness and privacy.
Flow matching variant (Carré du champ flow matching) that improves quality-generalization tradeoff in generative models through geometry-aware noise regularization.
Backdoor attack on vision-language-action models demonstrating vulnerability to behavioral hijacking via hidden training triggers.
Bayesian optimization method using LLM fine-tuning to perform Thompson sampling in large discrete spaces without gradient computation.
Training-free framework for improving vision-language models on information-dense images with text and graphical elements.