What Makes a Reward Model a Good Teacher? An Optimization Perspective
Optimization perspective on reward model quality in RLHF showing accuracy alone doesn't capture effective teacher properties.
Optimization perspective on reward model quality in RLHF showing accuracy alone doesn't capture effective teacher properties.
Domain decomposition approach for neural operators to improve geometry generalization and transferability in PDE solving.
LLM-empowered hierarchical RIC controller for O-RAN addressing cooperation, computational demands, and domain-specific adaptation.
FineScope framework using SAE-guided data selection for domain-specific LLM pruning and finetuning with maintained performance.
Feature selection method using permutation-invariant embeddings and policy-guided search for complex feature interactions.
Agentic Predictor using multi-view encoders for performance prediction in LLM-based agentic workflows without exhaustive evaluation.
Method bridging target-free and target-based deep reinforcement learning to reduce memory requirements and improve update propagation.
Framework converting generative multimodal LLMs into zero-shot discriminative embedding models without extensive pre-training.
OM2P offline multi-agent reinforcement learning using flow-based generative models with improved sampling efficiency.
Framework for improving LLM context-aided forecasting with diagnostic tools and reduced computational costs for practical deployment.
AC3 reinforcement learning framework for long-horizon robotic manipulation using continuous action chunking with sparse rewards.
LumiMAS framework for real-time monitoring and observability of multi-agent systems with LLMs, addressing system-wide failure detection.
SAT reduction approach for automating input/output logics, a family of deontic logics for reasoning over norms and obligations.
Latent Self-Consistency: method for reliable majority voting in LLM outputs handling both short and long-form reasoning tasks consistently.
Once4All: LLM-synthesized test generator for SMT solver fuzzing using skeleton guidance to uncover bugs in evolving solver versions.
Veritas: pattern-aware deepfake detection system with HydraFake dataset bridging gap between academic benchmarks and industrial deployment.
Draw-In-Mind: multimodal model rebalancing designer and painter roles to improve precision in image editing tasks.
LLaDA diffusion-based large language model applied to automatic speech recognition with deliberation-based post-processing for Whisper transcripts.
E-CIT: plug-and-play ensemble framework for conditional independence testing to reduce computational bottlenecks in constraint-based causal discovery.
Investigation of in-context learning emergence in world models for environmental dynamics prediction beyond static zero-shot performance.
Study of activation function design's role in preventing plasticity loss during continual learning, beyond catastrophic forgetting.
Meta-weighted online sampling approach for aligning LLMs by reducing distribution mismatch between offline preference data and evolving model policy.
MobileLLM-R1: sub-billion parameter language models with chain-of-thought reasoning and open training recipes, challenging assumptions about model size requirements.
DataMind: scalable data-analytic agent system with open-source training recipes for multi-step reasoning over diverse-format, large-scale data files.
BEV-VLM: trajectory planning approach using vision-language models with bird's-eye-view representations from fused camera and LiDAR data.
VoiceBridge: one-step latent bridge model for general speech restoration from diverse distortions with energy-preserving VAE design.
Max-V1: vision-language model framework for autonomous driving that formulates trajectory planning as next waypoint prediction via language.
FLOP: score-based causal discovery algorithm for linear models using fast parent selection and Cholesky-based updates to find optimal causal graphs.
CMT-Benchmark: 50 expert-level condensed matter theory problems for evaluating LLMs on advanced scientific reasoning and code generation.
Permutation-invariant feature selection method using generative models to capture feature interactions while improving robustness and privacy.
Flow matching variant (Carré du champ flow matching) that improves quality-generalization tradeoff in generative models through geometry-aware noise regularization.
Backdoor attack on vision-language-action models demonstrating vulnerability to behavioral hijacking via hidden training triggers.
Bayesian optimization method using LLM fine-tuning to perform Thompson sampling in large discrete spaces without gradient computation.
Training-free framework for improving vision-language models on information-dense images with text and graphical elements.
Study of continual pre-training for adapting LLMs to low-resource French dialects under tight compute and data constraints.
Investigation of in-context learning across transformer, state-space, and hybrid LLM architectures using behavioral and intervention methods.
Study of misconceptions novice programmers have about LLM-based coding assistants, examining impact of tool capabilities and extensions.
Training method combining supervised learning and reinforcement learning to improve multi-step reasoning in open-source LLMs.
Agentic multimodal model framework enabling tool invocation (code execution, web search) and reasoning integration for vision-language tasks.
Analysis of how LLMs shift moral judgments under persona role-play, introducing benchmark metrics for moral susceptibility and robustness.
Open benchmark for deep learning-based event reconstruction in neutrino telescope data using inverse problem solving.
Diffusion language model using Mamba backbone for efficient inference, achieving higher throughput than transformer-based alternatives.
Benchmark for evaluating embodied AI agents on interaction with physical interfaces (switches, panels, GUIs) in complex environments.
Machine learning system for automating data quality monitoring and anomaly detection in particle physics collider experiments.
SocialNav foundation model for socially-aware embodied navigation with hierarchical architecture trained on 7M samples for human-compliant trajectory generation.
Heterogeneous multi-agent reinforcement learning with attention mechanism for automated feature transformation on structured data.
QKAN-LSTM combining quantum-inspired Kolmogorov-Arnold networks with LSTM for improved sequential modeling with reduced parameter redundancy.
WisPaper end-to-end agent system for academic literature discovery and organization combining semantic search verification with workflow integration.
FRIEDA benchmark evaluating vision-language models on multi-step cartographic reasoning with map interpretation for disaster response and urban planning.
Generalized Primal Averaging optimizer extending Nesterov's method for faster LLM training, unifying DiLoCo and schedule-free approaches with reduced memory requirements.