Design Conditions for Intra-Group Learning of Sequence-Level Rewards: Token Gradient Cancellation
Analysis of sequence-level reward learning in reinforcement learning for reasoning models, addressing gradient cancellation and learning efficiency.
Analysis of sequence-level reward learning in reinforcement learning for reasoning models, addressing gradient cancellation and learning efficiency.
Langevin Gradient Descent algorithm with generalization guarantees for hyperparameter tuning in convex regression via learning to learn.
Graph-based hierarchical reinforcement learning for automated co-design of thermodynamic cycle parameters.
Pareto-optimal offline reinforcement learning via Tchebysheff scalarization for multi-objective LLM alignment and optimization.
KV Packet enables context-independent key-value caching for LLMs without recomputation, improving inference latency.
Studies whether dimensionality reduction via random projections preserves landscape features for exploratory landscape analysis.
Context-dependent anomaly detection framework for multimodal data recognizing that anomalies depend on contextual factors.
Bias-corrected adaptive conformal inference for multi-horizon time series forecasting with distribution shift adaptation.
Counterfactual invariant prediction framework prevents shortcut learning in TCR-pMHC binding neural prediction models.
Binomial gradient-based meta-learning approach to reduce computational overhead in gradient-based meta-learning.
Twin-pass chain-of-thought ensembling method to improve confidence estimation reliability in telecommunications LLMs.
MOONSHOT framework for multi-objective one-shot pruning of vision and large language models without retraining.
Combines active learning and input denoising to improve robustness of neural operators against adversarial perturbations.
Multi-task LLM framework with LoRA fine-tuning for automated cancer staging and biomarker extraction from pathology reports.
Uses LLMs to enrich knowledge graphs for medical concept representation in EHR mining and clinical prediction tasks.
TabDistill method leverages tabular foundation models to identify feature interactions for generalized additive models on tabular data.
Orthogonal Backfill compression for LLM multi-agent systems reducing KV cache relay costs while preserving communication context and information.
BioTrain framework enabling sub-50mW on-device fine-tuning for MCU-based wearable edge AI on biosignals addressing domain shift and privacy.
Diffusion sequence models and Transformer meta-models for in-context robot dynamics learning addressing distributional shifts and real-time constraints.
Fine-grained non-determinism evaluation in diffusion language models showing dataset-level metrics mask run-to-run variations and condition sensitivity.
WIN-U: Woodbury-informed Newton method for machine unlearning in LLMs enabling 'right to be forgotten' without requiring retain set data.
Forward-only KL-based sensitivity analysis for mixed-precision quantization of hybrid SSM-Transformer LLMs targeting edge device deployment.
Proposes Chain of Uncertain Rewards method for designing reward functions in RL using LLMs, addressing inefficiencies in manual reward design and capturing intermediate uncertainties.
Research on SFT-GRPO data overlap as a post-training hyperparameter for Lean 4 autoformalization using Qwen3-8B, ablating overlap percentages (0%, 30%, 100%).
Analysis of multi-timescale PPO revealing surrogate hacking when fusing multi-scale signals, proposing representation-focused solutions.
Study of self-supervised learning and predictive representation learning comparing alignment and reconstruction approaches.
Confidence-based test-time voting mechanism for latent recurrent neural networks enabling test-time scaling without explicit energy functions.
DynamicGate MLP architecture permitting concurrent learning and inference by separating routing from representation parameters.
Parameter-efficient quantum multi-task learning with shared backbone and task-specific heads for quantum neural networks.
RL approach for radiology report generation using evidence-aware rewards and self-correcting preference learning for clinical alignment.
Analyzes reward hacking vulnerabilities in RLHF and alignment approaches for LLMs, examining mechanisms and emergent misalignment issues.
Bayesian mitigation strategy for safer AI agents using expanded subjective reward range to prevent reward hacking via risk aversion.
Studies how learning rates regulate catastrophic overtraining in LLM fine-tuning through catastrophic forgetting lens.
Bayesian framework for uncertainty-aware explainable AI in power quality disturbance classification with instance-specific interpretations.
Python package for surrogate-model-based optimization using Kriging, Expected Improvement, and multi-objective extensions for expensive functions.
Combines Vision-Language-Action models with RL for robotic manipulation, enabling efficient long-horizon task learning with sparse rewards.
Multi-step off-policy soft Q-learning using eligibility traces for entropy-regularized RL with formal n-step formulation.
arXiv paper on role-playing evaluation in audio LLMs using reinforcement learning to align character attributes in speech dialogue systems.
arXiv paper on DASH-Q: post-training quantization for LLMs using stable diagonal curvature estimates for robust ultra low-bit compression.
arXiv paper on RPS: reinforcement prompt selection method for LLMs to elicit concealed information in interactive conversations.
arXiv paper on UI-Copilot: MLLM-based GUI agent framework with tool-integrated policy optimization for long-horizon automation tasks.
arXiv paper on behavior consistency in text-based world models for evaluating agent planning and offline evaluation beyond single-step metrics.
arXiv paper on SparseBalance: load-balanced training for long-context LLMs using dynamic sparse attention to address sequence length and sparsity imbalance.
arXiv paper on neuro-symbolic networks using Exp-Minus-Log operator for hardware-efficient inference in resource-constrained settings.
arXiv paper on deep reinforcement learning for adaptive autonomous braking detecting driver drowsiness using EEG data.
arXiv paper on principles and pitfalls in evaluating supervised ML models, covering metric selection and real-world performance assessment.
MolCryst-MLIPs: Open database of machine-learned interatomic potentials for molecular crystals. MACE models for nine systems with automated ML pipeline.
DiPO: Fine-grained exploration-exploitation tradeoff for reinforcement learning with verifiable rewards. Improves LLM reasoning training by handling hard/easy samples.
ASTER: Unsupervised time-series anomaly detection via latent pseudo-anomaly generation. Addresses heterogeneous anomalies and labeled data scarcity.
Unsupervised anomaly detection in complex industrial time-series using real production data. Empirical evaluation of existing methods on heterogeneous processes.