Studies thermodynamic regulation of finite-time Gibbs chain training in Restricted Boltzmann Machines, analyzing energy landscape evolution during learning.
Establishes theoretical connection between classifier-free guidance in diffusion models and Anderson acceleration via Hopfield dynamics.
Presents EdgeFLow, a federated learning framework using sequential model migration in edge networks to reduce communication bottlenecks in IoT systems.
Derives Wasserstein Proximal Policy Gradient using optimal transport geometry for continuous-action entropy-regularized RL without policy log-density evaluation.
Develops parameter-free temporal difference learning for RL that avoids requiring problem-dependent quantities like feature covariance eigenvalues.
Studies DNN partitioning and resource allocation for device-edge collaborative inference under jamming attacks on resource-constrained systems.
Introduces HACRL, a collaborative reinforcement learning paradigm where heterogeneous agents share verified rollouts during training but execute independently at inference.
Proposes diffusion actor-critic with flow matching for real-time autonomous driving policies, addressing inference latency in generative RL approaches.
Theoretical analysis of implicit regularization in Deep Linear Discriminant Analysis for metric learning objectives.
Multi-objective reinforcement learning method for extracting Pareto fronts of policies in continuous control tasks, addressing trade-offs between multiple objectives.
Bandit-based prompt optimization for multi-agent systems using graph neural networks to improve LLM-powered workflow performance without modifying workflows.
Heterogeneous analog-digital computing approach for efficient Mixture-of-Experts inference with theoretical generalization guarantees and hardware nonideality mitigation.
SaFeR-ToolKit formalizes multimodal safety as checkable protocol using virtual tool calling for vision-language models to prevent jailbreaks.
HomeAdam variant of Adam optimizer improves generalization bounds to match SGD convergence rates for deep learning model training.
SAGE method improves diffusion planners for offline RL by using latent consistency signals to penalize dynamically inconsistent plans at inference-time.
Two-Stage Causal-GRPO framework addresses shallow safety alignment in LLMs vulnerable to adversarial prefix attacks through semantic intent pinning.
Paradigm for causal structure learning from observational data leveraging human causal knowledge to address combinatorial explosion of possible graphs.
Unified framework addressing both missing and noisy modalities in multimodal learning to improve robustness on low-quality real-world data.
Empirical evaluation of uncertainty-based selective prediction reliability in multimodal clinical condition classification using ICU data.
Theoretical analysis of factorized gradient descent for low-tubal-rank tensor recovery from noisy linear measurements under t-product framework.
Training recipe enabling FP4 efficiency for large-scale Mixture-of-Experts models on Hopper GPUs without native 4-bit support.
BoGA framework combines evolutionary search with Bayesian optimization for protein sequence design and function prediction.
Domain generalization for time series via structure-stratified calibration addressing heterogeneous dynamical systems with distinct feature distributions.
NE-Dreamer agent uses temporal transformer to predict next-step embeddings for improved model-based reinforcement learning in high-dimensional domains.
LLMs for automated algorithm design benefit from strong priors; investigates token-wise attribution of prompts guiding algorithm generation.
Fine-tuning time series foundation models using data mixtures and LoRA for improved zero-shot forecasting on new domains.
Deep reinforcement learning approach for flexible job-shop scheduling using memory-enhanced improvement heuristics for manufacturing optimization.
RL algorithm design for MDPs with exogenous dynamics where only subset of state variables are affected by agent actions.
Interpretable time series forecasting approach using polynomial learning to improve trust and debuggability for predictive maintenance applications.
Method for obtaining numerical predictive distributions from LLMs without autoregression for efficient regression tasks like time series and tabular forecasting.
Study on structural irreversibility in weight-based neural model adaptation, proposing reversible behavioral learning as alternative to parameter fine-tuning.
Framework introducing contextual latent world models for offline meta-reinforcement learning with improved task representation learning via self-supervised methods.
Approach combining LLMs with Graph Neural Networks for zero-shot graph learning using adaptive subgraph denoising to handle cross-modal alignment issues.
Framework for continual learning in GUI agents using multimodal LLMs with reinforcement fine-tuning to adapt to new tasks without catastrophic forgetting.
Method addressing class imbalance in semi-supervised learning using Proportion Loss regularization to align predictions with global class distribution.
Research on federated learning combining homomorphic encryption and synthetic data to improve privacy and learning quality while reducing computational costs.
arXiv: Optimization algorithm combining Bayesian optimization with trust region methods through adaptive competition.
arXiv: Theoretical explanation for reinforcement learning from AI feedback through latent value hypothesis.
arXiv: Federated contrastive learning framework addressing prototype bias in imbalanced distributed data.
arXiv: Information-theoretic feature selection for multi-view multi-label learning using structural entropy.
arXiv: Step-level sparse autoencoders for interpreting LLM reasoning processes in chain-of-thought outputs.
arXiv: Continuous progressive neural networks handling streaming time series with concept drift and temporal dependencies.
arXiv: Incremental k-NN graph construction method improving robustness of spectral clustering on text embeddings.
arXiv: Reinforcement learning approach using symbolic reward machines to eliminate manual labeling functions.
arXiv: Study of Transformer expressive power proving universal approximation for maxout networks and piecewise linear functions.
arXiv: Theoretical analysis explaining Adam optimizer's empirical advantage over SGD through second-moment normalization.
arXiv: Multi-scale adaptive transformer for graph fraud detection using neighborhood awareness.
DynFormer: Transformer architecture for solving PDEs that respects scale separation in complex dynamics, replacing expensive classical solvers.
Training strategy applying global top-k activation sparsity constraints cyclically to improve neural network generalization across dense/sparse regimes.
Torus embeddings: research on representing deep learning embeddings on toroidal manifolds instead of Euclidean space for efficiency.