Self-Improving Skill Learning for Robust Skill-based Meta-Reinforcement Learning
Hierarchical meta-reinforcement learning approach using self-improving skills to handle noisy offline demonstrations.
Hierarchical meta-reinforcement learning approach using self-improving skills to handle noisy offline demonstrations.
Reversible Runge-Kutta solvers for neural differential equations in generative models with improved numerical stability.
Continual learning approach for foundation models addressing stability-plasticity trade-off during post-training on new classes/domains.
Parameter-efficient fine-tuning method using orthogonal adaptation on principal subspaces for adapting large models efficiently.
Survey demystifying common beliefs about oversmoothing, oversquashing, heterophily and long-range tasks in graph neural networks.
Analyzes KL-regularization design choices in policy gradient algorithms for LLM reasoning, comparing forward/reverse KL variants.
Strict Subgoal Execution method for hierarchical RL that improves long-horizon planning by validating subgoal feasibility.
AXLearn production system for scalable hardware-agnostic training of large models with modular software architecture.
Extends reinforcement learning to continuous-time systems using Hamilton-Jacobi-Bellman equations for irregular interaction frequencies.
First watermarking method for diffusion language models that generate tokens non-sequentially, addressing unique DLM challenges.
Analyzes overthinking in reasoning LLMs and proposes early exit using entropy after thinking tags to improve efficiency.
Extends prompt optimization to multimodal LLMs, proposing methods to optimize across text, images, video and other modalities.
Introduces ConDA, a contrastive learning layer for organizing diffusion model latent spaces to enable controllable generation.
Proposes pi-Flow, a policy-based approach to improve few-step diffusion model distillation by predicting network-free policies.
System for accelerating neural network inference on mobile devices through fine-grained CPU-GPU co-execution and synchronization optimization.
LRT-Diffusion applies risk-aware sequential hypothesis testing to improve diffusion policy guidance for offline reinforcement learning.
Semi-supervised preference optimization framework for LLM alignment using limited labeled paired feedback data.
PREPO method improving data efficiency for LLM reinforcement learning with verifiable rewards using intrinsic exploration signals.
Model-agnostic local explanation method using MARS and N-ball sampling for high-fidelity black-box model interpretability.
Distributionally robust reinforcement learning approach using general function approximation for policy robustness under environment shift.
Unified framework connecting physics-informed neural networks and neural operators for learning PDE solvers.
Temporal graph pattern machine for learning transferable representations in dynamic networks without restrictive assumptions.
Method for jointly optimizing data mixture and model architecture configurations during LLM training to avoid suboptimal individual choices.
Automated black-box pipeline for detecting unverbalized biases in LLM reasoning traces and chain-of-thought explanations.
Analysis showing logit distance bounds representational similarity in discriminative models including autoregressive language models.
HPMixer model for long-term multivariate time series forecasting using hierarchical patching to capture periodic patterns and residuals.
Framework for embodied AI agents to infer user goals from open-ended dialog using LLMs for efficient task accomplishment.
Knowledge distillation pipeline to compress Dust3r foundation model for efficient 3D reconstruction and visual localization.
Perceiver architecture for auto-regressive language modeling reducing attention complexity from quadratic to semi-linear.
Model-based data filtering framework for multilingual LLM pretraining that identifies diverse, high-quality training samples.
Combining Self-Organizing Maps with Vision Transformers to improve performance on smaller datasets through explicit inductive biases.
Learning user-specialized reward models for reinforcement learning from human feedback to capture individual preference disagreement.
Framework for handling unstructured data feature extraction with neural networks while accounting for measurement bias in economic analysis.
ReplaceMe: training-free depth pruning method replacing transformer blocks with linear operations for efficient model compression.
LLM fingerprinting via semantically conditioned watermarks that survive finetuning and quantization without being easily detected.
∞-THOR framework for long-horizon embodied AI tasks with Needle(s) in Embodied Haystack benchmark for testing long-context reasoning in agents.
Using persona-driven prompting to simulate European Parliament voting behavior with LLMs, addressing political bias in model responses.
SPECS: method for faster test-time scaling in LLMs through speculative drafts, balancing reasoning accuracy with user-facing latency.
Bongard Problems benchmark using real-world images to test abstract visual reasoning and fine-grained concept identification in models.
Proposes inference-time search algorithm that guides diffusion model sampling with side information for improved image reconstruction in inverse problems.
LayerSync regularizes diffusion models using their own intermediate layer representations to improve generation quality and training efficiency.
Uses LLMs for automated assessment of critical thinking skills in educational contexts, addressing evaluation of evidence and claim reliability.
Extends wireless foundation models to accept multiple input modalities for improved task performance and adaptation across varying conditions.
Introduces Block-Recurrent Hypothesis explaining Vision Transformer depth as block-recurrent computational flow for mechanistic interpretation.
Proposes world model approach for offline multi-agent RL using local-to-global puzzle solving to overcome conservative policies and improve generalization.
Framework for treating representation reliability as a first-class property in machine learning, beyond traditional predictive uncertainty quantification.
Proposes implementing optical neural networks using linear optical resources and phase-shift encoding for neuromorphic machine learning hardware.
SpikeScore method for detecting LLM hallucinations that generalizes across domains, addressing the gap in cross-domain hallucination detection for real-world deployment.
Proposes exploration-exploitation optimization for dataset distillation to compress large datasets into synthetic versions while maintaining model performance.
Addresses temporal leakage in clinical NLP models for discharge planning, proposing methods to prevent overconfident predictions from deployment artifacts.