Principled Analysis of Deep Reinforcement Learning Evaluation and Design Paradigms
Survey of design paradigms and evaluation practices in deep reinforcement learning research spanning algorithms from DQN to model-free methods.
Survey of design paradigms and evaluation practices in deep reinforcement learning research spanning algorithms from DQN to model-free methods.
Theoretical proof of robustness law for two-layer neural networks without weight restrictions, extending Bubeck-Sellke conjecture.
Framework for continual learning in LLMs distinguishing between domain adaptation and competence improvement as world conditions change.
NFTR method for offline goal-conditioned reinforcement learning using normalizing flows to address optimistic bias and mode collapse in subgoal selection.
Physics-informed ML methodology for small-dataset manufacturing using abrasive waterjet milling, addressing data cleaning and curation with physics integration.
Theoretical analysis of gradient descent dynamics in deep scalar linear networks showing optimal learning rate scaling depends on data properties.
Survey of multimodal unlearning methods for VLMs, DMs, LLMs across vision, language, video and audio with datasets and benchmarks.
Latent Personality Alignment method for efficient LLM safety training using 66 harm-agnostic statements instead of large adversarial datasets.
Python package implementing PathBoost gradient boosting for interpretable graph-level predictions via discovered path-based features.
Evaluation of time series foundation models on wildfire PM2.5 forecasting, assessing generalization under extreme out-of-distribution conditions.
Comparative study of linear attention architectures versus softmax attention, analyzing efficiency trade-offs for long context processing.
KronQ post-training quantization framework using Kronecker-factored Hessian for improved LLM compression beyond standard PTQ methods.
Provably efficient learning algorithms for assistance games where informed and uninformed agents repeatedly interact with shared objectives.
Federated learning system optimizing inference latency for heterogeneous edge devices in real-time collaborative neural network training.
Neural operator architecture using precomputed geometry decomposition for scaling physics simulations to million-scale 3D problems.
Systematic evaluation of quantization for small vision-language models with hardware-aware deployment on edge devices like Jetson Orin.
Rate-distortion framework for memory compaction in LLMs and agents, analyzing KV cache, prompt, and state compression trade-offs.
Theoretical analysis of how Bayesian diffusion models avoid overfitting through information-theoretic lens with analytically tractable models.
Framework for validating LLM-assisted safety analysis tools using Constitutional Meta-STPA to identify hallucinations in system analysis.
Research on optimizing token generation order in diffusion models for text-to-image synthesis and multimodal understanding tasks.
Survey on system-aware KV cache optimization techniques for efficient LLM serving, addressing memory-intensive inference bottlenecks.
Empirical characterization of uncertainty quantification in vision language models with chain-of-thought reasoning.
Modular pretraining framework enabling fine-grained access control over AI capabilities without training multiple models.
Symbolic regression framework with dynamic pruning and Pareto selection for discovering interpretable equations from data.
ML research on zero-shot model size interpolation via layer patching and boomerang distillation for language models.
RL research using multi-modal LLMs to inspect agent policies and design open-ended curricula for training complex agents.
RhyMix lightweight adaptive network for time series forecasting capturing multiple simultaneous temporal patterns through multi-rhythm modeling.
DAG structure learning approach for causal discovery on clustered data accounting for cluster-specific variations common in scientific applications.
CASL-VAE deep contrastive latent variable model learns structured generative factors from unpaired data for clustering and paired sample generation.
Theoretical work on learning constant-depth circuits under locally sampleable graphical models extending low-degree algorithm results to correlated distributions.
Analyzes structural limitations in language-grounded world models for robotics, identifying safety constraints for LLM/VLM feature integration with symbolic systems.
ArtMine system for discovering and formalizing artistic creative processes beyond finished artifacts using generative AI and iterative decision reasoning.
AutoAnchor method for diffusion model unlearning using cross-attention as manifold surrogate to mitigate harmful/copyrighted content generation.
Spectral analysis of dueling Q-learning algorithm extending Q-learning with value/advantage function decomposition for high-dimensional RL problems.
Proposes certified interventional fidelity framework for anytime-valid adaptive evaluation of causal claims in mechanistic interpretability research.
Reinforcement learning approach for self-adaptive anomaly detection in connected vehicles handling system evolution from updates and configuration changes.
Framework for calibrating eigenvalues of semantic embeddings from LLMs for uncertainty quantification in reliable model deployment.
Theoretical analysis of gradient descent dynamics with large step sizes near flat minima manifolds, addressing violations in deep neural network training.
Demonstrates gradient-free training of deep neural networks using Monte Carlo method on GPU, avoiding vanishing/exploding gradient problems of backpropagation.
MatBind creates shared embedding space for multimodal materials characterization integrating atomic structures, diffraction patterns, density of states, and language.
Proposes Ensemble Diversity Optimization (EDO) framework for NLP tasks with annotator disagreement, jointly optimizing ensemble weights and calibration via differentiable objective.
Study systematically evaluates learning rate scheduling strategies across 30 neural network architectures (CNNs and transformers) to understand impact on classification accuracy.
Efficient model evaluation method using adaptive sample sizes to match diverse evaluation objectives and reduce computational cost.
Anomaly detection framework incorporating causal relationships and structural consistency for industrial system failure diagnosis.
Theoretical analysis showing strong alignment of minimal DNN solutions to hard tasks through contravariance, comparing networks and brains.
Framework for guiding neural network training using interpretable constraints based on partial dependence analysis for faithful explanations.
Binary spherical coding compression method for extreme low-bit LLM quantization, achieving 2-bit weight representation without lookup tables.
Multi-environment inverse RL approach for learning robust reward functions that generalize across diverse operational contexts for autonomous agents.
Budget-aware test-time routing mechanism for selecting among LLMs, balancing response quality against serving cost via resampling.
Investigates relaxed speculative decoding to accelerate LLM sampling with faster auxiliary models, exploring speed-capability trade-offs.