Delayed Homomorphic Reinforcement Learning for Environments with Delayed Feedback
Reinforcement learning approach for handling delayed feedback by replacing state augmentation with homomorphic methods to reduce sample complexity.
Reinforcement learning approach for handling delayed feedback by replacing state augmentation with homomorphic methods to reduce sample complexity.
Mechanistic interpretability method for discovering repeated attention patterns in large language models at scale without resource-intensive controlled settings.
CountsDiff: diffusion model framework for generating and imputing count-based discrete ordinal data using survival probability schedules.
Framework for automated mathematical conjecture resolution combining LLMs with formal verification to improve reliability of research-level mathematical problem solving.
Research on representational collapse in multi-agent LLM committees using majority voting, measuring agent diversity via cosine similarity and effective rank on mathematical reasoning tasks.
k-Maximum inner product attention for graph transformers addressing quadratic complexity while maintaining expressive power of GraphGPS.
DDCL-Attention: Prototype-based readout layer for transformer encoders using soft probabilistic token matching for compact summaries.
Bayesian information-theoretic approach to training data attribution for tracing model predictions to influential training examples.
Method for input-dependent layer selection in steering vectors to improve LLM alignment at inference time, adapting intervention layer per input.
SODA: Semi on-policy knowledge distillation method for LLMs balancing off-policy simplicity with on-policy effectiveness without adversarial training instability.
Theoretical research on multi-task representation learning for reinforcement learning with shared representations across related RL tasks with different rewards.
Framework combining structure pretraining with diffusion models for generating molecular dynamics trajectories with limited MD data.
ACES: method for selecting LLM-generated code using LLM-generated tests via leave-one-out AUC consistency without determining test correctness.
Low-bit mixed-precision attention kernel using MXFP for efficient transformer inference with reduced memory bandwidth.
BWTA: binarized transformer quantization scheme with ternary activations and algorithm-hardware co-design for efficient inference.
Analysis of LLM reasoning models under noisy labels in reinforcement learning with verifiable rewards, identifying label noise vulnerabilities.
ArrowFlow: novel ML architecture operating in permutation space using ranking filters and permutation-matrix updates without gradients.
Generalization analysis of stochastic bilevel optimization with applications to hyperparameter optimization, meta-learning, and RL.
Spectral Path Regression using directional Chebyshev harmonics for interpretable learning on tabular data without exponential scaling.
Analysis of geometric alignment cost in scientific foundation models for biology/physics, showing discrete tokenization degrades continuous geometry preservation.
Framework for uncertainty-aware foundation models on clinical data, addressing incomplete and irregular measurements in healthcare.
ClawArena benchmark for evaluating AI agents in dynamic environments with evolving information, contradictions, and implicit user feedback.
Graph-assisted retrieval framework for reasoning about defects in laser powder bed fusion manufacturing using structured scientific knowledge.
Framework using Temporal Behavior Trees to repair suboptimal trajectories before using them for robot control policy learning.
Analysis of token routing in Mixture-of-Experts models reveals three-phase training trajectory for load balance evolution.
Method for constrained model steering of LLMs addressing safety/privacy requirements via spectral subspace optimization.
Risk scoring system optimizing net benefit using sparse integer linear programming for high-stakes decision-making.
Study of vulnerabilities in large reasoning models when applying machine unlearning techniques to remove influence of specific data.
Federated RLHF method for fair LLM alignment across diverse human preferences without centralizing preference data.
Dual-step generative framework combining causal structure learning with tabular data synthesis using directed acyclic GANs.
Computer vision and neural network approach for classifying human brain activity from EEG data during hand movement.
Pipeline combining LSTM, synthetic data, and fine-tuning for EEG classification on implicit visual stimuli tasks.
Analysis of distributional reinforcement learning for complex domains like healthcare, addressing heterogeneous groups under uncertainty.
Tutorial on using flow- and score-based generative models for decision-making under distributional shift in operations research.
System for generating editable design variations using decoder-only language model with Creative Markup Language representation.
Theoretical analysis of Q-value iteration convergence in multi-agent Stackelberg games using control-theoretic perspective.
Research on aligning LLMs with human preferences using relative density ratio optimization without assuming specific preference models, improving statistical consistency.
Study questioning necessity of prompt selection in task-free online continual learning for non-stationary data streams.
Ablation framework to estimate contributions of central, peripheral, and temporal visual information to human decision-making in Atari games.
DP-OPD: differentially private on-policy distillation method for compressing LLMs on sensitive data while maintaining privacy guarantees.
MAVEN: mesh-aware volumetric encoding network for simulating 3D flexible deformation using graph neural networks on mesh structures.
Discrete Prototypical Memories approach for federated time series foundation models using LLMs while preserving data privacy.
Isokinetic Flow Matching introduces pathwise acceleration regularization to improve few-step sampling in flow-based generative models.
SLaB: sparse-lowrank-binary decomposition framework for efficient LLM compression maintaining performance at high compression ratios.
Multi-objective controllable language models framework enabling personalized alignment with varying human preferences beyond fixed reward optimization.
GAIN: multiplicative modulation technique for domain adaptation in LLMs, preventing catastrophic forgetting through feature re-emphasis.
Reproducibility study on spurious correlations and shortcut learning in DNNs, comparing frameworks for ensuring models use causally relevant features.
Revisits learning from equivalence queries model for modern ML systems like generative models and recommendation systems with periodic updates.
FlashSAC: off-policy reinforcement learning algorithm for stable, fast robot control in high-dimensional action spaces.
Detection method for free-riders in federated learning via simulated attack patterns, improving the WEF-based approach.