Towards General Preference Alignment: Diffusion Models at Nash Equilibrium
Preference alignment for text-to-image diffusion models using game-theoretic approach at Nash equilibrium instead of reward modeling.
Preference alignment for text-to-image diffusion models using game-theoretic approach at Nash equilibrium instead of reward modeling.
Analysis connecting power sampling, self-reward RL, and self-distillation for improving LLM reasoning and performance.
Model-based reinforcement learning for data-efficient HVAC control in buildings using counterfactual building models.
HeterSEED decouples semantics from structure in heterogeneous graph learning to handle heterophilic graphs with dissimilar connected nodes.
Queueing-theoretic framework for analyzing LLM inference stability under KV cache memory constraints and computation bottlenecks.
FAAST enables forward-only test-time adaptation by analytically compiling labeled examples into fast weights without backpropagation overhead.
Threshold-guided optimization for visual generative models using implicit reward comparisons to enable scalar rating-based preference alignment.
ITBoost improves gradient boosting robustness to noisy labels by using information-theoretic evaluation of sample reliability.
SPHERE mitigates plasticity loss in mixture-of-experts networks for continual deep reinforcement learning by preserving spectral properties.
Re-evaluation of attention-based programming knowledge tracing models reveals sensitivity to implementation details and experimental protocols.
OSAQ addresses weight outliers in post-training LLM quantization to reduce model size and inference latency through selective absorption.
KFCA mechanism rewards federated learning client contributions without ground truth labels, evaluated on LLM adapter tuning and PCB inspection tasks.
AxMoE characterizes the impact of approximate multipliers on mixture-of-experts DNN architectures for efficient edge inference.
Personalized Thinking Model for AI-supported education capturing learner behavior, cognition, and metacognition in hierarchical structure.
Bilinear Mamba-Koopman Neural MPC for model predictive control with time-varying dynamics from historical data.
Uses graph neural networks and hypergraph representation learning for unsat-core prediction in Boolean satisfiability problems.
Improves FMQA black-box optimization through better initial training data design considering marginal bit coverage in one-hot encoding.
Federated label distribution learning approach handling heterogeneous annotation quality across clients in privacy-sensitive settings.
Analyzes phase transitions in diffusion models through symmetry breaking and nonlocality during generation dynamics.
Continual learning approach for physics-informed neural operators using replay to address out-of-distribution performance degradation.
ALL-IN method enables graph neural network transferability across datasets with different input feature spaces, enabling graph foundation models.
QpiGNN introduces quantile-free uncertainty quantification for graph neural networks without requiring costly resampling or post-hoc calibration.
Uncertainty-aware DPO method for reducing hallucination in multimodal LLMs by allocating fine-grained supervision to vision tokens.
Theoretical analysis proving limitations of symmetric spectral methods for diagnosing attention failures in language models.
Geometric analysis of sampling errors in language models, showing curvature of token embedding geometry correlates with semantic properties.
Delta-Code Generation: LLMs generate compact unified diffs to refine baseline architectures rather than complete models from scratch.
In-context learning approach for tabular data generation balancing quality and privacy without model-specific training.
Reinforcement learning approach for compositional generalization using outcome-level optimization instead of token-level training.
Learning method for avoiding undesired futures using order structure instead of graph structure in forecasting scenarios.
KernelBench-X benchmark evaluating LLM-generated GPU kernels across 176 tasks, identifying where LLM kernel generation capabilities break down.
EP-GRPO algorithm addressing credit assignment failures in reinforcement learning for LLM reasoning, improving upon Group Relative Policy Optimization.
Method for extending LLM capabilities to new skills using soft tokens (skill neologisms) to avoid catastrophic forgetting while maintaining expressiveness beyond context limits.
Conceptor-based steering method for controlling LLM behavior using soft projection matrices that preserve multidimensional concept subspaces.
Self-Induced Outcome Potential method for turn-level credit assignment in long-horizon LLM agents without process-level verifiers.
Theoretical comparison of in-context learning vs agentic learning for task approximation under neural network realizability constraints.
Training method for decentralized learning where heterogeneous nodes learn to compose predictions effectively at inference time.
Graph-SND metric for measuring behavioral diversity in multi-agent RL using sparse graph aggregation instead of complete pairwise distances.
CuBridge uses LLMs to understand and generate high-performance CUDA attention kernels with improved correctness and efficiency.
Empirical study showing optimal predictive encoders fail at causal fidelity, with analysis of 2695 neural networks on linear-Gaussian dynamics.
Self-distillation method for on-policy LLM training using preference-based reward regularization beyond KL matching.
Imitation learning approach for plasma stabilization control in Vlasov-Poisson systems with partial state observability constraints.
ORDERED algorithm for unsupervised domain adaptation that reduces variance in discrepancy estimation through optimal data reordering.
Memini system for continual knowledge updating in LLMs using multi-timescale memory dynamics inspired by biological memory mechanisms.
Theoretical framework for analyzing regret distribution in multi-armed bandits and episodic RL with probabilistic guarantees across confidence levels.
RL method for improving sample efficiency in agentic tasks by controlling pass-rate to optimal 50% for maximizing reward signal informativeness.
Theoretical study of signal propagation in finite-width linear recurrent models, analyzing approximation accuracy as sequence depth and width grow jointly.
Analysis of jailbreak attack difficulty on LLMs, showing random search effectiveness challenges assumptions about prompt structure necessity.
Low-cost black-box method to detect LLM hallucinations by treating model as dynamical system and analyzing embedding manifolds.
Case study of AI tools assisting high school students in financial forecasting research project with human-AI co-mentorship.
Mechanistic interpretability study of transformer time series forecasting using sparse autoencoders, questioning superposition hypothesis.