Memorization to Generalization: Emergence of Diffusion Models from Associative Memory
Dense associative memories framework showing emergence of diffusion models from memory storage, bridging memorization and generalization in neural networks.
Dense associative memories framework showing emergence of diffusion models from memory storage, bridging memorization and generalization in neural networks.
VERINA benchmark for evaluating LLM code generation with jointly generated specifications and proofs, addressing correctness verification challenges.
Coded robust aggregation method for distributed learning resilient to Byzantine attacks, improving gradient aggregation in federated settings.
Graph model merging technique for combining GNN models pre-trained on different domains with distribution discrepancy to create generalized models.
Knowledge graph embedding methods for feature learning on large-scale graphs as external knowledge for downstream ML tasks, optimizing beyond link prediction.
Dynamic weighting approach combining supervised fine-tuning and reinforcement learning for LLM post-training, unifying on-policy and off-policy learning paradigms.
Tool-augmented LLM agents trained with synthetic code environments via RL to improve generalization on tool-use tasks, addressing brittleness with new tools and unseen workflows.
LANCE: low-rank activation compression method for efficient on-device continual learning, reducing memory costs during backpropagation in resource-constrained environments.
NanoFlux: adversarial dual-LLM framework for generating targeted training data to improve reasoning, achieving strong results with <200 examples through competitive Attacker-Defender dynamics.
Attribution-Guided Decoding uses interpretability to improve LLM instruction-following and factual accuracy.
PolyGraph Discrepancy metric provides absolute performance measure for graph generative models.
Tree search guidance method for controllable graph generation with diffusion models.
Theoretical bounds connecting Jensen-Shannon and Kullback-Leibler divergences for representation learning.
Cluster-PFN extends Prior-Data Fitted Networks to Bayesian clustering with uncertainty quantification.
AGRAG improves graph-based RAG for LLMs by addressing hallucination, reasoning, and answer quality issues.
FedSDWC applies causal learning to federated learning for handling out-of-distribution data shifts.
MAVA accelerates masked auto-regressive diffusion inference for practical reinforcement learning applications.
Derives tail distribution bounds for regret in optimism-based reinforcement learning algorithms.
Empirical comparison of flow matching variants with diffusion models for privacy-preserving tabular data synthesis.
PRISM complex-valued encoder explores phase relationships in semantic representations of neural sequence models.
Method to identify and measure social biases in text-to-image diffusion models via automated prompt search.
Domain-adapted LLM fine-tuned for educational QA in space weather and heliophysics.
TRACE framework uses autoregressive density estimation for causal discovery in single event sequences.
MetaDOAR meta-controller applies multi-agent reinforcement learning to large-scale cyber-network security games.
TRC² architecture enables LLMs to continually learn and adapt without catastrophic forgetting through specialized decoder design.
Framework combining ML and contextual stochastic optimization for transit network design under demand uncertainty.
AOI framework enabling LLM agents to improve from failed cloud diagnosis trajectories in SRE automation with safety constraints.
KV cache optimization using low-dimensional attention selection to reduce transformer memory with O(log N) key dimensions.
Study examining many-shot prompting for test-time LLM adaptation, analyzing reliability and limits of in-context learning scaling.
Active feature acquisition method for biomedical applications optimizing measurement selection under temporal and cost constraints.
Theoretical analysis proving attention sinks are functionally necessary in softmax transformers for certain tasks.
Research addressing multimodal model underperformance in context-aided forecasting via improved context quality assessment.
Machine learning method using hypergraph pre-training to improve atrial fibrillation prediction in stroke patients.
Interactive benchmark environment for synthesizing flat-foldable origamis, testing AI systems' planning and causal reasoning in physical domains.
Neural compression framework using SIREN auto-decoders for high-fidelity compression of multi-structural seismic velocity models.
Graph-based verifier for LLM task planning that identifies and corrects hallucinations and flaws in agent-generated plans.
Post-hoc model-agnostic explanation method using informative perturbation selection for uncertainty-aware interpretability of black-box ML models.
Convergence analysis of Muon optimizer under heavy-tailed noise for nonconvex optimization in large-scale deep neural network training.
Analysis showing wider beam search in LLMs can degrade output quality due to overestimation bias in noisy scorer outputs, with theoretical grounding.
Large-scale benchmark for AI agents combining partial observability, game-theoretic reasoning, and long-horizon planning in Pokemon battle environment.
Physics-informed neural networks and neural operators for simulating EUV electromagnetic wave diffraction in lithography mask applications.
Annotation-free method for reconstructing controllable 3D Gaussian splats of articulated objects from monocular video using flow derivatives.
Data synthesis engine using scene graphs to improve compositional generalization and semantic alignment in text-to-image generation models.
CHARM method calibrating reward models using Chatbot Arena scores to mitigate model preference bias, improving alignment of LLMs through RLHF.
PhysioOmni foundation model for multimodal physiological signals handling arbitrary missing modalities across EEG, ECG, EOG, EMG for healthcare and brain-computer interfaces.
BiomedSQL benchmark for text-to-SQL generation requiring scientific reasoning over biomedical knowledge bases, evaluating LLM capability for complex analytical tasks.
DP-Powered LLMs for privacy-preserving radiology report classification, enabling differential privacy in healthcare diagnosis and abnormality classification workflows.
TempCore benchmark analyzing whether video QA models genuinely require temporal frame selection, introducing Frame Selection Sensitivity metric for VLM diagnostic evaluation.
Comparison of statistical and logic-based XAI techniques for interpreting ML security alerts in 5G intrusion detection systems, enabling actionable incident response.
ERGO framework for efficient high-resolution image processing in vision-language models using coarse-to-fine reasoning pipeline with two-stage visual token reduction.