Study showing LLMs struggle to effectively reason over text-attributed graphs despite their capabilities in natural language and cross-modal understanding.
Bayesian approach to offline reinforcement learning using posterior over world models without explicit conservatism constraints.
Method to improve activation sparsity in LLMs by addressing representational instability caused by suppressing hidden activations.
NRGPT reframes GPT inference as energy-based model exploration, proposing minimal architectural modifications to unify with EBM framework.
Cascaded Flow Matching approach for generating tabular data with mixed discrete and continuous features using diffusion models.
Expert-Sample method leverages fine-grained MoE routing patterns for test-time scaling in LLMs without temperature tuning.
Analysis of GCG adversarial attacks on LLMs examining token position beyond suffix-based approaches for jailbreak robustness evaluation.
Riemannian MeanFlow framework reduces neural network evaluations needed for diffusion models on Riemannian manifolds for generative modeling.
SOAR algorithm for multi-armed bandits with heterogeneous noise sources across multiple data sources with adaptive source selection.
RAT+ introduces recurrence-augmented attention enabling sparse dilated inference while maintaining accuracy, reducing FLOPs and KV cache with flexible configuration.
Comparative analysis of UMAP against PCA, Kernel PCA, SIR, and t-SNE for dimensionality reduction techniques.
Method using spectral heat diffusion to discover continuous abstraction levels in knowledge graphs and GraphRAG systems.
Analysis of alignment evaluation gaps showing concept detection differs from routing behavior using Chinese language model case study.
Hierarchical RL approach for allocating limited public health resources across asynchronous disease outbreak clusters.
Learned neural improvement policy for TSP that iteratively applies local modifications conditioned on candidate solutions.
LLM-guided RL approach for drug lead optimization that ensures synthesizable molecular modifications through action space constraints.
Question augmentation framework using partial solutions as hints to optimize RL training efficiency for LLM reasoning.
Unified convergence theory for adaptive first-order optimization methods including AdaGrad, AdaNorm, Shampoo, and Muon.
Unified framework for preference optimization revealing common update dynamics and preventing degradation of chosen responses.
ARFBench benchmark evaluating multimodal foundation models on time series anomaly understanding for software incident response.
Training method using weak supervision to prevent LLM sandbagging and elicit best performance without full output verification.
Investigation of internal confidence signals and error detection mechanisms in LLMs using decision neuroscience frameworks.
Lightweight online RL fine-tuning method for vision-language-action models using RL tokens for robot manipulation.
Intrinsic reward method using entropy centroids for test-time compute scaling and response selection in large language models.
Adaptive multimodal networks that handle runtime variations in modality quality, input complexity, and compute resources.
Method to achieve learning rate transfer across model sizes in Normalized Transformers using alignment exponents.
Reinforcement learning algorithms using data-driven Koopman operator to linearize nonlinear dynamics for tractable control.
DCT-based approach to improve Vision Transformer efficiency and initialization for self-attention mechanisms.
Comprehensive review of bias sources, evaluation methods, and mitigation strategies in large language models.
Post-training method using latent control to adapt frozen goal-conditioned policies without discrete text prompts.
Research paper analyzing internal representations and mechanisms in LLMs, addressing theoretical disagreements about how these models function.
Mathematical analysis of Mixture of Experts networks using mean-field theory and gradient flow, studying convergence properties as expert count increases.
Theoretical analysis of fundamental statistical limits in aligning LLMs with diverse human preferences, connecting preference aggregation to game theory concepts.
DiffMI diffusion-based model inversion attack on face recognition systems that recovers identity information from embeddings without iterative optimization.
TokenWeave technique for efficient distributed LLM inference via compute-communication overlap in tensor parallelism, reducing 20% overhead.
ML-Agent framework for autonomous ML engineering using reinforcement learning to improve LLM agents' ability to learn from execution trajectories.
Benchmark evaluation of multimodal foundation models (GPT-4o, Gemini, Claude, Llama) on standard computer vision tasks beyond question answering.
Analysis revealing low intrinsic dimensionality and redundancy in electronic structure datasets, reducing computational requirements for ML model training.
Method for zero-shot geospatial reasoning in vision-language models using indirect rewards from metadata to overcome supervision scarcity.
MIST foundation models family for molecular property prediction and chemical space discovery, trained on large unlabeled datasets for materials innovation.
Memory-augmented framework enabling LLM agents to learn classification functions from labeled examples without parameter updates using semantic and episodic memory.
Active learning framework for LLM classification systems minimizing costly human feedback while maintaining error guarantees through efficient labeling strategies.
PORTool algorithm for training tool-use LLM agents with improved credit assignment through importance-aware policy optimization on multi-tool reasoning tasks.
Method for reducing mode collapse in text-to-image diffusion models through noise optimization rather than guidance mechanisms.
Theoretical analysis of AdamW-style Shampoo optimizer convergence rates, unifying preconditioning approaches with nuclear norm guarantees.
Research on chain-of-thought reasoning failures in LLMs when reasoning steps exceed training distributions, analyzing internal mechanisms and proposing improvements.
Token Sparse Attention reduces quadratic complexity of LLM long-context inference via dynamic layer-wise token selection mechanism.
Framework for building and optimizing multi-agent conversational shopping assistants with evaluation metrics for multi-turn interactions.
Black-box domain adaptation method using dual-teacher distillation and subnetwork rectification to handle semantic gaps without source data.
Agent factory pipeline using general-purpose coding agents to autonomously optimize hardware designs from high-level specifications.