Symmetry Guarantees Statistic Recovery in Variational Inference
Research on symmetry properties in variational inference guaranteeing statistic recovery in approximate target density distributions.
Research on symmetry properties in variational inference guaranteeing statistic recovery in approximate target density distributions.
Study of RLVR fine-tuning effectiveness for LLMs in low-data and low-compute regimes, evaluating practical applicability in resource-constrained settings.
Bandit algorithm framework for online learning on graph-structured payoffs with applications to content recommendation systems.
DESPITE benchmark evaluates safety risks of LLMs used as planners for robotic systems across 12,279 tasks, finding planning ability doesn't guarantee safety.
Method combining learned safety filters with adaptive conformal inference to guarantee safety in control systems with learned components.
Deep operator network approach for learning Riccati equation solutions to replace repeated numerical solving in LQR control problems.
Framework for stable integration of reinforcement learning with uniform discrete diffusion models, addressing training instability in GRPO application.
Unsupervised method for improving LLM verification by ensembling verifiers without labeled data, useful for training and deployment.
Gumbel-Softmax based scalar quantization technique for LLM weight compression achieving sub-4-bit accuracy without vector quantization overhead.
Method for controlling protein conformational states in OpenFold3 via latent variable manipulation to capture biologically relevant alternate conformations.
Benchmark comparing cloud-based and open-source LLMs on causal loop diagram extraction and system dynamics discussion tasks.
Empirical study examining whether neural networks trained on different modalities converge to shared representations, challenging the Platonic Representation Hypothesis.
Large-scale multimodal multilingual benchmark of Olympiad-level math problems for evaluating reasoning in LLMs and embedding-based retrieval systems.
Framework for enforcing hard convex constraints on neural network outputs during training and inference for robotics applications.
Theoretical analysis of score estimation optimization and generalization in diffusion models trained via gradient descent on neural networks.
Method for pruning Mixture-of-Experts layers in LLMs to reduce memory requirements while maintaining performance through condensation rather than removal.
Technique using AI feedback to improve text-to-video model generation, specifically addressing dynamic object interactions and physics violations.
Method for real-time LLM personalization via hypothesis reweighting that adapts model outputs to individual user preferences with minimal labeled examples.
Research on efficient uncertainty estimation in LLMs using single-sequence measures instead of computationally expensive multi-sequence methods for trustworthiness evaluation.
Theoretical study of score smoothing in diffusion models enabling interpolation and generation of new data.
Evaluation of multi-modal LLM prompting strategies for zero-shot handwritten document transcription without fine-tuning.
Two-stage pruning method for reducing LLM parameters with regularization to preserve knowledge during deployment.
Parameter-efficient fine-tuning method using column space projection as alternative to LoRA for foundation models.
Theoretical analysis of gradient descent dynamics in deep ReLU networks escaping saddle points with low-rank bias.
Analysis of bias in vision-language models on objective visual tasks like counting and identification.
Parameter-efficient fine-tuning method using rank reduction for LLM reasoning tasks with reduced computational cost.
Study of vision-language model privacy concerns and degenerate outputs after machine unlearning.
Framework for optimizing multi-agent search systems coordinating LLM agents with tools using reinforcement learning.
Method for quantifying uncertainty in LLM-based systems and prompt sensitivity without access to model internals.
LLM-based time series forecasting using reinforcement learning and reasoning approach instead of fast pattern extraction.
Research on limitations of fidelity-based explanations in XAI, introducing linearity score to measure linear decodability of neural network behavior.
Demonstrates 8:16 semi-structured sparsity for LLM compression, outperforming N:M sparsification methods while handling outlier weights.
Introduces EvoCoT to address exploration bottleneck in RL with verifiable rewards for LLM post-training, improving reasoning on hard problems without teacher models.
Studies multi-step reasoning in LLMs using cellular automata framework to understand how models learn and execute reasoning beyond memorization.
Negative regularization technique addressing over-shrinkage in small-data regression through controlled anti-shrinkage.
Bi-LoRA integrates Sharpness-Aware Minimization with LoRA for efficient fine-tuning of large models with improved generalization.
RefineStat method for probabilistic program synthesis combining small language models with statistical exploration under domain constraints.
Low-rank orthogonalization technique for matrix optimization improving foundation model training efficiency.
PiERN architecture enables token-level routing to integrate high-precision numerical computation within LLM reasoning.
Flow Marching algorithm combining neural operators with flow matching for generative PDE foundation models.
Theoretical analysis establishing central limit theorems for asynchronous Q-learning with convergence rate bounds.
CaTS-Bench benchmark for evaluating language models on time series captioning and temporal reasoning across 11 domains.
Framework for LLM reasoning via reinforcement learning that optimizes semantic space exploration rather than token-level proxies.
Method for improving robustness of LLM unlearning by simplifying optimizers, addressing fragility of privacy/safety interventions.
OptunaHub platform for distributed black-box optimization components with unified interface, supporting AutoML and hyperparameter tuning.
Theoretical convergence analysis of Graph Neural Differential Equations combining GNNs with continuous-depth architectures.
RACE Attention proposes linear-time attention layer to handle contexts exceeding 4M tokens, addressing quadratic complexity of softmax attention.
Mechanistic analysis of training failures with low-precision transformers and Flash Attention, revealing root causes of loss explosion.
Research on full Gauss-Newton preconditioning for LLM training, establishing practical iteration complexity bounds up to 150M parameters.
PaTaRM: Reward modeling approach bridging pairwise and pointwise preference signals for improved RLHF alignment of LLMs with human preferences.