Pro-KLShampoo: Projected KL-Shampoo with Whitening Recovered by Orthogonalization
Combines Kronecker-factored preconditioning with gradient momentum orthogonalization for LLM pre-training optimization.
Combines Kronecker-factored preconditioning with gradient momentum orthogonalization for LLM pre-training optimization.
Proposes adaptive task graphs to improve coordination efficiency in LLM agent teams, reducing error propagation and resource waste.
Measures evaluation-context divergence in open-weight LLMs using paired prompts to detect alignment pipeline-specific behavioral differences.
Fine-tunes small language models for Windows event log analysis with actionable remediation, addressing privacy and resource constraints of LLMs.
Uses LLMs with evolutionary algorithms to design heuristics for coupled optimization problems across multiple interdependent subproblems.
Analysis of feedback loops between humans and LLMs as coupled dynamical systems, studying knowledge collapse from recursive AI-generated training.
Decision-theoretic framework for LLM cascades analyzing optimal deferral thresholds between models to balance cost and quality tradeoffs.
Topological analysis of grokking phenomenon in neural networks using persistent homology on embedding matrices.
Memory-efficient gradient computation method for evaluating adversarial robustness of diffusion and Langevin-based defense mechanisms.
Generative modeling framework extending flow matching with arbitrary auxiliary distributions for flexible trajectory generation.
Framework using visual explanations as regularization to improve model robustness under distribution shifts, addressing spurious correlations interpretably.
Continuous-time distribution matching for accelerating diffusion model distillation with sparse supervision.
On-policy distillation method addressing token-level learning between exploitation and imitation with improved gradient properties.
Convolutional learning framework for infinite-dimensional signals on manifolds using Hilbert bundles and cellular sheaves.
Dimensionless control parameter predicting Mixture-of-Experts model collapse, unifying expert ecology across vision and language.
Study of LLM agents failing to maintain structural constraints in backend code generation, analyzing non-functional requirement violations.
Bayesian hyperparameter optimization addressing acquisition estimation noise through orthogonal sampling.
Off-policy evaluation framework using recursive reweighting and moment matching for finite-horizon MDPs.
Custom SIMD kernels for running ternary neural networks on consumer CPUs, enabling LLM inference without GPUs.
RL method for continuous spaces discovering value-preserving structures via operator-guided invariance learning.
PAC-private zeroth-order mechanism for fine-tuning LLMs via sign quantization with strong membership inference attack resistance.
Mechanistic study of inference dynamics in tabular transformer foundation models, analyzing layerwise prediction emergence.
RL framework for Benders decomposition that adaptively selects cuts using neural network policies for stochastic programming.
Study of implicit reward overfitting in reinforcement learning with verifiable rewards, analyzing low-rank dynamics in model reasoning.
Hierarchical latent diffusion language model for text generation using non-autoregressive approach with improved efficiency and semantic modeling.
Coordination-aware evaluation metrics for cooperative multi-agent reinforcement learning with process-level diagnostics.
GONO optimization framework identifying decoupling between directional alignment and loss convergence in deep learning.
Neural architecture for graph matching that estimates graph edit distance using GNNs with focus on encoder geometry.
Vision-language pretraining method combining DINOv3 distillation with ranking consistency for improved CLIP models.
Multi-agent reinforcement learning approach for embodied navigation using specialized agents for different sensory modalities.
Unified self-distillation framework for LLMs enabling adaptation without external teachers through self-generated trajectories and supervision.
Language model agent reconstructs security vulnerabilities from Linux binary patches using local binary evidence without source code access.
Open-source AI agent for computational fluid dynamics discovery combining LLMs with physics simulators to automate scientific discovery loops.
Mechanistic explanation of attention sink phenomenon in LLMs, tracing root causes to variance discrepancy and dimension disparity in self-attention.
Theoretical analysis explaining when sign-based optimization algorithms like SignSGD outperform vanilla SGD for training large models.
Reinforcement learning approach enabling agents to recursively spawn and delegate sub-tasks for divide-and-conquer inference-time scaling.
Method for generating concept-based causal explanations of vision models using abductive and contrastive reasoning.
Framework for training LLM agents for long-horizon decision making using strategic trajectory abstraction to improve exploration and credit assignment.
Comprehensive benchmark study on multimodal domain generalization methods, examining whether performance improvements reflect genuine algorithmic progress.
Retrieval-augmented agent system treating retrieval as active expert navigation rather than black-box queries for knowledge base interaction.
Framework for validating comparative LLM safety scores without labeled benchmarks, enabling safety evaluation across diverse languages and domains.
Study showing that LLM fine-tuning with the same optimizer as pretraining reduces catastrophic forgetting while maintaining performance on new tasks.
Method for generating valid, challenging mathematical problems using LLM verifiers and self-play, enabling automated problem creation for LLM training.
Training-free bias mitigation method for GUI grounding in agents, improving performance on complex screen interaction tasks.
Knowledge distillation method for multi-modal models focusing on transferring modality-level relationships from teacher to student networks.
Formal game-theoretic framework for safety evaluation of AI deployment protocols through red-teaming exercises.
Framework for aligning LLM-based agents with human preferences through goal inference from multi-turn dialogue, addressing challenges in collaborative agent interactions.
Benchmark for evaluating differentially private text generation methods using LLMs, enabling secure sharing of sensitive datasets across institutions.
Research on how reinforcement learning enables LLMs to develop multi-step reasoning capabilities, using statistical physics framework to explain emergence of slow thinking.
Learning reasoning reward functions for LLMs via inverse reinforcement learning from expert demonstrations, avoiding manual reward specification.