Position-Aware Sequential Attention for Accurate Next Item Recommendations
Improves sequential attention for recommendation systems by integrating position information into attention mechanism beyond additive embeddings.
Improves sequential attention for recommendation systems by integrating position information into attention mechanism beyond additive embeddings.
Proposes dual-model training framework inspired by neuroscience motivation states with alternating base and larger model activation.
Extends projection pursuit tree classifier with visual diagnostic methods for high-dimensional multi-class classification problems.
DEEPSYNTH benchmark evaluates LLM-based agents on complex multi-source information synthesis tasks beyond fact retrieval.
Implements tensor parallelism for selective state-space models on multi-GPU systems to scale SSM inference beyond single GPU limits.
Decomposes epistemic uncertainty in Bayesian deep learning into per-class contributions for safety-critical classification asymmetric costs.
Aletheia, a mathematics research AI agent powered by Gemini 3 Deep Think, autonomously solves 6 of 10 FirstProof challenge problems.
Combines off-policy and on-policy reinforcement learning for fast visual sim-to-real robotics training with minimal sample waste.
Proposes coalition-based partitioning approach for Shapley value attribution methods in explainable AI to resolve attribution conflicts.
Theoretical framework for disentangled representation learning when factors of variation are dependent rather than independent.
Proposes dynamic optimal transport minimization for discrete flow matching on categorical data with Kantorovich formulation.
Measurement-driven analysis of LLM inference energy and performance tradeoffs across workloads using GPU DVFS on 1B-32B parameter models.
Theoretical investigation of benign overfitting phenomenon in binary linear classification across over-parametrized neural networks.
Theoretical analysis of neural network-based optimal transport solvers using semi-dual adversarial formulations for generative modeling.
Theoretical analysis of differentially private shuffled gradient methods for convex ERM, addressing privacy-accuracy tradeoffs compared to standard DP-SGD.
Analyzes flaws in Integrated Gradients attribution method and proposes alternative path-based approach using model-induced geometry for better feature importance explanations.
Proposes Cauchy-Schwarz divergence for vision-language alignment to address distributional differences and alignment-uniformity conflicts in multimodal models.
Semantic Parallelism optimizes MoE LLM inference by co-scheduling model device placement and request routing, reducing communication overhead in multi-device serving.
Comprehensive survey of Federated Learning combined with LLM fine-tuning (FedLLM), covering privacy-preserving collaborative model adaptation methods.
Survey of trustworthy GUI agents built on LLMs, identifying execution gap challenges in real-world digital environment automation with irreversible actions.
RefLoRA improves LoRA fine-tuning of large models by identifying optimal low-rank factorizations to address convergence and performance degradation issues.
Analysis of performance asymmetry in Model-Based RL agents on Atari100k, showing dramatic variance across task types despite high average performance.
Wasserstein Barycenter Soft Actor-Critic algorithm improves sample efficiency in off-policy reinforcement learning via directed exploration.
CausalFM framework trains Prior-Data Fitted Networks as foundation models for causal inference via in-context learning on tabular data.
Frequency-domain occlusion method for interpreting time series neural networks, benchmarking frequency-based attribution approaches.
Efficient fine-tuning method for LLMs using entropy-based complexity detection to apply chain-of-thought reasoning selectively on difficult examples.
Theoretical analysis of transfer learning in infinitely wide neural networks under gradient flow, quantifying pretraining benefits.
One-Step Flow Q-Learning accelerates Diffusion Q-Learning for offline reinforcement learning by enabling single-step denoising without auxiliary modules.
Uncertainty Propagation Networks extend neural ODEs to model both state trajectories and uncertainty quantification in continuous-time systems.
MCTD-ME combines masked diffusion models with Monte Carlo Tree Search for protein design, addressing long-range dependencies and search space challenges.
Probabilistic Scenarios paradigm for time series forecasting generates finite scenario sets instead of samples to address computational and coverage limitations.
RHYTHM framework uses LLMs as spatio-temporal predictors with hierarchical temporal tokenization for human mobility prediction.
Polychromic objectives framework for reinforcement learning fine-tuning preserves policy diversity during RLFT to prevent mode collapse.
Recursive Self-Aggregation (RSA) test-time scaling method combines parallel and sequential inference to improve LLM reasoning capabilities.
Cautious Weight Decay (CWD) optimizer modification applies weight decay only to parameters aligned with optimizer updates.
TeamFormer proposes shallow parallel Transformer architecture with progressive approximation for efficient training and inference.
Latent-Augmented Discrete Diffusion (LADD) improves discrete diffusion models for fast language generation by modeling cross-token dependencies.
Framework for scalable AI oversight by partitioning complex multi-domain evaluation tasks among domain-specific human experts.
ContextPilot accelerates long-context LLM inference by enabling context reuse via KV-cache optimization for RAG and agent memory layers.
Framework for evaluating anomaly detection in 5G networks accounting for non-IID data and adaptive attackers.
Research on scaling laws for LLM pretraining with different optimizers beyond AdamW, examining new optimizers like Muon, Shampoo, and SOAP.
Research on optimization algorithms for LLM reinforcement learning, comparing SGD vs Adam optimizers and their effectiveness in RL training phases.
AceGRPO method for autonomous ML engineering agents using adaptive curriculum and group relative policy optimization to overcome behavioral stagnation.
VESPO algorithm for stable off-policy LLM training via importance sampling with variance reduction to prevent policy divergence and collapse.
Vector quantization compression for MoE LLMs using KLT-guided SVD and bias-corrected quantization for ultra-low-bit model deployment.
Multi-tenant ML serving system handling seamless model updates while maintaining decision thresholds across clients with distribution shifts.
Pawsterior framework for simulation-based inference using variational flow matching with structured domain constraints for bounded parameters.
Analysis of worker-level optimization misalignment in data-parallel LLM fine-tuning despite parameter synchronization, termed silent inconsistency.
Numerical analysis showing GLU variants scale asymptotically faster than MLPs, explaining architectural dominance in frontier LLMs.
Study of multilingual data curation across 13 languages identifying interference patterns and optimal training strategies for 20-trillion-token dataset.