QSIM: Multi-agent reinforcement learning method addressing Q-value overestimation via action similarity weighting. Value decomposition improvement.
Research on diffusion models for end-to-end autonomous driving in real-world settings. arXiv paper exploring decision-making applications.
TARAZ benchmark for evaluating cultural competence of LLMs in Persian with short-answer format and morphological analysis.
Unsupervised continual learning framework for amortized Bayesian inference handling sequential data and distribution shifts.
SPD Learn Python library for symmetric positive definite matrix neural networks in neural decoding with geometric deep learning.
OmniGAIA benchmark for evaluating omni-modal AI agents with unified vision, audio, and language perception capabilities.
SIGMA system applying LLMs to multi-task recommendation at scale, handling diverse business requirements beyond traditional next-item prediction.
Parameter-efficient fine-tuning combining multiple domain expert models for visual adaptation tasks via prompt tuning.
Tree-based system for analyzing semi-structured documents with mixed content types enabling question-answering over complex layouts.
LLM agent framework (SALA) for evaluating deanonymization risks through stylometry-assisted analysis of textual data with interpretable pipeline.
Study of information-theoretic limits in multimodal LLMs showing modality-specific information is discarded during decoding despite surviving encoding layers.
arXiv paper: Fairness-aware mixed-precision quantization for neural network compression in medical imaging, explicitly addressing algorithmic fairness during model compression.
arXiv paper: Theoretical analysis of fine-tuning effects on in-context learning in linear attention transformers, balancing downstream and few-shot task performance.
arXiv paper: Plug-and-play diffusion framework with ADMM for medical image reconstruction addressing memory issues in PnP solvers.
arXiv paper: Large-scale app store search ranking system augmented with LLM-generated textual relevance judgments to address scarcity of expert labels.
arXiv paper: Zeroth-order optimization for leader-follower Stackelberg control in combinatorial congestion games with discrete strategy selection.
arXiv paper: Training-free hierarchical manifold guidance for dataset distillation using diffusion models without requiring gradient computation.
arXiv paper: Zero-shot and one-shot adaptation of small language models for leader-follower role assignment in human-robot interaction with resource constraints.
Analyzes parameter-efficient fine-tuning for continual learning using Neural Tangent Kernel theory to understand model adaptation and catastrophic forgetting mitigation.
Physics-inspired neural framework using graph neural networks to solve large-scale graph coloring combinatorial optimization problems near algorithmic phase transitions.
Proposes label unlearning method for vertical federated learning using representation-level manifold mixup to enable privacy-preserving model unlearning.
Model-agnostic explanation technique integrating concept-based approaches with diverse explanation forms beyond attribution methods.
Reinforcement learning framework for aligning few-step diffusion models with downstream objectives using stepwise policy optimization.
Neuro-symbolic framework automating discovery of analytical solutions to differential equations using formal grammars and continuous search.
Continual learning method using sample compression theory to provide computable guarantees while avoiding catastrophic forgetting.
Compression technique for large language models using global rank and sparsity optimization with layer-wise weight allocation.
Large-scale benchmark for evaluating machine learning solvers on combinatorial optimization using real-world industrial datasets.
Research on whether language models can learn to evade latent-space safety monitors, with implications for LLM alignment and security.
Statistical framework for evaluating LLM-as-a-judge systems, addressing reliability and bias issues in automated LLM output evaluation.
Silent Gradients approach for training VAEs by restricting decoder architecture to reduce gradient estimation variance.
Online time series forecasting method using feature adjustment to handle distribution shift in sequential data deployment.
rBridge framework enabling small proxy models (≤1B params) to predict reasoning performance of larger LLMs, optimizing dataset scaling.
Theoretical analysis of softmax attention mechanisms in LLMs through single-location regression task, addressing why softmax dominates alternative activation functions.
Study of compute-optimal allocation between full-precision and quantization-aware training phases for neural networks.
Dynamic safety monitoring system for LLMs that adaptively adjusts computation based on input difficulty.
Method for improving masked diffusion models by learning better unmasking policies beyond rule-based position scheduling.
Imitation learning approach for training models with multiple acceptable answers from limited correct demonstrations.
Method for learning probability distributions on simplexes via smooth bijections and Aitchison geometry.
UniQL framework combining quantization and low-rank compression with adaptive on-device pruning for edge LLM deployment.
Post-training method achieving 99.6% attention sparsity without performance loss for mechanistic interpretability research.
WebGym: largest open-source environment with 300k tasks for training visual web agents on realistic websites.
Confidence-Variance theory framework for improved pseudo-label selection in semi-supervised learning beyond fixed thresholds.
Method for optimizing interaction between feature alignment and target fitting in cross-modal model fine-tuning.
Framework for learning Hamiltonian flow maps to enable stable large-timestep molecular dynamics simulations.
Data-free early stopping framework for federated learning using task vector growth rate monitoring.
EPIAGENT agentic framework that automatically synthesizes and calibrates epidemiological simulators via iterative program synthesis.
Framework for score-based density ratio estimation addressing path-variance issues in practical training objectives.
Theoretical analysis of phase transitions in neural network feature learning on multi-index models.
Technique to detect misbehaviors and hallucinations in large vision-language models using evidential uncertainty quantification.
Method for training ensemble models that quantify epistemic uncertainty via distributionally robust optimization.