Self-Pruned Key-Value Attention: Learning When to Write by Predicting Future Utility
Self-Pruned Key-Value Attention mechanism reduces KV cache size in transformers by predicting future utility for efficient long-sequence generation.
Self-Pruned Key-Value Attention mechanism reduces KV cache size in transformers by predicting future utility for efficient long-sequence generation.
Comparative study of ML approaches for financial distress prediction under class imbalance constraints using classical and neural methods.
SurF generative model for forecasting irregular multivariate time series using Time Rescaling Theorem as learnable bijection.
Study on fairness and calibration in toxicity detection models using training interventions and safety mechanisms.
Research challenging cosine similarity as a metric for assessing layer relevance in LLMs, proposing alternatives for mechanistic interpretability.
Fleet of small foundation models for geospatial hydrologic analysis enabling accessible agentic environmental reasoning systems.
Reinforcement learning approach for tool-calling LLM agents reasoning over healthcare FHIR resource graphs for clinical queries.
Metacognitive framework leveraging LLM self-monitoring signals to control test-time inference and improve problem-solving.
Principled scaling rules for mixture-of-experts architectures analyzing hyperparameter relationships with network and expert dimensions.
Parameter-efficient finetuning method optimizing prefill-only updates to improve inference throughput for personalized LLMs.
Identifies and diagnoses training-inference mismatches in LLM reinforcement learning from implementation differences in token probabilities.
Empirical evaluation of quantum entanglement benefits in decentralized multi-agent reinforcement learning with variational quantum policies.
Active learning approach to improve pairwise ranking prompting reranking from LLMs with noisy and intransitive judgments.
Evaluation of AI-generated text detection methods' resilience to paraphrasing attacks across fine-tuned and classifier approaches.
Router for LLM agents selecting among functionally equivalent tool providers based on latency and quality tradeoffs.
Framework for predicting and optimizing energy consumption during multi-GPU LLM inference without expensive profiling.
Auditing protocol to verify that deep neural network heatmap explanations faithfully reflect actual model decision drivers in industrial inspection.
Analysis of how computation propagates through transformer layers by studying residual stream dynamics and spectral geometry in LLMs.
MetaMoE unifies independently trained domain-specialized experts into privacy-preserving MoE using public proxy data for federated settings.
KV-cache compression study with diversity-penalty survivor method for efficient LLM inference on long-form reasoning tasks.
Mixed gradient policy optimization for hybrid discrete-continuous action spaces in robotics and control problems.
Domain adaptation framework leveraging expert textual descriptions as language-induced priors to prevent negative transfer in cold-start scenarios.
Matrix-Space RL reuses local transition geometry through positive semidefinite matrix descriptors for compositional generalization.
Continuous semantic alignment framework for GUI critic models in test-time scaling for generalist AI agents beyond binary classification.
Dynamic Latent Routing composes optimal sub-policies temporally for MDPs and proposes LLM post-training method based on General Dijkstra Search.
AIM-DDI predicts drug-drug interactions using model-agnostic multimodal integration with focus on unseen-drug generalization.
Exemplar Partitioning constructs interpretable feature dictionaries from LLM activations using Voronoi partitioning with 1000x fewer tokens than sparse autoencoders.
Distributionally robust multi-task reinforcement learning addresses imbalanced learning across tasks via adaptive task sampling.
RQ-MoE combines residual quantization with mixture-of-experts for efficient input-dependent vector compression of embeddings.
MoRe proposes modular representations for continual learning on sequential data with minimal catastrophic forgetting.
LoMETab extends rank-1 ensembles to rank-r multiplicative factorizations for improved tabular deep learning performance.
Coherent Coordinate Descent optimizes zeroth-order optimization for memory-constrained on-device learning without backpropagation.
Optimal Pattern Detection Tree introduces interpretable rule-based classification model for pattern discovery in healthcare and maintenance domains.
Multi-agent exploration strategy for imperfect-information games like StarCraft using data-augmented game starts to accelerate policy gradient learning.
NodeSynth generates socially aligned synthetic data for AI model evaluation using a fine-tuned taxonomy generator anchored in real-world evidence.
Novel OOD detection method using class-wise Mahalanobis distance variance for neural networks in safety-critical applications.
Counterfactual time series forecasting method incorporating textual conditions for future events influence.
Federated actor-critic framework for collaborative policy training with shared representations and personalized local policies.
FrontierSmith: Method for synthesizing open-ended coding problems at scale to train stronger LLM coders.
QAOD: Single-pass white-box framework for hallucination detection in LLMs using question-answer orthogonal decomposition.
LiSA: lifelong safety adaptation framework for AI agents executing workflows, maintaining guardrails against contextual safety failures in tool use and data access.
Method for positive-unlabeled learning from highly imbalanced datasets, applicable to disease identification, fraud detection, and recommender systems.
EvoLib: test-time learning framework enabling LLMs to accumulate and evolve knowledge across problem instances via extracted modular skills without parameter updates.
Offline-to-online reinforcement learning method using bi-level optimization for adaptive data mixing between datasets.
Autonomous agentic workflow for end-to-end machine learning interatomic potential development with adaptive active learning.
Mechanistic interpretability study using activation patching to examine how LLMs process relative geographic space.
Benchmark for evaluating LLM medication recommendations at prescription level with per-timepoint predictions and information-rich inputs.
Weight space analysis of neural PDE operators identifying reusable physical structure across multiple regimes in pretrained models.
Multi-objective prompt optimization technique using pure-exploration bandits for efficient LLM prompt selection.
Token-level credit assignment method for agentic RL using energy-based perspectives to improve PPO and GRPO training signals.