VeRO: An Evaluation Harness for Agents to Optimize Agents
VeRO: Evaluation framework for assessing coding agents that optimize other agents through iterative edit-execute-evaluate cycles.
VeRO: Evaluation framework for assessing coding agents that optimize other agents through iterative edit-execute-evaluate cycles.
Flow matching generative model that adapts to manifold structures, offering simulation-free alternative to diffusion models.
Theoretical analysis of scaling limits from shallow Bayesian neural networks to Gaussian processes with scalable inference methods.
Dynamic dense retrieval with routing strategy for adapting information retrieval models across domains without full retraining.
CourtGuard: Model-agnostic multi-agent framework for zero-shot LLM safety policy adaptation using retrieval-augmented debate.
Search-P1: Path-centric reward shaping for training agentic RAG systems with improved sample efficiency via RL.
Item Response Theory approach to correct systematic rater biases in human evaluations for AI model assessment.
SideQuest: Model-driven KV cache management technique for long-context agentic reasoning tasks with multi-hop retrieval.
Hybrid ML framework combining autoencoders and transformers for accelerator beam diagnostics simulations.
Trie-based constrained decoding optimization for LLM generative retrieval on accelerators. Improves business logic constraints in recommendations.
dLLM: Unified framework for diffusion language models. Standardizes components across research implementations for reproducibility.
GR4AD: Production generative recommendation system for large-scale advertising using LLMs. Architecture, learning, and serving optimization.
AMA-Bench: Benchmark evaluating long-horizon memory for LLM-based agentic applications. Addresses gap between dialogue benchmarks and real agent scenarios.
QSIM: Multi-agent reinforcement learning method addressing Q-value overestimation via action similarity weighting. Value decomposition improvement.
Research on diffusion models for end-to-end autonomous driving in real-world settings. arXiv paper exploring decision-making applications.
TARAZ benchmark for evaluating cultural competence of LLMs in Persian with short-answer format and morphological analysis.
Unsupervised continual learning framework for amortized Bayesian inference handling sequential data and distribution shifts.
SPD Learn Python library for symmetric positive definite matrix neural networks in neural decoding with geometric deep learning.
OmniGAIA benchmark for evaluating omni-modal AI agents with unified vision, audio, and language perception capabilities.
SIGMA system applying LLMs to multi-task recommendation at scale, handling diverse business requirements beyond traditional next-item prediction.
Parameter-efficient fine-tuning combining multiple domain expert models for visual adaptation tasks via prompt tuning.
Tree-based system for analyzing semi-structured documents with mixed content types enabling question-answering over complex layouts.
LLM agent framework (SALA) for evaluating deanonymization risks through stylometry-assisted analysis of textual data with interpretable pipeline.
Study of information-theoretic limits in multimodal LLMs showing modality-specific information is discarded during decoding despite surviving encoding layers.
arXiv paper: Fairness-aware mixed-precision quantization for neural network compression in medical imaging, explicitly addressing algorithmic fairness during model compression.
arXiv paper: Theoretical analysis of fine-tuning effects on in-context learning in linear attention transformers, balancing downstream and few-shot task performance.
arXiv paper: Plug-and-play diffusion framework with ADMM for medical image reconstruction addressing memory issues in PnP solvers.
arXiv paper: Large-scale app store search ranking system augmented with LLM-generated textual relevance judgments to address scarcity of expert labels.
arXiv paper: Zeroth-order optimization for leader-follower Stackelberg control in combinatorial congestion games with discrete strategy selection.
arXiv paper: Training-free hierarchical manifold guidance for dataset distillation using diffusion models without requiring gradient computation.
arXiv paper: Zero-shot and one-shot adaptation of small language models for leader-follower role assignment in human-robot interaction with resource constraints.
Analyzes parameter-efficient fine-tuning for continual learning using Neural Tangent Kernel theory to understand model adaptation and catastrophic forgetting mitigation.
Physics-inspired neural framework using graph neural networks to solve large-scale graph coloring combinatorial optimization problems near algorithmic phase transitions.
Proposes label unlearning method for vertical federated learning using representation-level manifold mixup to enable privacy-preserving model unlearning.
Model-agnostic explanation technique integrating concept-based approaches with diverse explanation forms beyond attribution methods.
Reinforcement learning framework for aligning few-step diffusion models with downstream objectives using stepwise policy optimization.
Neuro-symbolic framework automating discovery of analytical solutions to differential equations using formal grammars and continuous search.
Continual learning method using sample compression theory to provide computable guarantees while avoiding catastrophic forgetting.
Compression technique for large language models using global rank and sparsity optimization with layer-wise weight allocation.
Large-scale benchmark for evaluating machine learning solvers on combinatorial optimization using real-world industrial datasets.
Research on whether language models can learn to evade latent-space safety monitors, with implications for LLM alignment and security.
Statistical framework for evaluating LLM-as-a-judge systems, addressing reliability and bias issues in automated LLM output evaluation.
Silent Gradients approach for training VAEs by restricting decoder architecture to reduce gradient estimation variance.
Online time series forecasting method using feature adjustment to handle distribution shift in sequential data deployment.
rBridge framework enabling small proxy models (≤1B params) to predict reasoning performance of larger LLMs, optimizing dataset scaling.
Theoretical analysis of softmax attention mechanisms in LLMs through single-location regression task, addressing why softmax dominates alternative activation functions.
Study of compute-optimal allocation between full-precision and quantization-aware training phases for neural networks.
Dynamic safety monitoring system for LLMs that adaptively adjusts computation based on input difficulty.
Method for improving masked diffusion models by learning better unmasking policies beyond rule-based position scheduling.
Imitation learning approach for training models with multiple acceptable answers from limited correct demonstrations.