AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation
AgentLens benchmark evaluates coding agents on full trajectory quality including instruction-following, tool-use, and error recovery, not just pass/fail.
AgentLens benchmark evaluates coding agents on full trajectory quality including instruction-following, tool-use, and error recovery, not just pass/fail.
Dynamic-in-Few-Step combines dynamic computation with distillation to accelerate video diffusion model inference.
Analysis of how specification-grounded tests improve LLM-generated code quality on edge cases compared to model self-testing.
Retrieval-Augmented Generation applied to medical QA reduces LLM hallucinations and grounds answers in current public health data.
Method to detect and mitigate backdoor attacks in decentralized model training without full recomputation overhead.
POPS method recovers knowledge in multimodal LLMs after machine unlearning, addressing privacy-utility tradeoffs in MLLMs trained on sensitive data.
Vision-language-action model integrating perception, future prediction, and action planning with attention-based generalization without task-specific tuning.
Review of Vision Language Action models for embodied AI in robotics and manipulation tasks from camera images.
Theoretical framework unifying gradient-based optimization as coupled evolution of parameters, particles, and time-varying Riemannian metrics.
Diff-aware ML model for predicting deployment risk at Prime Video to reduce unnecessary code freezes.
Reinforcement learning application using policy gradient methods on masked language models for e-commerce ad headline generation.
Framework extending symplectic learning to robotic systems with actuation and dissipation for stable dynamics prediction.
Gradient-based method for speech-to-text alignment compatible with CTC, transducer, and speech LLM models.
Low-rank tensor train approach for high-dimensional sampling in diffusion models using score-based methods.
Analysis of video diffusion model representations revealing structured latent spaces across network depth and noise levels.
Theoretical framework modeling AI-augmented computation as probabilistic Turing machines interacting with stochastic oracles.
Parameter-efficient fine-tuning method extending LoRA to vision foundation models with spatial awareness for downstream tasks.
Survey of mathematical foundations in reinforcement learning, covering MDPs, Bellman operators, and convergence guarantees for value/policy iteration algorithms.
Flow-ERD combines flow matching with entropy-regularized distillation for diverse and realistic multi-agent traffic simulation in autonomous driving.
MILES proposes modular instruction memory with learnable selection enabling LLMs to accumulate and reuse reasoning experience across sequential problems.
EdgeCompress framework combining dynamic image cropping and model compression techniques to reduce CNN computational overhead for resource-constrained embedded devices.
Tensorized algorithms and scalable filtering methods for hidden Markov models and factorial HMMs to efficiently represent multi-factor time-series systems.
Analysis of behavior shift in multi-teacher on-policy distillation for LLM agents learning tool use, showing invisible distribution shifts in aggregate losses.
Active deep probabilistic subsampling approach with prior-aware context-guided group sampling for reducing measurement overhead and data transfer.
Comparative evaluation of 8 open-source VLMs on document visual question answering, assessing robustness and transferability across document domains.
Privacy-preserving continual learning framework with auditable buffering-aggregation recipe for federated and streaming systems under adaptive interaction.
DiPhon applies diffusion models to large-scale graph generation using graphon theory to study structural statistics across different node scales.
Multi-fidelity Bayesian optimization framework for tuning genetic algorithm hyperparameters using FFT, CNN surrogates, and Gaussian processes for lattice material design.
PA-SciML introduces physics-audited verification workflow for LLM agents discovering scientific surrogate models, ensuring predictions satisfy physical constraints beyond error metrics.
TF-Engram proposes SSD-backed memory architecture for LLMs to enable efficient knowledge expansion without retraining, using engram-style hidden-state injection.
ARGTCA improves confidence calibration in vision-language models by using graph-based attribute reasoning to represent class relationships, addressing overconfidence from prompt tuning.
GIFT: geometry-informed low-precision gradient quantization for distributed LLM pretraining reducing communication bottlenecks.
PALS: layer-aware sparsity pruning for LLMs using percentile-based activation thresholds, improving perplexity over uniform pruning methods.
Studies activation patterns in Polish Bielik LLMs to identify entity familiarity and factual reliability signals before answer generation across model scales.
MedPMC framework scales high-fidelity medical multimodal data from PubMed Central for foundation model development with clinical validation.
Studies generalization of variable-size input models (point clouds, sequences, graphs) from small training sizes to unseen larger sizes.
Co-LMLM externalizes factual knowledge to continuous-query knowledge bases during LLM generation, enabling knowledge control beyond conventional weight-based approaches.
Analyzes adversarial Rademacher complexity of DNNs to provide generalization guarantees for models robust to adversarial perturbations.
Studies off-policy adversarial imitation learning convergence and sample complexity, showing sample reuse improves efficiency without importance sampling correction.
Analyzes robust overfitting in adversarial training using dynamical systems and PAC-Bayesian theory, providing mechanistic interpretation of generalization in robust models.
ContrastiveCFG improves diffusion model sampling by contrasting positive and negative concepts, addressing limitations of naive negative prompting in classifier-free guidance.
Proposes theoretical framework analyzing deep neural network learning dynamics through dynamical systems theory, introducing order-preserving transformations at neuron level.
PB-OEL framework for real-time safety assessment using online ensemble learning with mixed feedback and performance guarantees under concept drift.
Studies silent neurons and plasticity in deep reinforcement learning for adaptive video streaming to improve generalization across heterogeneous network conditions.
Protocol Models proposes communication-efficient model parallelism for decentralized training by compressing activations and activation gradients in distributed deep learning setups.
FPTQuant introduces function-preserving transforms to enable efficient quantization of LLMs by handling magnitude outliers, improving inference efficiency without significant performance degradation.
Theoretical analysis of two-layer neural networks with smooth activation functions using Taylor series and spline methods.
L-GTA: VAE-based generative model for time series augmentation using Bi-LSTM and temporal self-attention.
Quantization technique for vision encoders using register prefixing to handle outliers and reduce inference cost.
Theoretical analysis proving affine identifiability in nonlinear CCA under latent distributional priors.