Variational Deep Learning via Implicit Regularization
Analysis of implicit regularization in overparametrized deep neural networks and improved out-of-distribution generalization via variational methods.
Analysis of implicit regularization in overparametrized deep neural networks and improved out-of-distribution generalization via variational methods.
Dynamic benchmark framework (NetArena) for evaluating AI agents in network automation with production-level complexity and reduced contamination risk.
Adaptive multi-objective reinforcement learning method for balancing exploration and skill diversity in skill-based RL pretraining.
Benchmark for evaluating multimodal LLM-based front-end code generation with modern development frameworks and evaluation metrics.
Curriculum learning approach scheduling tasks from easy to hard to improve LLM reasoning via reinforcement learning, inspired by DeepSeek-R1.
BIS Reasoning 1.0: Japanese benchmark with 1K+ syllogistic problems evaluating belief bias and inconsistent reasoning in LLMs.
AVA-Bench: systematic evaluation benchmark for vision foundation models addressing blind spots in VQA evaluation protocols.
TRACED: unsupervised environment design using regret approximation for co-learning to improve deep RL agent generalization.
Rationale-Enhanced Decoding improves chain-of-thought prompting in vision-language models by optimizing intermediate reasoning generation.
Lumos-1: LLM-based autoregressive video generation using discrete diffusion with efficient architecture avoiding external encoders.
SOAR: self-improving method integrating language models into evolutionary program synthesis for challenging tasks like ARC-AGI.
FingerTip 20K: benchmark for proactive mobile LLM agents with 20K tasks, evaluating multimodal agents using contextual data without explicit instructions.
Neural Combinatorial Optimization solver for min-max heterogeneous vehicle routing with multiple vehicles using novel decoding approach.
EvolvR: self-evolving method for story evaluation using LLM-as-judge with pairwise reasoning to improve generation guidance.
Novel benchmarking system evaluating LLM-based agent capabilities for single-cell omics data analysis, assessing planning and code generation.
Systematic study of post-training quantization methods for diffusion LLMs to enable edge device deployment, comparing compression techniques.
UTRL: reinforcement learning framework training LLMs to generate high-quality unit tests automatically, addressing test generation challenges.
Research evaluating Law-Following AI framework for embedding legal compliance in advanced AI agents, analyzing legal personhood constructs and technical feasibility.
Reinforcement learning approach for radiology report generation using FactScore-based rewards with reduced data requirements.
Framework evaluating robustness of Vision-Language-Action models under real-world physical variations for robotic tasks.
Matched-compute study evaluating synthetic data interventions for in-context learning in language models. Tests mechanism-targeted pretraining effects.
Method for reducing LLM agent inference costs through trajectory reduction. Addresses token cost efficiency in multi-turn agent systems for software engineering.
Technique reducing LLM reasoning model overthinking through decoupled rewards and curriculum scheduling. Addresses excessive token generation without performance gain.
Reinforcement learning framework (SFPO) for improving LLM reasoning by decomposing policy optimization into slow and fast components to reduce training instability.
Information-theoretic framework measuring higher-order structure and emergence in multi-agent LLM systems. Tests for dynamical emergence in agent coordination.
Model-native technique to explain LLM hallucinations using layer-wise semantic maps. Traces concept flow through residual streams via unembedding.
Theoretical analysis of knowledge distillation in neural networks from a functional perspective. Decouples compression from architecture reduction.
Justitia scheduling algorithm for fair and efficient execution of task-parallel LLM agents on shared GPUs. Resource scheduling optimization for agent systems.
StreamingTOM training-free token compression for streaming video understanding. Efficiency optimization for video vision-language models.
GlobalRAG reinforcement learning approach for multi-hop question answering with improved global reasoning. RAG system enhancement via RL for better query planning.
Gaze-guided object detection framework for egocentric videos using vision transformers. Computer vision with attention mechanisms.
Zero-shot robotic grasp detection using vision-language models without training data or retraining. Robotics application leveraging VLMs.
Semi-supervised intent detection framework with active learning correction for voice dialog agents. LLM application for dialog system improvement.
Vision-language model for unified gaze understanding combining detection, target, and object recognition. Multimodal model for attention estimation.
Study of multi-agent LLM cooperation on math problems with analysis of adversarial robustness. Evaluates agent collaboration and vulnerability to perturbations.
Privacy-preserving explainable AI for IoT applications using SHAP entropy regularization. Focuses on interpretability and privacy in edge devices.
Training-free stabilizer (PAS) fixing temporal inconsistency in video LLMs caused by rotary position embedding ripples. Technical improvement for video understanding.
Decoupled action expert for vision-language-action models using diffusion/flow-matching for manipulation policies. Computer vision and robotics research.
Fast inference acceleration for masked auto-regressive diffusion models enabling practical reinforcement learning. Optimization technique for generative models in RL contexts.
Pipeline for speaker-attributed civic simulation using LLMs from ASR transcripts. LLM application for multi-agent deliberation modeling.
Concept bottleneck models for explainable visual anomaly detection with semantic interpretability. Computer vision research with interpretability focus.
MapReduce LoRA and RaTE methods for multi-preference optimization in generative models using RLHF. Advances alignment of LLMs through parameter-efficient fine-tuning techniques.
Review of uncertainty quantification and data efficiency methods for AI systems in robotics, telecommunications, and healthcare. Theoretical ML research with practical applications.
Combines large language models with transformer encoders for financial news classification with limited labeled training data.
Rough set theory applied to explain results of spectral graph clustering algorithms for text document analysis.
Method for watermarking deep neural networks to protect intellectual property and verify model ownership using chaos-based approaches.
Geo-4D: training-free geometric approach for 4D LiDAR panoptic segmentation without deep networks or dedicated modules.
CARE: failure-centric post-training framework using contrastive learning for verifiable multimodal reasoning in reinforcement learning.
EgoGrasp: method for reconstructing world-space hand-object interactions from egocentric videos supporting open-vocabulary objects.
Agentic Retoucher: hierarchical agent for fixing distortions in text-to-image generation with spatial grounding.