Slow-Fast Policy Optimization: Reposition-Before-Update for LLM Reasoning
Reinforcement learning framework (SFPO) for improving LLM reasoning by decomposing policy optimization into slow and fast components to reduce training instability.
Reinforcement learning framework (SFPO) for improving LLM reasoning by decomposing policy optimization into slow and fast components to reduce training instability.
Information-theoretic framework measuring higher-order structure and emergence in multi-agent LLM systems. Tests for dynamical emergence in agent coordination.
Model-native technique to explain LLM hallucinations using layer-wise semantic maps. Traces concept flow through residual streams via unembedding.
Theoretical analysis of knowledge distillation in neural networks from a functional perspective. Decouples compression from architecture reduction.
Justitia scheduling algorithm for fair and efficient execution of task-parallel LLM agents on shared GPUs. Resource scheduling optimization for agent systems.
StreamingTOM training-free token compression for streaming video understanding. Efficiency optimization for video vision-language models.
GlobalRAG reinforcement learning approach for multi-hop question answering with improved global reasoning. RAG system enhancement via RL for better query planning.
Gaze-guided object detection framework for egocentric videos using vision transformers. Computer vision with attention mechanisms.
Zero-shot robotic grasp detection using vision-language models without training data or retraining. Robotics application leveraging VLMs.
Semi-supervised intent detection framework with active learning correction for voice dialog agents. LLM application for dialog system improvement.
Vision-language model for unified gaze understanding combining detection, target, and object recognition. Multimodal model for attention estimation.
Study of multi-agent LLM cooperation on math problems with analysis of adversarial robustness. Evaluates agent collaboration and vulnerability to perturbations.
Privacy-preserving explainable AI for IoT applications using SHAP entropy regularization. Focuses on interpretability and privacy in edge devices.
Training-free stabilizer (PAS) fixing temporal inconsistency in video LLMs caused by rotary position embedding ripples. Technical improvement for video understanding.
Decoupled action expert for vision-language-action models using diffusion/flow-matching for manipulation policies. Computer vision and robotics research.
Fast inference acceleration for masked auto-regressive diffusion models enabling practical reinforcement learning. Optimization technique for generative models in RL contexts.
Pipeline for speaker-attributed civic simulation using LLMs from ASR transcripts. LLM application for multi-agent deliberation modeling.
Concept bottleneck models for explainable visual anomaly detection with semantic interpretability. Computer vision research with interpretability focus.
MapReduce LoRA and RaTE methods for multi-preference optimization in generative models using RLHF. Advances alignment of LLMs through parameter-efficient fine-tuning techniques.
Review of uncertainty quantification and data efficiency methods for AI systems in robotics, telecommunications, and healthcare. Theoretical ML research with practical applications.
Combines large language models with transformer encoders for financial news classification with limited labeled training data.
Rough set theory applied to explain results of spectral graph clustering algorithms for text document analysis.
Method for watermarking deep neural networks to protect intellectual property and verify model ownership using chaos-based approaches.
Geo-4D: training-free geometric approach for 4D LiDAR panoptic segmentation without deep networks or dedicated modules.
CARE: failure-centric post-training framework using contrastive learning for verifiable multimodal reasoning in reinforcement learning.
EgoGrasp: method for reconstructing world-space hand-object interactions from egocentric videos supporting open-vocabulary objects.
Agentic Retoucher: hierarchical agent for fixing distortions in text-to-image generation with spatial grounding.
WebCoderBench: benchmark for evaluating LLM-generated web applications with comprehensive metrics and interpretable results.
LAMB: LLM-based audio captioning framework bridging modality gap between audio and text using Cauchy-Schwarz divergence.
GeoMotionGPT: LLM framework for motion understanding using geometry-aligned discrete motion tokenization and embeddings.
Imagine-then-Plan framework for agent learning using world models and adaptive lookahead for complex task planning.
RAG-3DSG: retrieval-augmented generation approach for constructing 3D scene graphs with uncertainty estimation for robotics tasks.
Multi-agent LLM framework for generating research limitations by identifying methodological issues beyond superficial statements.
Research on decomposable inference in large models showing gradient updates are localized, reducing inference costs and complexity.
Jacobian Scopes: gradient-based methods for token-level causal attribution in LLMs to identify which prior tokens influence predictions across layers and attention heads.
VibeVoice-ASR framework for speech understanding in long-form audio using single-pass processing to handle context fragmentation and multi-speaker scenarios.
NaVIDA improves vision-language navigation agents by explicitly modeling action-grounded visual dynamics for better planning and generalization in embodied environments.
Benchmark for evaluating reasoning in baby language models trained on child-directed speech; developmentally-inspired evaluation methodology.
Expert-panel study on human detection of LLM-generated Korean text using rubric-based calibration framework for attribution.
Mixture-of-Experts approach for time-series forecasting transformers using segment-wise routing to improve scaling and temporal dynamics.
MDial framework for generating multi-dialectal dialogue data; addresses LLM performance gaps for non-standard English speakers.
Adversarial attacks against search-enabled LLM fact-checking systems; proposes DECEIVE-AFC for testing robustness of retrieval-augmented verification.
First labeled dataset of 98,380 malicious agent skills characterizing security threats in LLM-based agent registries and ecosystems.
Open-source singing voice synthesis system with zero-shot generalization and controllable generation capabilities.
System for natural language graph analytics over large property graphs using LLMs; enables querying complex heterogeneous datasets efficiently.
Token reduction method for multi-modal LLMs in autonomous driving to improve efficiency while maintaining human-vehicle interaction.
Physics-based tropical cyclone estimation using spline-parameterized KAN for efficient edge device deployment on satellite data.
Framework for on-policy supervised fine-tuning of LLMs using Distribution Discriminant Theory to improve generalization over standard SFT.
Generic object tracking method using joint-embedding predictive architecture with model adaptation and occlusion reasoning.
Research on identifying missing persona dimensions for user simulation in dialogue systems to improve validity of simulation results.