Probing Dec-POMDP Reasoning in Cooperative MARL
Analyzes reasoning capabilities in cooperative multi-agent reinforcement learning under Dec-POMDP framework with partial observability and decentralized coordination.
Analyzes reasoning capabilities in cooperative multi-agent reinforcement learning under Dec-POMDP framework with partial observability and decentralized coordination.
Proposes regret-guided search control for AlphaZero to improve learning efficiency through targeted state revisitation in reinforcement learning.
Introduces transcoder adapters to interpret MLP computation differences in reasoning models, comparing Qwen2.5-Math-7B and DeepSeek-R1-Distilled variants.
Semantic-guided expert forest method for class-incremental learning that organizes adapter knowledge through task relationships.
Extension of maximal update parameterization for hyperparameter transfer across different optimizers in large language model training.
Theoretical work connecting law of robustness to robust generalization in neural networks through Lipschitz constraints.
Theoretical analysis of multi-agent imitation learning showing impossibility and hardness results for learning low-exploitable policies offline.
Domain adaptation method for offline reinforcement learning addressing dynamics mismatch through localized similarity matching.
Federated semi-supervised learning framework using proxy guidance to handle data heterogeneity across and within distributed clients.
Evaluation of DeepSpeed distributed training framework for scaling Vision Transformer models to address computational and memory demands.
Analysis of Graph Neural Network activation patterns through graph topology and curvature to understand oversmoothing and oversquashing artifacts.
Tokenization method combining vector quantization with Self-Organizing Maps to create structured discrete codebooks for interactive generative models.
Self-evolving LLM agent framework using uncertainty-aware rewards to guide multi-step decision-making and improve learning signal for agent training.
Study on Predictor-Corrector samplers for discrete diffusion models to improve multi-step generation quality beyond ancestral sampling methods.
Research on Pass@k metric optimization for LLMs showing trade-offs between multi-sample and single-sample inference performance in code generation and reasoning tasks.
Statistical query lower bounds for learning halfspaces under Gaussian perturbations in the smoothed agnostic learning model.
Untied Ulysses: memory-efficient context parallelism technique for long sequences via headwise chunking in Transformers.
Reflective Test-Time Planning for embodied LLMs combining in-action and post-action reflection with test-time scaling for robot task learning.
Analysis revealing that test-time training with KV binding can be expressed as learned linear attention mechanism.
Benchmark comparing knowledge-distilled small language models against vanilla and proprietary models for resource-constrained environments.
Leakage-aware benchmarking framework for early patient deterioration prediction under realistic emergency triage sensing constraints.
Deep learning model for de novo peptide sequencing with explicit mass consistency constraints using regressor-guided diffusion.
Theoretical analysis of gap-dependent regret bounds for reinforcement learning algorithms with linear function approximation.
LLM-based framework for rare disease phenotyping from clinical notes, extracting and standardizing features to Human Phenotype Ontology terms.
Circuit tracing framework for vision-language models using transcoders and attribution graphs to analyze multimodal reasoning mechanisms.
QueryBandits: model-agnostic contextual bandit framework to mitigate hallucinations in closed-source LLMs through post-hoc detection and mitigation.
Research: Zero-overhead online continual learning framework for deep neural network OFDM receivers adapting to time-varying communication channels.
Research on detecting and mitigating group bias in heterogeneous treatment effect predictions when aggregating ML model outputs to subgroups.
Diffusion-based planning method for offline reinforcement learning with environment mechanism modeling to ensure trajectory consistency.
Heterogeneity-aware client selection for federated learning to improve model accuracy across statistical heterogeneous clients.
Prior-agnostic incentive-compatible exploration in bandit settings with multiple sequential agents.
ActionEngine: training-free framework for GUI agents using state machine memory to reduce costs, latency, and improve accuracy vs reactive vision-language model approaches.
Inner speech guides steerable imitation learning for human-AI coordination, capturing behavioral diversity and non-Markovian human behaviors.
STAR-LDM integrates latent diffusion planning with autoregressive generation, enabling semantic planning before token commitment.
Theoretical proof that standard Transformers achieve minimax approximation rate for Hölder functions in nonparametric regression.
Personal information memorization in language models: detector suite for email, phone, IP addresses outperforms existing regex baselines.
Learning theory characterizing online and private learnability under distributional constraints via generalized smoothness.
Interpretable open-world object detection framework using concept decomposition to distinguish known from unknown objects.
Theoretical analysis of SGD convergence with perturbations in forward-backward passes through sequential operators.
DANCE method for conformal prediction uncertainty quantification using adaptive neighborhood estimation with pre-trained deep learning models.
Vision-language models applied to ergonomic assessment by estimating hand distances from RGB video for NIOSH lifting equation.
Communication-inspired discrete image tokenizer for vision transformers optimized for semantic structure over texture reconstruction.
SibylSense enables adaptive reward rubric learning for open-ended generation via memory tuning and adversarial probing to prevent reward hacking.
Tail-aware divergence for language model distillation that decouples top-K probabilities to improve knowledge transfer from teacher to student models.
DRESS framework addresses computational complexity of higher-order Weisfeiler-Lehman graph analysis using continuous dynamical systems.
Functional Continuous Decomposition framework for non-stationary time-series analysis with parametric optimization and guaranteed continuity.
SpatiaLQA benchmark evaluates spatial logical reasoning capabilities in vision-language models for real-world scenarios.
Economic analysis of AGI impact on labor, marginal costs, and human verification as binding constraint on growth.
PyTorch framework for medical image processing with support for volumetric data and domain-specific training procedures.
Studies multi-distribution learning complexity with bounded label noise to determine if single-task 1/ε rates extend to multi-source settings.