Actor-Critic Pretraining for Proximal Policy Optimization
Actor-critic pretraining approach for PPO that leverages expert data to reduce environment interactions required for RL training.
Actor-critic pretraining approach for PPO that leverages expert data to reduce environment interactions required for RL training.
Theoretical investigation of offline reinforcement learning with general function approximation and parametric policies beyond state-wise methods.
Q-learning approach for learning safe policies from expert demonstrations with unknown constraints in constrained MDPs.
FedNSAM addresses sharpness-aware minimization in federated learning under high data heterogeneity, ensuring both local and global model flatness.
LK Losses directly optimize acceptance rate in speculative decoding for LLM inference, improving upon KL divergence proxy objectives for draft model training.
Hierarchical Concept Embedding Models improve interpretability of deep neural networks by mapping inputs to human-interpretable concept representations with inter-concept relationships.
Learns optimal generation orders for masked discrete diffusion models via variational inference to balance parallel generation and sample quality.
Outlines vision for foundation world models as persistent compositional representations enabling agents to learn and adapt in open worlds.
Proposes RewardUQ framework for uncertainty-aware reward models in LLM alignment that reduces annotation costs and prevents overoptimization.
Introduces pathsig, a PyTorch-native GPU-accelerated library for computing path signatures as trainable features for sequential data.
Proposes ACWI framework that adaptively balances intrinsic and extrinsic rewards online for sparse reward reinforcement learning exploration.
Surveys agentic AI systems with planning, tool use, and self-management capabilities applied to Open RAN network control and optimization.
Studies best arm identification problem with heterogeneous resource costs and constraints across multiple resource types.
Proposes explainable AI method for discrete token inputs like text using attribution highlighting to identify important tokens in transformers.
Applies multi-objective reinforcement learning to optimize container consolidation in human-robot collaborative fulfillment centers.
Proposes federated learning approach for anomaly detection in heterogeneous IoT networks while preserving privacy through distributed training.
Investigates trade-off between regret minimization and statistical power in combinatorial multi-armed bandits using Pareto optimality framework.
arXiv paper benchmarking general-purpose time-series foundation models for zero-shot transportation forecasting across multiple datasets.
arXiv paper on Latent Manifold Compaction for unsupervised harmonization of histopathology images across different batch effects and scanners.
arXiv paper proposing Web-Knowledge-Web pipeline for discovering suppliers in specialized industries via iterative web crawling and knowledge base integration.
Analyzes limitations of standard identifiability metrics (MCC, DCI, R²) on synthetic benchmarks, revealing implicit structural assumptions in representation learning evaluation.
Memory caching architecture enabling RNNs with growing memory capacity and subquadratic complexity as alternative to Transformers for sequence modeling.
Low-rank approximation method (LoRA-Pre) for optimizer memory efficiency in Adam and Muon, reducing overhead for large language model training.
Agentic RL system using LLMs for high-performance CUDA kernel generation at scale, outcompeting traditional compiler-based approaches.
Graph reinforcement learning approach to moderate opinion polarization in social networks under Friedkin-Johnsen model with improved scalability.
Methodology for dynamic neural networks using isotropic activation functions enabling real-time architectural growth and shrinkage via symmetry-principled primitives.
Improves Laplace mechanism for differentially private SGD in high-dimensional models using majorization theory, applicable to LLM fine-tuning.
Few-shot continual learning approach for 3D brain MRI using frozen foundation models with task-specific LoRA modules for tumor segmentation and brain age estimation.
Completeness and bounding results for causal identification using counterfactual data from Layer 3 of Pearl's Causal Hierarchy.
Probabilistic framework for symbolic regression using variational inference over soft symbolic trees for scientific discovery with uncertainty quantification.
Hybrid reinforcement learning approach (RL-CMSA) for solving min-max multiple traveling salesman problems with iterative construction and adaptation.
Efficient image captioning via hyperdimensional cross-modal alignment of frozen language and vision models without multimodal fine-tuning.
Causal discovery method distinguishing whether variables influence the mean versus variance of other variables in heteroscedastic data.
Retrieval system learning node-specific Riemannian metrics for geometry-aware semantic search on citation graphs.
General Bayes framework for policy learning where decision rules are the target rather than outcome prediction.
Platform for standardized access to remote sensing foundation model embeddings across heterogeneous model formats and interfaces.
Dynamic benchmarking protocol where AI agents autonomously generate, validate, and solve problems to evaluate LLM reasoning capabilities beyond static datasets.
Open-source interpretability tool for analyzing gated activation functions (SwiGLU) in transformer neurons across recent language models.
Addresses catastrophic forgetting in continual fine-tuning of LLMs for vulnerability detection in source code, using selective replay with LoRA on temporal distribution shifts.
MI²DAS: multi-layer intrusion detection framework for IIoT with incremental learning to detect novel attacks in dynamic environments.
Benchmark study evaluating cross-domain transferability of flow-based feature sets across IoT and IIoT datasets for intrusion detection.
RF-Agent: automated reward function design for reinforcement learning using LLM-based tree search to optimize low-level control tasks.
BUSD-Agent: cascaded multi-agent framework for breast ultrasound screening and diagnosis reducing biopsy referrals through selective decision-making.
Autonomous robotic assembly framework using reinforcement learning for constructing stable structures without predefined plans.
Benchmarking 10 BERT variants for Nepali sentence-level topic classification, evaluating multilingual and Indic-specific models.
Jailbreak Foundry: multi-agent system translating jailbreak papers into executable modules for unified benchmarking and reproducible LLM robustness evaluation.
Data-driven optimization pipeline for GPU efficiency in distributed LLM adapter serving, maximizing throughput with concurrent adapter hosting.
Two-stage unsupervised pipeline for IoT device traffic profiling with incremental model adaptation using density-based clustering.
Analysis of monoculture in LLMs showing agreement metrics depend on subjective baseline assumptions for independence.
Artificial Agency Program: research agenda for building resource-bounded AI agents driven by curiosity-as-learning-progress and human-tool integration.