Toward a Multi-Layer ML-Based Security Framework for Industrial IoT
ML-based multi-layer security framework for Industrial IoT addressing resource constraints and threats across network layers.
ML-based multi-layer security framework for Industrial IoT addressing resource constraints and threats across network layers.
MedAidDialog multilingual multi-turn medical dialogue dataset for conversational AI in healthcare with improved realism over template-based systems.
TSRL framework uses reinforcement learning to dynamically optimize training curriculum for deepfake detection, modeling training as an MDP.
Visual study of UMAP projections examining geometric patterns in embedding difference vectors of antonym and synonym word pairs.
Applies quantum convolutional neural networks to solve partial differential equations on quantum simulators for scientific computing applications.
HEART-PFL framework for personalized federated learning using hierarchical directional alignment and adversarial knowledge transfer to handle data heterogeneity.
UniScale explores synergistic data and model scaling for search ranking, demonstrating that joint architectural and data design improvements outperform model scaling alone.
DVM enables real-time kernel generation for dynamic AI models, addressing compilation overhead and memory footprint issues in runtime compilation.
C-STEP introduces physics-informed safety measures for reinforcement learning in robotics, using intrinsic rewards for safe navigation in continuous domains.
CGRL framework addresses poor generalization of GNNs on out-of-distribution data using causal-guided representation learning to avoid spurious correlations.
Proposes method to quantify self-awareness in intelligent systems by identifying invariant cognitive processes that change slower than acquired skills.
Neuro-symbolic system using attention-based encoders and differentiable reasoning rules to detect human fatigue from eye-tracking and fNIRS signals.
Investigates joint effects of differential privacy and fairness constraints on federated classification systems across distributed servers.
Studies relationship between fair model representations and fair recommendations in recommender systems, examining demographic attribute classification.
Analysis of why self-distillation degrades LLM reasoning capability by suppressing epistemic verbalization and expression of uncertainty.
Composer 2 model specialized for agentic software engineering with long-term planning and coding abilities trained via continued pretraining and reinforcement learning.
Multi-agent framework with verification for improving calibration and accuracy in medical multiple-choice question answering.
Study evaluating RAG systems on AI policy analysis showing retrieval improvements don't guarantee better answers on complex regulatory documents.
Inverse-forward differentiation method to reduce memory requirements for backpropagation by avoiding activation storage.
Learning-theoretic framework for coded computing in distributed systems to handle slow, faulty, or compromised servers.
Visualization technique for understanding RNN internal dynamics during training using multislice PHATE algorithm.
Physics-informed neural networks using wavelet decomposition to improve training on differential equations with rapid oscillations and steep gradients.
arXiv paper on Symmetry-Guided Memory Augmentation (SGMA) improving efficiency of RL-based legged locomotion training.
arXiv paper on machine learning techniques to detect and localize power/radiation leakage of cryptographic keys from hardware implementations.
arXiv paper on multi-agent reinforcement learning for adaptive traffic signal control in heterogeneous urban networks.
arXiv paper: GraphOmni benchmark framework evaluating LLM reasoning on graph-theoretic tasks with diverse formats and serializations.
arXiv paper introducing Distance Explainer method for post-hoc interpretability of embedded vector spaces in ML models.
arXiv paper on Bottlenecked Transformers: KV cache consolidation technique for scaling inference-time reasoning in LLMs.
arXiv paper interpreting neural networks as dynamical systems on latent manifolds, analyzing autoencoder vector fields.
arXiv paper on scalable longitudinal patient pathway modeling from multimodal EHR data using neural networks for condition forecasting.
Research paper demonstrating LLMs perform in-context reinforcement learning during inference. ICRL prompting framework enables inference-time self-improvement.
TimeRecipe benchmarks module-level effectiveness of components in time-series forecasting architectures.
DART adds server-side robustness to federated learning for edge devices without expensive client-side computation.
Theoretical analysis of federated distillation with weighted aggregation of client predictions under class mismatch.
PromptLoop refines prompts for diffusion models using sequential reinforcement learning feedback during sampling.
Generative method for synthetic financial time series data to address data shortage in ML models for trading and investment.
Proposes future summary pretraining for LLMs as alternative to next-token prediction, addressing limitations in long-horizon reasoning and planning tasks.
Addresses distribution shift in time-series forecasting by identifying concept drift and temporal shift, proposing mitigation strategies for generalization.
OffSim proposes model-based offline inverse RL framework to learn environmental dynamics and reward functions from offline data without manual definition.
Applies deep RL to dynamic origin-destination matrix estimation in traffic simulations, addressing credit assignment across temporal vehicle dynamics.
Proposes curiosity-driven quantized Mixture-of-Experts framework using Bayesian uncertainty for deploying neural networks on resource-constrained devices.
ContagionRL is a Gymnasium-compatible RL platform for reward engineering in spatial epidemic simulations, enabling systematic study of learned behavioral strategies.
Develops Hessian-free actor-critic algorithm for bi-level RL optimization with applications to LLM fine-tuning, addressing second-order information requirements in policy optimization.
Introduces continual learning task for GUI agents that must adapt to shifting domains and resolutions over time, identifying failure modes in existing agent methods.
Study of variance in agentic system evaluations using 60,000 trajectories on SWE-Bench-Verified, showing pass@1 estimates vary significantly across runs, questioning single-run reliability assumptions.
AceGRPO proposes adaptive curriculum learning with group relative policy optimization for autonomous ML engineering agents, addressing behavioral stagnation in LLM-based agents through RL with efficient data selection.
Framework for learning inspectable alignment through inverse RL without direct policy modification, improving reusability and transparency.
Soft advantage policy optimization using smooth gate functions instead of hard clipping for stable LLM training and reasoning.
Comprehensive benchmark comparing state space models, transformers, and recurrent networks for US power grid electricity demand forecasting.
Continual learning architecture for LLMs preventing catastrophic forgetting during sequential updates using thalamically routed cortical columns.