Continuous-time q-learning for mean-field control with common noise, part-II: q-learning algorithms
arXiv paper on Q-learning algorithms for mean-field control with common noise in multi-agent settings.
arXiv paper on Q-learning algorithms for mean-field control with common noise in multi-agent settings.
arXiv paper on dialogue models with intent prediction using lightweight Temporal Bayesian Networks injected via prompts.
Perturbation probing identifies causal circuit structures in LLM feed-forward networks across 13 models, revealing opposition and other behavioral circuits from RLHF.
Cross-architecture analysis of adversarial attack transferability in vision-language models for autonomous driving systems.
Analyzes information contamination propagation in multi-agent workflows with heterogeneous data sources, showing how uncertainty redirects agent decision-making and execution paths.
Studies order sensitivity in LLM-based recommendation reranking and proposes position-invariant listwise ranking to fix decoder-only model vulnerabilities.
MEDS dataset maps how 14 LLM families reason about mathematics across 28,000 personas to study model mathematical abilities and biases for education applications.
Cognitive Digital Shadows dataset analyzing how 19 LLMs debate societal issues under different persona and sociodemographic prompts.
TFM-S3 method uses tabular foundation models to improve global exploration in continuous control robot policy learning.
Open-source Python library extracting subjects, verbs, and objects from text using cognitive network science and AI for NLP tasks.
Open-source Python toolkit using machine-learned potentials for automated phonon analysis and structural remediation of crystalline materials.
SISA-based architecture for machine unlearning enabling removal of class data from trained neural networks.
Privacy-preserving federated fine-tuning method addressing noise-induced prototype degradation with local differential privacy.
TwinGate defense mechanism against decompositional jailbreaks in LLMs using stateful asymmetric contrastive learning.
Empirical comparison showing in-context prompting with self-orchestration outperforms external agent frameworks for procedural tasks.
Mixture of Experts approach for semi-supervised inference combining multiple diverse prediction models with limited labels.
Conformal Abstention framework using conformal prediction to detect when LLMs lack knowledge and should abstain.
D3-Gym dataset with 565 verifiable environments for evaluating language models and agents on scientific discovery tasks.
TopBench benchmark for evaluating LLMs on implicit prediction and reasoning over tabular data.
DEFault++ automated fault detection system for identifying component-level failures in transformer architectures.
Token-aware clustering and hierarchical indexing method for efficient multivector retrieval models used in NLP applications.
Sequential inference methods for Gaussian Processes from signal processing perspective, adapting ML models for streaming/online scenarios.
ML classification of Vicsek flocking model phase diagram using K-means clustering on dynamical observables across parameter space.
PhyCo framework adds physical constraints to video diffusion models for realistic object motion, collision physics, and material responses in generated videos.
Scalable methodology for generating synthetic computer environments with realistic folder hierarchies and artifacts for productivity task simulation.
Physics-constrained inverse reinforcement learning using Fokker-Planck equation to infer reward functions from trajectories without known dynamics.
Analysis of LLM pruning showing small-magnitude pre-trained weights are critical for downstream task performance, contradicting redundancy assumptions.
Framework for variational inference in Bayesian neural networks with heteroscedastic uncertainty estimation.
VERA system for generating visual explanations of 2D embeddings from dimensionality reduction techniques via region annotation.
Network pruning framework for reducing parameters and computational costs in deep neural networks for edge device deployment.
Gradient-based optimization on Gödel logic for neurosymbolic systems, enabling discrete logical reasoning with continuous optimization.
Theoretical analysis of scaling laws in deep learning through implicit bias, explaining power-law relationships between model performance and resource growth.
Training expressive RL policies like diffusion and flow-matching models using online RL with offline data, addressing gradient stability challenges in complex policy parameterizations.
arXiv paper on evaluating AI systems without ground truth using information theory and mutual information, addressing adversarial manipulation in agent evaluation.
arXiv paper on constraint-aware flow matching for generative models addressing constraint violations.
arXiv paper presenting AEGIS, an edge augmentation framework for link prediction in sparse bipartite knowledge graphs.
arXiv paper on activation function design's role in preventing plasticity loss during continual learning.
ActiNet open-source tool using self-supervised deep learning and HMMs for activity intensity classification from wrist accelerometry.
Unified framework for neural network compression (pruning, quantization, low-rank) using vanishing contribution analysis.
NashPG policy gradient method for finding Nash equilibria in two-player zero-sum imperfect-information games.
Mixed precision training schemes for neural ODEs balancing low-precision efficiency with stability and accuracy.
Transformer Semantic Genetic Programming using pre-trained transformers as variation operators for symbolic regression.
Study of privacy, robustness, ethics and fairness implications of low-rank factorized LLM compression techniques.
Physics-Geometry Operator Transformer addressing geometric aliasing in complex PDE modeling on unstructured meshes.
DeepWeightFlow generative model using re-basined flow matching for generating complete neural network weights including symmetries.
Taxon framework using LLM expert guidance for hierarchical tax code prediction in e-commerce compliance automation.
VaR-CPO algorithm for sample-efficient constrained reinforcement learning with Value-at-Risk objectives and safe exploration.
BicKD bilateral contrastive knowledge distillation framework combining sample-wise and class-wise probability alignment.
Hinge Regression Tree using Newton methods for learning high-quality oblique decision tree splits with multivariate boundaries.
CausalCompass benchmark framework for evaluating robustness of time-series causal discovery methods under model misspecification.