Learning under noisy supervision is governed by a feedback-truth gap
Two-timescale analysis showing feedback-truth gap governs learning under noisy supervision across neural networks and human studies.
Two-timescale analysis showing feedback-truth gap governs learning under noisy supervision across neural networks and human studies.
Verbalized Action Masking method for controllable exploration in RL post-training of LLMs with chess case study.
Residual-aware theoretical analysis explaining position bias in transformer attention mechanisms and cumulative attention rollout.
Progressive Thought Encoding method for efficient training of large reasoning models via parameter-efficient RL fine-tuning.
Mechanistic analysis of how two-layer networks learn Fourier features for modular addition with theoretical training dynamics explanation.
Framework for detecting and reducing ballast information across structured, semi-structured, and unstructured multimodal datasets.
Dementia classification model for Brazilian adults using variable selection and multivariable analysis from ELSI-Brazil dataset.
Mixed-integer programming framework providing sound and complete certification guarantees for worst-case data poisoning attacks.
Symbolic alternative to GNNs with improved expressivity beyond 1-Weisfeiler-Lehman barrier and fine-grained interpretability.
NSGGM: neuro-symbolic framework for molecule generation combining neural proposals with symbolic constraints for controllability.
Communication-free decentralized multi-agent bandit protocol with Lipschitz-structured action spaces and hard collision constraints.
Unified framework for exploiting locality in scalable multi-agent RL with relaxed conditions on exponential decay property.
Studies grokking transition via loss-landscape geometry on sequence-learning tasks SCAN and Dyck-1 using commutator defect metrics.
Fail-closed alignment design principle for robust LLM safety through redundant refusal mechanisms across latent features.
UniLeak: mechanistic interpretability framework identifying universal activation directions that trigger PII leakage in language models.
Dynamic Delayed Tree Expansion improves multi-path speculative decoding for faster LLM token sampling verification.
Action Graph Policies: framework for learning action dependencies in multi-agent reinforcement learning to coordinate agent behavior.
WS-GRPO: weakly-supervised training method for language models on complex reasoning with improved rollout efficiency.
Using in-context learning and AI-enhanced tensor methods to automate behavioral neuroscience discovery pipelines.
Uncertainty-aware time-series ensemble method for proactive anomaly prediction providing early warning signals before anomalies occur.
Multi-agent reinforcement learning framework using spatio-temporal hypergraphs for human-centric multimodal traffic signal control.
Graph neural network architecture with adversarial synthesis and contrastive learning for resilient node classification under structural noise.
NAMO optimizer combining Adam's adaptive moments with Muon's orthogonalized momentum for efficient large language model training.
Machine unlearning approach using target feature disentanglement to balance privacy and model utility for right to be forgotten compliance.
Formal framework defining locality radius to determine when multi-hop reasoning is necessary for foreign key discovery in relational databases.
Federated fine-tuning approach for LLMs using low-rank Gram matrices and Procrustes alignment to enable collaborative adaptation without data sharing.
Serverless MLOps framework orchestrating complete ML lifecycle with event-driven pipelines for model-agnostic inference and rapid deployment.
Online learning framework for multiclass settings where agents can improve features to achieve better labels with budgeted costs.
Vector Perturbation VAE decouples representation learning from discretization by eliminating explicit codebook dependency in generative modeling.
Analysis of underfitting challenges in multi-expert learning to defer systems where classifiers abstain and defer to experts.
Gradient-free zeroth-order optimization method for memory-efficient fine-tuning of large-scale models via subspace gradient orthogonalization.
Empirical study comparing how transformers and linear attention models perform in-context learning on regression tasks, analyzing MSE, convergence, and generalization.
Deep reinforcement learning with domain randomization for robust control of mechanical systems with multiple uncertainties and nonlinear dynamics.
Open-source PyTorch library implementing GPU-accelerated Soft Dynamic Time Warping with improved memory efficiency and numerical stability.
ArXiv paper on CounterFlowNet for generating multiple minimal counterfactual explanations for tabular data with heterogeneous features.
ArXiv research on adapting EEG foundation models with limited supervision using prototype-guided fine-tuning for clinical settings.
ArXiv paper applying Wasserstein Autoencoders to control laser pulse shapes in Free-Electron Lasers via differentiable latent interface.
ArXiv research on Unified Latents framework combining diffusion priors with diffusion decoders for efficient latent representation learning.
ArXiv paper proposing RL framework for extremal graph theory problems, extending Deep Cross-Entropy methods to combinatorial optimization.
LexiSafe: Offline safe reinforcement learning method using lexicographic hierarchy to prevent safety violations in cyber-physical systems.
Research paper introducing Flickering Multi-Armed Bandits framework where available actions change dynamically, modeled via random graph processes.
Machine learning framework using deep learning on carotid ultrasound videos to detect vascular damage for cardiovascular disease risk assessment.
Subquadratic structure inference pipeline for immune repertoire analysis combining retrieval, affinity kernels, and multimodal fusion.
Test-time training method using prompts for Graph Neural Network out-of-distribution detection without accessing training data.
Study of shortcut learning in neural networks applied to knot topology classification with applications to protein folding and polymer physics.
ML framework for biomedical data emphasizing feature stability and interpretability under incomplete/missing data for trustworthy clinical decision-making.
Formulates MDP planning as Bayesian policy inference with return-based optimality prior for discrete action domains.
Uses LLMs for temporal credit assignment in self-evolving agents via retrospective in-context learning from sparse feedback.
Introduces CRAFT, parameter-efficient fine-tuning method using Tucker decomposition on frozen pre-trained transformer attention weights.
Identifies transformer attention heads functioning as membership testers/bloom filters across GPT-2 and Pythia models.