From Few-Shot to Zero-Shot: Towards Generalist Graph Anomaly Detection
Proposes generalist graph anomaly detection method enabling zero-shot transfer across datasets without dataset-specific training.
Proposes generalist graph anomaly detection method enabling zero-shot transfer across datasets without dataset-specific training.
Applies lottery ticket hypothesis to Bayesian neural networks, finding sparse subnetworks for uncertainty quantification with reduced computational cost.
Spectral graph neural networks via Cauchy factorizations enabling local-to-global spectral methods without expensive eigenbasis computation.
Rank-aware concentration inequality for attention logit magnitudes in transformers, improving stability of low-precision training via spectral bounds.
Analysis of task complexity measurement via random policies in non-tabular RL domains, examining statistical and information-theoretic metrics.
Variational Bayes-adaptive planning for deep reinforcement learning, balancing exploration-exploitation with scalable belief-state computation.
Theoretical analysis of boosting for vector-valued prediction and conditional density estimation under general divergences and stability conditions.
PCA-VAE replaces vector quantization with differentiable online PCA bottleneck via Oja's rule, eliminating codebook collapse and straight-through estimators.
Trustworthy Unified Explanation framework for interpreting LLM reasoning, revealing stability and systematic failure mechanisms across instances.
Generative recommendation framework using multi-modal LLMs to mine deep multi-interests beyond shallow behavioral signals for semantic ID prediction.
Framework for calibrating AI benchmark performance against world population baselines to provide human-anchored capability scales.
Membership inference attacks on ML models using model extraction in label-only settings without access to confidence scores or shadow models.
Theoretical analysis of gradient descent convergence rates for separable logistic regression under large step sizes and unstable regimes.
Theoretical analysis of neural network training complexity under Real-RAM vs bit-model computation, proving ERM for simple networks is ∃ℝ-complete.
Proposes Active Data Reconstruction Attack (ADRA) for detecting LLM training data through active model manipulation rather than passive membership inference.
Applies generative RL to inverse lithography for semiconductor manufacturing mask synthesis, replacing deterministic approaches with conditional sampling.
Analyzes model collapse in image generative models through iterative feedback loops using Markovian framework, revealing neural resonance phenomena in latent space.
Addresses intransitive preferences in LLM fine-tuning via preference learning, proposing methods to handle cyclic preference conflicts in multi-objective optimization.
Inverse distillation for diffusion language models. Accelerates discrete diffusion models for faster text generation inference.
Dynamic sample pruning for spatio-temporal forecasting. Optimizes training data efficiency for deep learning on large datasets.
Robust Bayesian random feature regression with contaminated priors. Studies double descent phenomenon under model misspecification.
Influence functions for detecting labeling bias in datasets. Addresses fairness issues from biased data collection.
Test-time learning method for causal structure discovery from interventional data. Combines test-time training with causal inference.
Celo2 learned optimizer with improved meta-generalization. Aims for practical adoption of learned optimization rules beyond hand-designed optimizers.
Analysis of how transformers learn sparse attention patterns incrementally. Studies information integration from multiple past positions.
Virtual Parameter Sharpening for inference-time reasoning enhancement. Dynamic low-rank perturbations for transformer adaptation without persistent parameters.
Theoretical analysis of realizable online regression under metric-like losses. Studies ReLU networks in adversarial setting.
Adaptive problem generation via symbolic representations for training small open-weight LMs on math tasks. Data generation using RL with verifiable rewards.
Dynamic rollout allocation and policy optimization for LLM reasoning with verifiable rewards. Improves RL training efficiency for reasoning tasks.
Combinatorial interpretability framework for understanding knowledge persistence in unlearning. Studies how information is retained in foundation models.
Evaluation of SAP's RPT-1 tabular foundation model on enterprise data. Compares in-context learning vs traditional ML on structured datasets.
Gradient-based optimization for explainable neuro-fuzzy systems. Addresses accuracy-explainability tradeoff in fuzzy AI models.
RL-steered graph diffusion for neural architecture search. Uses reinforcement learning to guide generative models for DAG-based NAS.
Addresses preconditioner drift in federated second-order optimizers on non-IID data through curvature alignment techniques.
Bridges GMPO and SAPO by combining sequence-level importance sampling with soft clipping alternatives for improved LLM policy optimization.
Benchmarks graph coarsening trade-offs for GNNs applied to clock tree synthesis in electronic design automation.
Introduces H-GRAMA for merging heterogeneous GNN architectures through routing and message aggregation without retraining.
Proposes Soft Adaptive Policy Optimization (SAPO) replacing hard clipping with smooth sigmoid gate functions to stabilize LLM training and reasoning in GRPO framework.
Framework combining active perception and disentangled representations for continual and few-shot learning without destructive interference.
Method for training LLMs to reason using off-policy reinforcement learning, addressing policy lag in distributed training architectures.
Analysis showing isotropic Gaussian representations improve stability in deep RL under non-stationary training dynamics.
Spiking graph neural network with predictive coding framework for out-of-distribution generalization in dynamic web environments.
Federated learning approach for causal representation learning in state-space systems enabling decentralized counterfactual reasoning across networked assets.
Benchmark combining general reasoning LLMs with domain-specific time-series knowledge for improved time-series diagnostic reasoning tasks.
Evaluation of conformal prediction methods for EEG classification handling distribution shifts in healthcare without standard i.i.d. assumptions.
Interactive browser-based educational platform for learning Federated Learning concepts with real-time visualization of heterogeneous data effects.
Investigation of grokking phenomenon in neural networks learning multiplication in finite-dimensional algebras beyond group operations.
Theoretical analysis of sample complexity bounds in replicable realizable PAC learning using Cayley graphs and spectral analysis.
Leap+Verify applies speculative execution to accelerate neural network training by predicting and validating future weights across detected regimes.
Compares XGBoost, Random Forest, and TabNet for radiation dose estimation in nuclear safety using interpolation-driven ML approaches.