From Gradients to Riccati Geometry: Kalman World Models for Single-Pass Learning
Proposes Kalman World Models, a method for training state-space models using recursive Bayesian filtering instead of backpropagation for online learning.
Proposes Kalman World Models, a method for training state-space models using recursive Bayesian filtering instead of backpropagation for online learning.
Flow matching-based approach for full-waveform inversion with generative priors to improve seismic imaging robustness.
Outcome-Aware Tool Selection method for semantic routers in LLM inference, reducing latency by offline interpolation without GPU cost.
Standardized benchmark dataset and evaluation framework for computational antibody design methods with unified metrics.
Framework for in-context learning on graphs without modality-specific encoders, enabling cross-domain adaptation for graph foundation models.
Multimodal diffusion model using flow matching for channel estimation, fusing LiDAR, camera, and location data.
Theoretical work on implementing higher-order mental-state dynamics in Transformers via triadic modulation for information pre-selection.
Framework reconciling in-context and in-weight learning in Transformers through dual representation space encoding to reduce conflict.
Probabilistic Gaussian Homotopy framework for nonconvex optimization using Boltzmann-weighted gradient aggregation.
Analysis of complex singularities in softmax cross-entropy loss that limit safe step sizes during optimization training.
End-to-end LLM method for auditing course information sheets at scale to identify GenAI vulnerabilities in academic assessments.
Survey of privacy-preserving machine learning for IoT devices, covering federated learning, differential privacy, and resource constraints.
Physics-informed CNN for precipitation nowcasting using volumetric radar data to estimate multi-altitude motion fields.
Federated learning approach for fraud detection in payment systems using NVIDIA FLARE, preserving privacy across institutions with non-IID data.
Systematic study of chemical language models for molecular property prediction, analyzing performance inconsistencies across benchmarks through controlled experiments.
SemRep: generative code representation learning using code transformations for semantic reasoning in software development.
PLUME: 140M-parameter foundation model for wireless packet traces using protocol-aware tokenization.
PDE-SSM: state-space block replacing attention in diffusion transformers using learnable convection-diffusion-reaction equations.
Dynamic curriculum learning approach using gradient-based difficulty estimation for adaptive example ordering.
GIP framework for selecting training examples for LLM fine-tuning by maximizing mutual information with task-specific signals.
IGU-LoRA: adaptive rank allocation for parameter-efficient fine-tuning of LLMs using integrated gradients and uncertainty scoring.
LLM-guided approach for interpretable dynamic graph clustering with semantic explanations of cluster formation and evolution.
Theoretical analysis extending Domingos interpolation formula to stochastic gradient descent, characterizing neural network generalization as kernel machines.
Deep metric learning approach (Siamese, Triplet, Vision Transformer) for identifying scribes in Chinese manuscript calligraphy datasets.
UVLM: unified framework for loading and benchmarking multiple vision-language models with standardized interface in Colab environment.
Self-training framework with label correction for handling noisy data in neural network training using bilevel optimization techniques.
FedPBS: federated learning algorithm addressing statistical heterogeneity and non-IID data for distributed ML training with privacy preservation.
Interpretable data augmentation for imbalanced learning that generates realistic, feasible samples with transparent procedures and adjustability.
Method for training 4-bit quantized CNNs on standard CPUs with PyTorch achieving full-precision accuracy parity for cost-effective deep learning.
CONSERVAttack method for testing high energy physics ML applications against physically motivated systematic uncertainties and adversarial robustness.
Chunk-Guided Q-Learning algorithm for offline reinforcement learning balancing bootstrapping error and policy flexibility over long horizons.
Aumann-SHAP framework for interaction-aware counterfactual explanations decomposing model transitions using cooperative game theory.
Benchmarking open-source PPG foundation models for biological age prediction, comparing task-specific vs general-purpose models across populations.
Gated graph attention networks for predicting duration of large-scale power outages induced by natural disasters.
Analysis of redundant features emerging in Transformer next-token predictors, identifying gradient components responsible for seemingly useless feature computation.
Hyperbolic control mechanism using parallel transport to steer text-to-image models away from unsafe content generation.
Framework for accelerating LLM inference using contextual sparsity predictors for ReGLU-based feed-forward networks with minimal accuracy loss.
Analysis of training-inference mismatch in neural networks with soft vs hard selection, using logic gate networks as test case.
TACTIC method for tabular anomaly detection using in-context learning with foundation models, advancing unsupervised learning for anomaly detection tasks.
Hybrid architecture combining classical ML for customer segmentation with RAG-enabled LLMs for personalized financial services marketing content generation.
Gradient modulation and projection techniques balance optimization across modalities in multimodal domain generalization tasks.
Pocket-K AI-ECG system using ECGFounder foundation model for non-invasive hyperkalemia detection with handheld deployment.
Efficient procedure for evaluating excess risk of empirical risk minimization using black-box access with minimal data and compute.
Self-Indexing KVCache predicts sparse attention from compressed keys to reduce KV cache bottleneck in long-context LLM inference.
Addresses domain skew in federated learning through feature decoupling and calibration across distributed clients with diverse data.
GoldenStart improves flow-matching RL policies via Q-guided priors and entropy control for faster inference and better exploration.
Mathematical foundation for sampling Boltzmann distributions via normalizing flows, proving existence of transport map approximations.
Unified functional analytic framework interpreting supervised and unsupervised learning as variational optimization over function spaces.
SIREN auto-decoder framework for high-fidelity compression of seismic velocity models using implicit neural representations.
Spectral clipping optimization technique for LLM training that addresses spectral norm instability and gradient noise issues in standard optimizers.