Local Inverse Geometry Can Be Amortized
Amortized learned surrogate (Deceptron) for solving nonlinear inverse problems by encoding curvature information into reusable reverse operator.
Amortized learned surrogate (Deceptron) for solving nonlinear inverse problems by encoding curvature information into reusable reverse operator.
Analysis of Muon optimizer showing spectral flattening mechanism enables larger learning rates and faster convergence through momentum orthogonalization.
Graph Information Bottleneck approach for multi-label graphs reducing over-squashing in deep message passing of graph neural networks.
Entropy regularization extension to multi-agent PPO addressing non-stationary observations in multi-dimensional cooperative reinforcement learning environments.
Game-theoretic analysis of collaborative multi-agent bandit learning where strategic agents balance information sharing against free-riding incentives.
Decision Pattern Shift framework analyzing how deep neural network internal decision mechanisms evolve from training to test data for understanding generalization failures.
Parameter-efficient fine-tuning method for LLMs using program memory to balance rapid adaptation and knowledge retention in continual learning settings.
Adversarial attack methods against multi-agent reinforcement learning systems using Jacobian-based gradient information to identify vulnerable communications.
arXiv paper: Hybrid Tucker-LSTM tensor network for battery state-of-charge prediction in electric vehicles.
arXiv paper: Hierarchical reinforcement learning with switching successor measures for zero-shot RL without fixed temporal abstractions.
arXiv paper: Bilingual pre-training outperforms hyperparameter tuning in data-constrained low-resource language settings.
arXiv paper: Teacher-Guided Policy Optimization for LLM distillation using Reverse KL to improve student-teacher convergence.
arXiv paper: EMO framework for progressive training of Mixture-of-Experts models addressing memory and communication efficiency bottlenecks.
arXiv paper: Generalization bounds analysis for Physics-Informed Neural Networks (PINNs) and variational variants.
arXiv paper: Chem-GMNet geometric transformer for molecular property prediction using domain-native architecture over generic SMILES models.
LightSplit privacy-preserving split learning using orthogonal projections to reduce communication overhead and prevent reconstruction attacks.
Exploration algorithm balancing uncertainty resolution with action budget using expected improvement and surprisal gating.
Contextual bandits algorithm optimized for resource-constrained devices using probabilistic learning.
Real-time AI agents using asynchronous I/O and speculative tool calling for sub-1-second latency in interactive applications.
Phasor Memory Networks architecture enabling stable backpropagation for explicit memory in language models via unitary dynamics.
Analysis of signal propagation in GNNs addressing information loss through oversmoothing and oversquashing phenomena.
Teaching framework for machine learning that accounts for deductive errors in learners like LLMs during few-shot learning.
Adversarial training approach addressing long-tail data imbalance through adaptive perturbations and theoretical analysis.
Novel encoder architecture replacing VAE encoders with diffusion models for improved latent representation learning.
Data augmentation method for offline reinforcement learning using trajectory-based techniques to train models from limited suboptimal data.
Research on warmstarting techniques for scaling language models, analyzing initialization constraints and growth strategies for training efficiency.
Data-free second-order preconditioning for differentially private deep learning without privacy budget consumption.
Pipeline combining pretrained LLM table extraction with fine-tuned small model error repair for low data requirements.
Asynchronous SGD with gradient rescaling for distributed optimization under data and system heterogeneity.
Flow-based policy for stable and expressive reinforcement learning without backpropagation through solvers.
Bijective representation learning framework for robust inversion of continuous forward processes.
Improved delta rule with online preconditioning for linear attention in state-space models.
Method for discovering hidden miscalibration in model confidence across different input types.
Analysis of how tokenization and representation choices affect transformer context window effectiveness and information exposure.
RL-based LLM approach for high-level synthesis code generation using comparative rewards for quality optimization.
Inference-time alignment technique using temperature adjustment to mitigate reward hacking in LLM outputs.
Foundation model for dynamic graphs across multiple domains using decoupled prompts for multi-domain pretraining.
Self-supervised contrastive reinforcement learning algorithm for discrete action spaces without hand-crafted rewards.
Neural Low-Degree Filtering (Neural LoFi) provides spectral theory of hierarchical feature learning in gradient-based deep learning.
Proves first Õ(ε⁻²) sample complexity for single-loop off-policy actor-critic methods under minimal assumptions.
Reward-Decorrelated Policy Optimization (RDPO) for multi-task and mixed-reward reinforcement learning with heterogeneous rewards.
Geometric and spectral study of low-rank pre-training for LLMs, comparing generalization to full-rank training beyond perplexity.
Three-stage learning approach for long-term time series forecasting using simple linear models and MLPs without complex architectures.
Proposes sampling method for flow language models using marginal-conditioned bridges to preserve posterior marginal structure.
Establishes scale-sensitive generalization of PAC learning fundamental theorem using fat-shattering dimension at optimal scales.
Analyzes hierarchical synthetic languages with exact k-gram ansatz to derive explicit scaling laws and benefits of reasoning in transformers.
Analyzes regret in online learning over combinatorial actions using convex relaxations, governed by polyhedral instability.
MILM extends large language models to handle multimodal irregular time series with asynchronous observations and textual channels.
Derives tight sample complexity bounds for best-policy identification in risk-sensitive reinforcement learning with entropic risk measures.
Learns POMDP world models from observation-action trajectories using language model priors for agents in partially observable environments.