Generalization Bounds and Statistical Guarantees for Multi-Task and Multiple Operator Learning with MNO Networks
Statistical analysis of multi-task and multiple operator learning architectures with generalization bounds and theoretical guarantees.
Statistical analysis of multi-task and multiple operator learning architectures with generalization bounds and theoretical guarantees.
Self-improving world models using forward-inverse asymmetry to improve robustness across suboptimal actions for policy evaluation and planning.
RL post-training framework for building general-purpose reasoning models across diverse domains with verifiable rewards, addressing multi-domain optimization challenges.
Active learning method using feature weighting for regression tasks to optimize sample selection from unlabeled pools under budget constraints.
ArXiv paper on dynamic weight generation for recursive transformers using input-conditioned LoRA modulation controller.
ArXiv paper on fast SVD-based compression for large language models without retraining, addressing distribution shifts.
ArXiv paper on graph attention network for multi-sensor object fusion and tracking in autonomous driving.
ArXiv paper introducing Universal Hypernetworks that generate weights for arbitrary model architectures using descriptors.
ArXiv paper applying diffusion denoising objectives to causal structure learning from observational data.
ArXiv paper on model-based reinforcement learning for control systems with time-varying dynamics.
ArXiv paper on in-context agentic reinforcement learning enabling LLM agents to internalize skills at inference time.
ArXiv paper on lightweight diffusion transformer for crystal structure generation using subatomic tokenization.
ArXiv paper unifying group-relative and self-distillation policy optimization for LLM post-training with improved credit assignment.
ArXiv paper proposing Head-Calibrated Clipped-Linear Softmax as efficient surrogate for attention softmax in edge inference.
ArXiv paper on exact parameterization of doubly stochastic matrices for learned mixing in neural networks.
Single-stage training paradigm for efficient LLM reasoning that reduces token consumption in chain-of-thought without degrading quality.
Learning-based cooperative coevolution framework addressing heterogeneous large-scale global optimization via adaptive low-dimensional optimizers.
Neural-symbolic framework for discovering constitutive closures in nonlinear PDEs from spatiotemporal data while avoiding spurious physical recovery.
Research on regularizing attention scores in vision transformers using bootstrapping to improve interpretability and reduce noisy attention maps.
Analysis of safety, security, and cognitive risks in world models used for autonomous decision-making in robotics, autonomous vehicles, and agentic AI systems.
Study of reliability gaps in AI-assisted medication systems, highlighting risks in healthcare decision support.
Framework combining LLMs with infeasibility detection for NP-hard combinatorial optimization problems.
Equivariant transformer architecture for modeling agent behaviors in autonomous driving with SE(2) symmetry.
New optimizer deriving design principles from Muon, improving LLM training efficiency through surrogate model analysis.
Method for efficiently adapting closed-box LLM APIs to target tasks by priming followed by local optimization.
Benchmark dataset for evaluating AI coding agents based on production workloads, addressing language distribution and codebase structure gaps.
Study of interactions between normalization methods and optimizers in LLM training at 1B parameters.
LiteInception: lightweight interpretable deep learning framework for fault diagnosis on edge devices.
LiveMathematicianBench: benchmark for evaluating LLM mathematical reasoning capabilities with proof sketches.
Study on language pre-training bias improving performance on general vision tasks through cross-modality transfer.
Analysis of permutation-invariant discrete representation learning for spatially aligned images using vector quantization.
Woosh: open-source sound effects foundation model from Sony AI with architecture, training details, and benchmarks.
Theoretical analysis of multi-head self-attention transformers using particle systems and homogenization limits.
Curia-2 foundation model using self-supervised learning for medical imaging analysis on CT and MRI data.
Comparison of centralized and decentralized RL controllers for traffic signal control in urban corridor networks.
Mining instance-centric vision-language contexts for human-object interaction detection, leveraging VLMs to improve semantic understanding and contextual reasoning.
LatentUM unified model for interleaved cross-modal reasoning combining visual understanding, generation, and world dynamics in latent space.
Hybrid framework combining LSTM workload prediction with game-theoretic heuristics for cloud cost optimization during dynamic workload changes.
AstroConcepts corpus of 21,702 astrophysics abstracts for multi-label classification research addressing extreme class imbalance with specialized terminology.
Shared task participation comparing lexical and contextual approaches for cross-document software mention coreference resolution under mention noise.
Analysis of Mixture-of-Experts LLM interpretability at expert level, comparing sparsity properties to dense feed-forward networks using probing methods.
Characterization of exact Pareto fronts in average-cost multi-objective MDPs, extending prior work on discounted settings.
Method for verifying operational design domain coverage for safety-critical AI systems in aviation, addressing EASA certification requirements.
ASK framework combines smaller language models with RL policies to enhance out-of-distribution generalization by gating LM assistance based on agent uncertainty.
Extended publication on learning state machines from data streams with PAC-bounds analysis and heuristic improvements.
Bayesian vertical federated learning approach for multimodal survival prediction with privacy preservation and uncertainty quantification.
De Jure is an automated pipeline for extracting structured regulatory rules from legal documents using LLM self-refinement, requiring no human annotation or domain-specific data.
Analysis of token initialization strategies for new vocabulary in language models used for generative recommendation systems.
Theoretical analysis showing transformers can solve non-linear non-Markovian stochastic filtering for conditionally Gaussian signals.
Moonwalk: inverse-forward differentiation technique addressing backpropagation's memory requirements for training deeper networks.