Analysis of proto-tokens in LLM one-step text reconstruction, studying semantic and syntactic information enabling multi-token generation from single forward pass.
Deep learning architecture using orthogonal hyper-connections to maintain identity mapping in residual networks while improving training stability.
Large-scale empirical study of transformer state tracking limitations in distribution, examining induction bias and generalization beyond length extrapolation.
Explainability methods for AutoML clustering systems, analyzing how dataset meta-features influence algorithm selection and hyperparameter choices.
Federated learning optimizer addressing non-IID data distribution and client drift with efficient client-side computation.
Byzantine-resilient federated learning framework using partial model sharing and conformal prediction for robust distributed training.
ML methods for vessel power prediction using SVMs, ANNs, Random Forests, and XGBoost compared against physics-based propeller law constraints.
Unified framework for analyzing graph neural network expressivity via Weisfeiler-Leman algorithm and substructural information.
Framework deriving RNN and Transformer architectures from subgroup structures of U(d).
Analysis of noise-agnostic diffusion models showing geometric concentration enables implicit noise estimation.
K-partition ensemble method for assigning confidence scores to cluster assignments.
IARPA TrojAI program final report on detecting and mitigating backdoor trojans in AI models.
Theoretical framework explaining persistent LLM failures like hallucination and sycophancy as rational behavior from model misspecification.
Benchmark dataset for evaluating visual document processing in scientific paper retrieval and QA tasks.
Analysis of bibliographic confounding in ML models for materials discovery, showing models learn spurious patterns.
Training approach for mutual adaptation between humans and AI agents that accounts for adaptive human behavior.
Physics-informed neural networks for multiscale Darcian dynamics using neural basis method approach.
Topological analysis of empirical risk landscapes for high-dimensional Gaussian models and phase retrieval problems.
Game-theoretic analysis of competitive markets with multiple generative model platforms and heterogeneous users.
Neuro-symbolic pipeline combining language models with formal mathematical ontologies to reduce hallucinations and improve grounding in reasoning tasks.
Method using conditional diffusion models to estimate drift functions in stochastic differential equations from trajectory data.
Study of targeted bit-flip attacks on LLM weights via DRAM exploitation and defense mechanisms against parameter faults.
Benchmark study evaluating fine-grained knowledge capabilities of vision-language models on visual reasoning and classification tasks.
Theoretical study of stochastic gradient descent for learning single-index models in sequential/interactive learning settings.
Analysis of why steering vectors for controlling LLM behavior show variable reliability and proposes geometric predictors for understanding limitations.
Method for optimal data collection from multiple heterogeneous sources under budget constraints and bias considerations.
Compact Vision Transformer removing fixed positional embeddings and class tokens for improved medical imaging generalization on edge devices.
Mean-field RL algorithm handling asynchronous multi-agent environments by using alternative summary statistics when agents are idle.
Framework combining Hamilton-Jacobi reachability analysis with deep Q-learning for safe autonomous vehicle interaction with cyclists.
Method using LLM agents to generate high-quality synthetic adversarial QA data for fine-tuning domain-specific language models.
Demonstrates that unsupervised embeddings using Self-Organizing Maps encode sensitive attributes like age and income despite explicit exclusion.
Simplified Vision-Language-Action baseline for robotic manipulation disentangling architectural innovations from training recipe contributions.
Web application for making mechanistic interpretability analyses of LLMs accessible to non-specialists through simplified interactive visualization.
Benchmark dataset of 500 Lean 4 formal verification proofs for evaluating LLM-based proof automation in software verification contexts.
arXiv paper proposing information-geometric framework for LLM in-context learning via quantum density operators.
arXiv paper applying reinforcement learning (PPO) to tune Pure Pursuit parameters for autonomous racing.
arXiv paper on equivariant neural networks for robust object recognition under symmetric transformations.
arXiv paper benchmarking graph neural networks against classical heuristics on hard constraint satisfaction problems.
Unified framework for analyzing meta-algorithms in online convex optimization across different feedback types and regret notions.
Leverages pretrained SciML foundation models for efficient inference of neural fluid fields from sparse real-world flow data.
Uses graph neural networks with co-evolutionary analysis to predict metal-binding residues in proteins for structural biology applications.
Evaluates efficiency of one-run auditing methods for certifying privacy guarantees in machine learning algorithms.
Formalizes generalization error bounds using Rademacher complexity and Dudley's entropy integral in Lean, advancing theoretical foundations of statistical learning.
Studies how fine-tuning changes model representations using crosscoders, a model diffing method that learns shared interpretable concepts across base and fine-tuned models.
Approach for using visual modalities instead of text for reasoning in multimodal LLMs, especially for vision-centric tasks.
Method for making CNNs self-explainable in medical imaging by improving interpretability beyond post-hoc attribution techniques.
Framework lifting autoencoders to distribution space (GDE) for multiscale representation learning on sets of samples instead of single points.
Parameter-free optimization for Sign-SGD addressing stepsize selection for memory-efficient LLM training and gradient compression in distributed learning.
Gradient-based data attribution method learning parameter importance weights to identify influential training examples without Hessian approximations.
Graph neural network methods for multi-objective routing on multigraphs with multiple distinct edges between node pairs.