OmniTabBench: Mapping the Empirical Frontiers of GBDTs, Neural Networks, and Foundation Models for Tabular Data at Scale
OmniTabBench: Large-scale benchmark comparing GBDTs, neural networks, and foundation models on tabular data with 100+ datasets.
OmniTabBench: Large-scale benchmark comparing GBDTs, neural networks, and foundation models on tabular data with 100+ datasets.
STQuant framework for adaptive quantization of optimizer states during large multimodal model training to reduce memory costs.
Decentralized multi-agent RL approach for vehicle-to-infrastructure systems using equivariant neural networks.
Efficient scaling technique for diffusion RL post-training using low-precision exploration and higher-precision training.
Neural method for learning search policies in Traveling Salesperson Problem, training models to iteratively improve solutions.
Transfer learning formalism using Outcome-Predictive State Representations for knowledge generalization across RL tasks.
Nonstationary classification approach using learned retrieval to condition classifiers on historical examples beyond training cutoff.
Study of expert specialization and routing behavior in sparse Mixture-of-Experts architectures for large language models at small scale.
Offline reinforcement learning approach addressing epistemic uncertainty through ensemble-based conservative value estimation.
Method to enhance LLM task performance by amplifying task-relevant neurons at inference time without parameter modification.
Theoretical framework addressing catastrophic forgetting in continual learning through informational structural alignment rather than external mechanisms.
Research on using multi-turn reasoning LLMs with deep reinforcement learning for task offloading decisions in mobile edge computing systems.
Research on calibrating uncertainty quantification in LLMs for question-answering through token-level temperature scaling, addressing gaps in existing confidence measures.
Mixture proportion estimation from unlabeled data using conditional independence assumptions. Application to PU learning, label noise, and domain adaptation.
Computational complexity analysis of ML model expressiveness for complex systems. Studies how ML manages complexity through probability on sampleable distributions.
Theoretical study of differential privacy cost for language identification and generation. Establishes algorithms and lower bounds quantifying privacy-utility tradeoff.
Categorical framework formalizing deep learning model architectures using array broadcasting and morphisms. Mathematical notation for neural network composition.
Comparative analysis of SHAP explainability method applied to different ML models. Reviews interpretability for black-box model predictions.
Training method for Android UI agents improving RL efficiency using single state multiple actions paradigm to reduce sample inefficiency and emulator latency.
Split learning framework with frequency-aware compression reducing communication overhead in distributed neural network training on resource-constrained edge devices.
Data deletion scheme predicting model behavior after training data exclusion. Fast approximation for understanding data influence on learned models.
Multi-agent system using RL for dynamic specialist routing in medical diagnosis. LMM agents route diagnostic queries to appropriate specialists for precision diagnosis.
Theoretical analysis explaining why entropy dynamics in LLM internal representations correlate with reasoning correctness. Proposes stepwise informativeness assumption.
DOVE benchmark evaluates LLM cultural value alignment through open-ended generation. Addresses limitations of discriminative multiple-choice formats and subcultural heterogeneity.
Multi-fidelity optimization framework combining VCG incentive mechanisms with efficient sampling to optimize LLM advertising while managing advertiser strategic behavior.
Framework to distill hallucination detection signals into transformer representations during training, enabling inference-time hallucination detection without external verification systems.
FedSpy-LLM demonstrates data reconstruction attacks on LLMs in federated learning, highlighting privacy risks in gradient sharing.
WebSP-Eval benchmarks web agents on website security and privacy task execution, filling gap in agent evaluation frameworks.
ForkKV is a system for efficient multi-LoRA agent serving using copy-on-write KV cache disaggregation to reduce memory overhead.
Research analyzing whether latent chain-of-thought reasoning in LLMs actually enables superposition of multiple solutions.
ProofSketcher combines LLMs with formal proof verification to improve mathematical and logical reasoning accuracy and reliability.
Proposes TinyML-based intrusion detection for CubeSats addressing cybersecurity vulnerabilities from COTS components and open-source software.
Evaluates LLM ability to integrate long-form text information through novel summarization task, comparing human and model-authored novel summaries.
Studies how offline recommendation system performance scales with training dataset size and identifies saturation points in data effectiveness.
Operator learning surrogate model for wave-induced forces as alternative to expensive numerical wave models in storm surge prediction.
Defines learning debt and actionable staleness metrics, derives Bayes retraining rule for optimal forecasting model retraining schedules.
Activation Prompts improve visual prompting for vision model adaptation, closing performance gap between prompting and conventional fine-tuning.
Applies reservoir computing to anticipate critical tipping points in complex spatiotemporal dynamical systems via machine learning.
Soft-quantum algorithms combining quantum operations with classical simulation for variational quantum circuits on few-qubit problems.
Tensor-network autoencoder using multiscale MERA architecture for reconstruction-based anomaly detection in particle physics collider jets.
Demonstrates LLMs fail at reliable stochastic sampling required for agentic systems, identifying critical failure point in distribution sampling from inferred data.
ExplainFuzz generates test inputs using probabilistic circuits, improving on grammar-based fuzzers and LLM approaches for constraint-conditioned software testing.
Guardian Parser Pack uses LLMs for schema-guided extraction and normalization of missing-person intelligence from heterogeneous investigative documents.
DynLP algorithm for efficient parallel dynamic batch updates in graph-based semi-supervised learning label propagation with incremental data arrival.
Empirical study showing 52-88% of chain-of-thought tokens in LLMs are generated after answer is already recoverable, revealing the detection-extraction gap.
Proposes Holistic Optimal Label Selection (HopS) for prompt learning with partial labels in vision-language models using pre-trained feature encoders.
Survey of David Blackwell's mathematical theorems (Rao-Blackwell, Approachability, Informativeness) and their foundational relevance to AI.
Feature compression framework for model-specific representations; prevents cross-model transfer and unauthorized data reuse.
Watermarking technique for generated content robust against removal/forgery attacks; addresses copyright protection for diffusion models.
Foundry: CUDA graph template system for fast LLM serving cold start; reduces graph capture time from tens of seconds to milliseconds.