Taking the GP Out of the Loop
Scaling Bayesian optimization beyond Gaussian processes for problems with abundant cheap observations.
Scaling Bayesian optimization beyond Gaussian processes for problems with abundant cheap observations.
Online conformal abstention for LLM factuality control with adversarial bandit feedback in interactive systems.
InvisibleInk framework for differentially private long-form text generation with RAG and inference-time scaling.
CAPTCHA system combining GANs and RL for adaptive multi-modal verification challenges.
Aligning inductive bias in state space models for improved data-efficient generalization.
Federated unsupervised multi-source domain adaptation framework addressing scalability and privacy.
LLM agents automated data quality improvement for text-attributed graphs using graph optimization techniques.
Empirical study showing equivariant neural architectures scale better than non-equivariant models for interatomic potential learning.
Study of how prediction horizon affects learned representations in predictive models and world models.
Zeroth-order optimization for on-device fine-tuning of AI agents without backpropagation, reducing memory constraints.
Deep RL with Transformers optimizes eVTOL aircraft takeoff trajectories for minimum energy consumption.
Conformal regression method using Laplace approximation for adaptive prediction intervals with uncertainty quantification.
RL post-training with canonical action order hints improves Transformer performance on constraint satisfaction puzzles like Zebra puzzles.
Analysis of embedding condensation in small LLMs and dispersion loss technique to improve generalization.
Prism: test-time scaling method for discrete diffusion language models using hierarchical search and self-verification.
Analysis showing test-time training with KV binding operates as learned linear attention, not memorization.
S2O algorithm enabling fine-grained sparse attention via online permutation for efficient long-context inference.
Speculative speculative decoding technique accelerating LLM inference by parallelizing draft model speculation.
Method for tractable training of LLMs for scientific discovery by decomposing hypothesis generation complexity.
Self-improvement framework for LLM personalization via mutual information maximization without additional labeled data.
Empirical study on incorporating explicit physical constraints into Vision-Language-Action robot models.
Theoretical framework analyzing Q-learning through switching systems and Bellman error representation.
Extends input convex neural networks from linear programming to second-order cone programming for improved representational capacity of surrogate functions.
Hospital readmission prediction framework addressing explainability, deployment reliability, and fairness using MIMIC-IV clinical database.
Reinforcement learning meta-framework for stable biomarker discovery through iterative feature selection in high-dimensional genomic data.
Diffusion-based approach for generating continuous-time stochastic processes conditioned on partial observations using non-Markovian bridges.
Analysis showing advanced jailbreaks of frontier language models impose minimal performance degradation compared to less complex attacks.
LoRA fine-tuning method using Fisher subspace initialization to improve LLM adaptation efficiency and downstream task performance.
Large-scale benchmark evaluating docking and scoring methods including AI-based DiffDock and NMDN on LIT-PCBA library for virtual screening.
Proposes Calibrated Size Ratio metric as alternative to Expected Calibration Error for measuring neural network confidence calibration.
Diffusion-based offline safe reinforcement learning with decoupled guidance balancing reward improvement and constraint satisfaction at deployment time.
Speculative decoding optimization for LLM inference using compression-aware adaptive gamma selection to accelerate token generation.
Analysis of using blockchain proof-of-work computing resources for machine learning model training instead of hash puzzle computation.
Scoping review of deep learning methods applied to photoplethysmography signal analysis for healthcare monitoring and wearable devices.
Open-source plug-and-play multi-modal feature extraction toolkit for psychology and social science research automating human data annotation.
Meta-review and taxonomy database organizing AI risks with standardized terminology across domains including privacy, safety, and alignment.
Benchmarking study evaluating YOLOv8, EfficientDet Lite, and SSD object detection models on resource-constrained edge devices across varying complexity.
Active learning method using diverse LLM feedback for generalized category discovery, recognizing both known and novel categories in unlabeled data.
Open-source low-cost bimanual mobile manipulator platform designed for embodied AI and Vision-Language-Action model training with diverse manipulation data.
Deep learning method for sparse semantic keypoint matching between images using normalized Transformers and hyperspherical normalization strategy.
Reinforcement learning approach for code generation that optimizes reasoning quality using process-level rewards and reasoning-aware training.
Hybrid neural-symbolic approach combining neural models with logical rules for improved generalization and reasoning in syllogistic logic tasks.
Memory-augmented framework enabling LLM-based agents to learn classification tasks from labeled examples without parameter updates using semantic and episodic memory.
Research extending supervised learning to optimal control problems as alternative to reinforcement learning in non-stationary environments.
LLM-assisted tool for translating conceptual queueing system descriptions into executable simulation code with mechanism verification.
Test-Time Scaling method recycling intermediate search rollouts from LLM reasoning to improve efficiency and reduce computational redundancy during inference.
AI agent system combining LLMs with operations research algorithms for inventory control, handling demand shifts and contextual information better than traditional OR methods.
Descent-Guided Policy Gradient method for scalable cooperative multi-agent reinforcement learning, reducing cross-agent noise using differentiable analytical solutions.
HiMAC framework for long-horizon LLM agents using hierarchical macro-micro learning to improve structured planning and execution reliability.
End-to-end pipeline for language-guided robotic grasping combining vision-language models with partial observations and collision-free execution on legged manipulators.