Rainbow-DemoRL: Combining Improvements in Demonstration-Augmented Reinforcement Learning
Rainbow-DemoRL combines multiple demonstration-augmented reinforcement learning strategies to improve sample efficiency using offline data.
Rainbow-DemoRL combines multiple demonstration-augmented reinforcement learning strategies to improve sample efficiency using offline data.
CarbonEdge framework for carbon-aware deep learning inference at network edge, extending model partitioning to optimize environmental impact alongside latency.
High-performance engine for low-bit matrix-vector multiplication enabling efficient inference in neural networks, vector databases, and LLMs.
Open-source benchmark evaluating four AI-powered people search platforms across 119 queries for recruiting, sales, and expert search use cases.
Introduces Hidden Ads backdoor attack class exploiting Vision-Language Models' recommendation behavior to inject unauthorized advertisements through natural triggers.
Extends Bellman Deviation Detection framework for model-free RL to detect man-in-the-middle attacks in cyber-physical systems with refined MDP attack models.
RTLSeek uses multi-stage reinforcement learning to improve LLM-based RTL/Verilog generation with diverse hardware design implementations.
Composer paradigm for test-time instance-specific parameter composition enabling adaptive generative models.
Neural Gaussian mixture model using energy score guidance for predictive uncertainty quantification in machine learning.
LVRPO framework for language-visual alignment in multimodal foundation models using GRPO for understanding and generation.
KAT-Coder-V2 agentic coding model using five expert domains with specialized fine-tuning and unified distillation for software engineering tasks.
GPU-accelerated JAX library for SGP4 orbital propagation of mega-constellations enabling efficient space situational awareness.
ImagenWorld benchmark with 3.6K condition sets for stress-testing image generation models across six core tasks with human evaluation.
Physics-informed neural networks framework (Deflation-PINNs) that identifies multiple distinct solutions to nonlinear PDEs.
Privacy-preserving federated learning framework using flow-matching generation to improve robustness and aggregation in distributed training.
Analysis of prompt injection attacks against LLM agents, tracking attack pipeline stages and defense mechanisms across five frontier models.
Modular transformer approach for efficient domain adaptation in optical character recognition with reduced computational requirements.
Research on how LLMs perform scientific reasoning tasks and how prompting affects their internal reasoning processes.
Deep reinforcement learning framework using PPO to train virtual agents for guiding fish school collective motion.
Multi-agent pipeline for literature analysis using Deleuzian ontology to identify non-linear patterns in research landscapes.
Framework for evolutionary GPU kernel optimization using evaluation-driven agent and evolutionary techniques for operator generation.
Analysis of prompt framing artifacts in vision-language model evaluation on clinical neuroimaging tasks.
Survey of network performance modeling approaches comparing traditional simulation with deep learning methods.
Diffusion distillation approach using reinforcement learning to improve student model performance beyond teacher anchoring.
Medical AI Scientist system that autonomously generates hypotheses, conducts experiments, and writes manuscripts in clinical medicine.
Genetic programming pipeline for automatically evolving interpretable composite features for music tagging tasks.
Method for improving diversity in text-to-image diffusion transformers through contextual space repulsion to address typicality bias.
Study on few-shot learning and RNNs applying asymptotic equipartition property from information theory to machine learning.
Theoretical analysis of inexact Langevin algorithm convergence for score-based generative models with KL divergence guarantees.
Survey on continual graph learning covering incremental learning from streaming graph data with experience and generative replay approaches.
Mathematical analysis of auto-differentiation reliability in neural-ODE training with high-order numerical methods.
Novel prior learning method for neural networks using structured posteriors to improve generalization and uncertainty estimation.
Proves asymptotic optimality of new restless bandit policies with O(1/√N) gap under unichain and aperiodicity conditions.
Theoretical analysis of sample complexity for model-based Q-learning, establishing finite-time convergence bounds for model-learning algorithms.
Paper proposing Explaining-Away Variational Autoencoders to improve uncertainty representations in deep generative models for visual inference tasks.
Survey of multimodal continual learning methods that enable models to learn from new data across multiple modalities while retaining previous knowledge without catastrophic forgetting.
Transformers learn variable-order Markov chains in-context with finite-sample accuracy analysis using context-tree weighting.
Wavelet subspace compression for optimizer states reduces memory during LLM training, improving upon low-rank approaches.
Steering vectors applied to LLM activations for bias mitigation across social dimensions like age, gender, and race.
Survey of LLM integration with Computer-Aided Design tools, covering applications in 3D modeling and design workflows.
Foundation models for time-series prediction often use simple parroting strategies rather than learning physics, revealing shared failure modes.
Structured Agent Distillation compresses large LLM-based ReAct agents into smaller models while preserving reasoning and action consistency.
CoDec kernel optimizes LLM decoding by sharing prefix computation across multiple prompts to reduce memory-intensive KV cache access.
MicroMix: mixed-precision quantization method using microscaling formats for efficient LLM inference on NVIDIA Blackwell hardware.
Novel fine-tuning mechanism for LLMs that addresses data quality/volume issues through controlled forgetting to improve domain adaptation.
PENGUIN: Transformer variant with periodic-nested group attention mechanism for improved long-term time series forecasting.
Empirical study of initialization schemes for Kolmogorov-Arnold Networks, proposing theory-driven approaches to improve training of spline-based KANs.
Training-free framework for deferring predictions to multiple experts using conformal prediction without retraining.
ReTrack enables data unlearning in diffusion models via importance sampling to remove memorized training data influence.
Algorithms for distributed RL with policy gradients under asynchronous parallel computation and communication.