Policy-based Tuning of Autoregressive Image Models with Instance- and Distribution-Level Rewards
Reinforcement learning approach for training autoregressive image models with policy-based tuning optimizing quality and diversity simultaneously.
Reinforcement learning approach for training autoregressive image models with policy-based tuning optimizing quality and diversity simultaneously.
AI system for automated spectroscopy interpretation in scientific discovery, reducing human bias in spectral analysis.
Polaris framework enabling self-improving agents for small language models through policy repair via experience abstraction and code modifications.
Adaptive prompt routing mechanism for selecting appropriate LLM or generative model based on input prompts, balancing fidelity and diversity.
Sparse packing format and CUDA kernels leveraging unstructured sparsity in LLM feedforward layers to reduce computational costs and model size.
Tight upper bounds on sample complexity for multi-group learning using one-inclusion graph prediction strategy and bipartite matching.
Foundational theoretical framework for learning under regime variation where learner, memory state, and evaluation conditions evolve over time.
Offline reinforcement learning method using guided expectation-maximization for action selection from multimodal action distributions in fixed datasets.
Model-based reinforcement learning using neural ODEs and SDEs to capture stochastic dynamics in fully and partially observed environments.
End-to-end reinforcement learning framework for heterogeneous DAG scheduling with gap-aware generation enabling rapid schedule adaptation across environments.
Mechanistic interpretability framework identifying and attributing safety circuits in LLMs responsible for alignment, jailbreak, and backdoor behaviors.
Off-policy value-based reinforcement learning framework for LLMs enabling improved data utilization and sample efficiency for long-horizon tasks.
Length-aware scheduling method accelerating reinforcement learning training for LLMs by optimizing rollout phase efficiency during chain-of-thought generation.
Continual learning framework using mixture-of-experts with similarity awareness for data-efficient adaptation to new tasks with limited samples.
Computationally efficient reinforcement learning algorithm for linear function approximation in MDPs satisfying linear Bellman completeness.
Federated learning approach combining differential privacy and Byzantine robustness to protect against both data leakage and adversarial server attacks.
Systematic evaluation of prompting strategies (zero-shot, few-shot, chain-of-thought) for chart question answering across GPT-3.5, GPT-4, and GPT-4o models on ChartQA dataset.
TIPS framework improves RL training for search-augmented LLMs via turn-level reward shaping, addressing sparse rewards and credit assignment in reasoning tasks.
Multi-agent reinforcement learning agents develop efficient private communication protocol; performance drops with human-comprehensible language enforced.
CHANRG benchmark reveals limited generalization of RNA secondary-structure prediction models. 170K structured RNA families dataset.
Quantitative assessment of reference retrieval errors from 5 LLM platforms on 2,000 medical literature references. Evaluates Grok-2, ChatGPT, Gemini, Perplexity, DeepSeek.
Theoretical analysis of low-rank knowledge distillation for LLMs with convergence and generalization guarantees. Covers compression techniques for efficient deployment.
Framework for computational arbitrage in AI model markets where arbitrageurs allocate inference budget across competing providers to undercut pricing.
First system enabling fully homomorphic encryption for end-to-end mmWave radar sensing with composable FHE kernels for signal processing and ML inference.
Token-level analysis of distributional shifts during RLVR fine-tuning of LLMs, examining mechanisms underlying reasoning improvements.
Functional component ablation framework analyzing specialization in hybrid language models combining attention with state space models or linear attention.
Verifiable synthetic benchmark for LLM-based insider threat detection using deterministic simulation engine to maintain ground truth and cross-artifact consistency.
Differential privacy framework for RLHF fine-tuning that decouples reward learning to preserve user privacy in LLM preference-based training.
Systematic benchmark comparing four multi-agent LLM orchestration architectures for financial document processing with cost-accuracy tradeoffs and scaling strategies.
Method leveraging intermediate layer representations in LLMs via Inter-Layer Structural Encoders to improve task-specific predictions beyond final-layer features.
Active learning approach using Rashomon ensemble for interpretable decision tree induction with direct hypothesis space characterization.
Quantitative model predicting when independently fine-tuned specialist LLMs can be fused post-hoc for improved performance using divergence metric.
Method addressing over-fragmentation in video object-centric learning through reconstruction-guided slot curriculum training approach.
Research on whether LLMs' step-by-step reasoning is genuinely used or post-hoc narrative generation through step-level evaluation of frontier models.
arXiv paper on brain-inspired object detection co-design. Algorithm-architecture optimization for CLIP-based task-oriented detection on edge devices.
arXiv paper analyzing hierarchical reasoning models for LLMs. Mathematical theory of recursive networks for algorithmic reasoning.
arXiv paper on black-box domain adaptation using dual-teacher distillation. Technical ML research on knowledge transfer without source access.
Critical review framework evaluating membership inference attacks and conditions under which they pose genuine privacy threats to ML models.
Investigation of LLMs' reasoning and optimization capabilities under physical and operational constraints using Optimal Power Flow problems.
Object detection framework combining YOLOv10 with Kolmogorov-Arnold networks and vision-language models for interpretability.
Zero-shot late fusion method combining audio-language models with specialist models for speech emotion recognition.
Systematic literature review of machine learning approaches for early detection of burnout in software engineers.
Scalable foundation model for automated knowledge graph generation from scientific literature using domain-specific optimization.
Active learning method leveraging vision-language foundation models for data-efficient visual recognition.
Safety monitoring approach for LLMs using activation watermarking to detect adaptive adversarial attacks during inference.
Evaluation of LLMs' ability to mimic authorial styles of literary and political figures using zero-shot prompting.
Framework for sharing memory systems across heterogeneous LLM-based agents via contrastive trajectory distillation to improve knowledge reuse.
Study evaluating whether six LLMs can emulate emotional expression and personality traits across English and Arabic languages.
Research on query-efficient jailbreak fuzzing for LLMs that identifies token importance during prompt mutation to reduce redundant searching under query constraints.
Perceptual optimization strategies for 3D Gaussian Splatting using distortion losses validated via large-scale human evaluation.