Efficient Multi-objective Prompt Optimization via Pure-exploration Bandits
Framework for multi-objective prompt optimization using pure-exploration bandits to select effective LLM prompts across multiple performance metrics.
Framework for multi-objective prompt optimization using pure-exploration bandits to select effective LLM prompts across multiple performance metrics.
Method for improving agentic reinforcement learning by optimizing token-level credit assignment in LLM trajectories using energy-based approaches.
Empirical study on aggregation strategies for visual document retrieval in RAG systems, evaluating information loss in financial document processing.
Research on backdoor attack vulnerabilities in deep reinforcement learning agents with plasticity interventions, examining security threats in DRL systems.
Theoretical analysis establishing fast convergence rates for entropy-regularized inverse reinforcement learning with linear reward classes.
Analysis showing defenses against malicious fine-tuning of foundation models fail under adaptive adversaries that account for defense mechanisms.
JetBrains IDE plugin toolkit for developers building AI features with LLMs and agentic workflows, enabling tracing, debugging, and evaluation in development loop.
SIRA: Training-free method to reduce hallucinations in vision-language models through internal contrastive decoding without external tools.
Energy efficiency analysis comparing neural combinatorial optimization solvers to CPU metaheuristics, accounting for amortized training costs.
Multi-label visual emotion analysis benchmark for evaluating multimodal large language models on image emotion prediction tasks.
AI agent framework (Automat) for automated design of material descriptors through iterative proposal, implementation, and evaluation for materials science.
Comparing neural machine translation and glossary-augmented LLM approaches for translating specialized rock art terminology in cultural heritage documents.
Benchmarking foundation models on EEG and brain-computer interface tasks with standardized datasets and metrics for clinical relevance evaluation.
arXiv paper on SceneFunRI: vision-language model benchmark for reasoning about occluded objects using spatial reasoning and context inference.
arXiv paper on IntentVLA: vision-language model for robot manipulation handling multimodal demonstrations and intent disambiguation.
arXiv paper investigating task-aware layer pruning for LLMs; shows pruning improves out-of-distribution performance while maintaining in-distribution accuracy.
arXiv paper on governance failures in LLM-based financial systems; proposes metrics for auditable decision-making compliance beyond task accuracy.
arXiv paper on Video2GUI: automated framework synthesizing large-scale GUI interaction data for pretraining generalized GUI agents using video.
arXiv paper extending intervention methods for LLMs beyond linear approaches to capture non-linear feature representations in model internals.
arXiv paper on EVA: defense against LLM jailbreaks through editing techniques to improve model safety alignment without computational overhead.
Knowledge distillation approach for identifying student misconceptions in educational settings with noisy labels.
Theoretical analysis of compositional sparsity as inductive bias enabling deep networks to overcome curse of dimensionality.
SpeechLLM system for streaming speech-to-text translation in real-time without waiting for complete utterances.
Data selection scheduling method that dynamically adjusts training data volume throughout model training.
Security analysis showing LLM browser agents can be fingerprinted through UI interaction traces and timing patterns.
LLM-based research idea generation using citation evolution graphs as structural supervision signal.
Autonomous AI agents for scientific discovery in cosmology using LLM-guided code evolution and multi-agent research labs.
Parameter-efficient fine-tuning method improving upon LoRA through isometric global parameter partitioning.
Dynamic weight quantization method for efficient LLM inference with adaptive codebook sizing and quality targets.
Multi-agent framework integrating operational plan generation and verification for complex battlefield planning.
Evaluation of whether coding agents understand least-privilege authorization principles for safe deployment.
Multi-agent AI system with recursion-of-thought for root cause localization in microservice systems.
Multi-step reasoning approach for text-to-image generation with closed-loop verification to handle complex semantics.
Analysis of noise in CLIP vision-language model embeddings using spectral decomposition of covariance matrices.
Distillation method for converting deep RL policies to interpretable surrogate models using Voronoi quantization.
MHSA: lightweight framework for mitigating hallucinations in vision-language models via steered attention mechanisms.
Viverra: text-to-code system generating verifiable, correct code with guarantees, reducing developer review burden.
MicroscopyMatching: framework for automated microscopy image analysis across diverse biological and imaging conditions.
Second-order actor-critic methods for discounted MDPs using policy Hessian decomposition for accelerated convergence.
Study quantifying premature closure in frontier LLMs—inappropriate commitment under uncertainty—with mitigation strategies.
Method for improving sample efficiency in RLVR by using randomly selected few-shot guidance for LLM chain-of-thought tasks.
COTCAgent: LLM-based clinical decision support using probabilistic chain-of-thought reasoning on longitudinal EHR data.
Generalized Priority-Aware Shapley Value: extension of Shapley value for valuation on arbitrary weighted priority graphs.
SemaTune: LLM-based system for online OS tuning that models cross-knob policy structure for long-running service optimization.
WARD: defense framework against prompt injection attacks on web agents, improving robustness across unseen domains and attack patterns.
Study of linguistic adaptation in multi-agent LLM systems under social observation, examining LLM behavior as communicative actors.
SpeakerLLM: audio-specialized LLM for speaker understanding, verification, and reasoning in audio-first agents and conversational robots.
TFGN: architectural method for continual pre-training of LLMs without replay buffers or task labels, addressing catastrophic forgetting at scale.
Survey of training algorithms for spiking neural networks with taxonomy and benchmarking framework.
Research identifying cultural anachronism in Vision-Language Models interpreting historical artifacts with temporally inappropriate concepts.