Empirical study of multi-agent LLM debate showing isolated self-correction outperforms unguided homogeneous debate across teams of 10 models on high-difficulty benchmarks.
Survey of generative AI use in qualitative research, discussing suitability for different research approaches (small-q positivist vs Big Q non-positivist) in software engineering.
StyleShield: demonstrates fragility of AIGC detectors through continuous controllable style transfer, exposing reliability paradox as language models improve.
Code World Model preparedness report: Meta's code generation and reasoning model assessed for frontier AI risks; found no additional catastrophic risks beyond current ecosystem.
Interpretable experiential learning model based on state history and global feedback for reinforcement learning in resource-constrained environments, evaluated on Atari.
E-MIA: black-box membership inference attack against RAG systems that infers whether documents are in retrieval corpus by analyzing LLM response interactions.
Certainty-aware RAG system enabling LLMs to express appropriate confidence and say 'I don't know' to improve user trust.
Ablation study isolating contributions of LLM, vision perception, and control components in human-robot interaction for object detection.
Evaluation of retrieval-augmented generation chatbots in realistic multi-turn information-seeking workflows with compliance constraints.
Physics-aligned rotary positional encoding method for wireless foundation models in channel state information tasks.
Medical audio QA dataset with diverse clinical scenarios for benchmarking language and audio reasoning models.
Visual analytics workbench for exploring weather and climate data using embedding-based representations and similarity search.
Method using perplexity differencing to reveal finetuning objectives in LLMs, enabling detection of intentionally introduced behaviors.
Framework assessing how ambiguity and uncertainty in real-world medical queries degrade LLM reasoning compared to simplified benchmarks.
Benchmark and investigation of multimodal LLM behavior for emotion recognition under modality conflict and missingness conditions.
Certified purity architecture converting governance enforcement into structural boundaries for cognitive workflow execution systems.
Multi-agent reinforcement learning approach for tactical deconfliction and separation assurance between heterogeneous unmanned aerial systems.
Unlearning method to suppress hallucinations in LLMs, specifically addressing supply-chain attacks from fictional package recommendations.
Framework for computing optimal policies and safety filters using value functions with temporal logic specifications.
Lightweight policy orchestration framework using ML to select cache eviction policies for nonstationary object caches in cloud services.
Pretraining method combining layer-aligned distillation with convergence-based early exit for efficient transformer inference.
Defense mechanism against malicious instruction injection attacks in retrieval-augmented generation and tool-integrated LLM agents.
Legal analysis of EU AI Act governance gaps for autonomous agents in critical infrastructure systems like traffic and power management.
Knowledge tracing approach for AI tutoring systems using LLM-powered dialogue to assess student performance and provide personalized support.
Component-aware self-speculative decoding accelerates hybrid language models with architectural heterogeneity without external drafters.
DUET framework enables efficient inference by decomposing tasks: capable model produces reasoning signal, lightweight model interprets for prediction.
Forager testbed for continual reinforcement learning with partial observability, addressing plasticity loss in non-stationary environments.
Approach using multi-perspective transformers with test-time training and products of experts on ARC-AGI-2 visual reasoning benchmark.
Analysis of contradictory productivity claims in AI-assisted coding: controlled studies show gains but RCTs and telemetry show slowdowns and longer review times.
Method to minimize unintended side effects (collateral damage) when using activation steering to control LLM behavior through internal representation intervention.
CNN-based multi-input multi-output model for efficient spatiotemporal prediction addressing RNN parallelization and error accumulation.
Position paper arguing LLM inference serving systems need mathematical optimization and algorithmic foundations beyond current heuristics like FIFO and LRU.
Pre-trained deep learning model for automated plant leaf disease classification to enable early detection in agriculture.
Chain of Evidence method adds pixel-level visual attribution to iterative retrieval-augmented generation for multi-hop question answering.
Framework for autonomous learning systems handling concept drift and non-stationary data streams beyond traditional temporal shifts.
GraphSculptor method for constructing efficient pre-training coresets in graph self-supervised learning by exploiting dataset redundancy.
Sequential experimental design framework for vision-language models to overcome perceptual bandwidth bottleneck via active vision.
MAD-OPD uses multi-agent debate to overcome single-teacher capability ceiling in on-policy distillation for agentic tasks.
Self-contrast decoding strategy for diffusion large language models to improve generation quality via high-information-density context modeling.
Empirical study of how developers use LLMs in software design tasks, surveying practitioners on benefits and drawbacks.
LiveFMBench: systematic study of LLM and agent-based formal specification generation for C programs with contamination awareness.
Verbal-R3 bridges retrieval and LLM reasoning via verbal annotations—analytic narratives connecting queries to retrieved contexts in RAG systems.
Framework for capturing expert cognition in practice-based domains to enhance AI-driven educational systems.
Medmarks: open-source benchmark suite with 30 benchmarks for evaluating LLMs on medical tasks including QA, information extraction, and calculations.
HepScript DSL enabling human-AI collaborative data analysis for high-energy physics workflows using agentic LLMs and domain-specific knowledge.
Theoretical analysis of generalization properties in multimodal metric learning with incomplete or redundant data.
Analysis of adversarial attacks on vision-language models, distinguishing between output perturbation and precise injection attacks.
Multi-agent autonomous testing system using LLM with LangGraph orchestration for UI test repair in enterprise applications.
FT-RAG: fine-grained retrieval-augmented generation framework for complex table reasoning using entry-level decomposition and semantic comprehension.
ProMORNA: multi-objective reinforcement learning framework for designing full-length therapeutic mRNA from protein sequences balancing stability and safety.