Who Benefits from RAG? The Role of Exposure, Utility and Attribution Bias
Study of fairness impacts in RAG-augmented LLMs, examining if certain demographic groups receive systematically different response quality.
Study of fairness impacts in RAG-augmented LLMs, examining if certain demographic groups receive systematically different response quality.
Multi-agent LLM workflow for automated penetration testing of networked cyber-physical systems and robotic infrastructure using environment grounding.
DVM runtime kernel generation system for efficient compilation of dynamic AI models with variable tensor shapes and control flows.
Probabilistic time series forecasting method embracing heteroscedasticity for uncertainty quantification.
Heterogeneous caching accelerates diffusion-based video editing by reusing features across denoising timesteps.
Studies coordination failures when multiple LLM-based code agents implement parts of same class without explicit specification.
Graph neural network layer (CSNA) for heterophilous graphs with cost-sensitive neighborhood aggregation.
Reviews neural motion planning approaches for robotic manipulators, discussing challenges in generalist manipulation policies.
Automated reward design framework using LLMs for cooperative multi-agent reinforcement learning with aligned incentives.
Coarse-to-fine visual processing reduces computational costs in document parsing with vision-language models.
GameplayQA benchmark for evaluating multimodal LLMs as perceptual backbones for autonomous agents in 3D environments.
Improves deepfake audio detection using neuron-level mechanisms and neuroplasticity. Builds on Wav2Vec and LLMs.
Studies emergent self-awareness in continual robot learning by quantifying invariant cognitive structures.
MolEvolve combines LLM guidance with evolutionary search for interpretable molecular optimization, addressing activity cliffs.
LLMs assess teacher-child interactions in Chinese preschools for scalable early childhood education monitoring.
Studies fairness in recommender systems, examining relationship between fair model representations and fair recommendations.
ClawKeeper adds safety mechanisms to OpenClaw autonomous agent runtime, addressing vulnerabilities in tool integration and command execution.
OneSearch-V2 improves generative retrieval for search systems with latent reasoning and self-distillation. Industrial-scale framework.
Large-scale annotated video demonstration dataset for computer-use agents enabling automation of complex desktop workflows with continuous video sequences.
Integration of causal machine learning into clinical decision support systems with clinician-facing interfaces for interpretable treatment-specific reasoning.
Autoresearch pipeline using Claude Code LLM agent to autonomously discover novel white-box adversarial attack algorithms outperforming 30+ existing methods.
Multi-dimensional evaluation framework for uncertainty attribution methods in explainable AI with aligned proxy tasks and metrics.
Mobile GUI agent using rejection fine-tuning to learn from failed trajectories and improve credit assignment for long-horizon tasks.
Video-language foundation model pretraining on surgical procedure videos for zero-shot event recognition in intraoperative settings.
Framework combining diffusion-based world models with selective enhancement for temporally coherent augmented reality applications.
Empirical study comparing chunking strategies for RAG systems in oil and gas documents, evaluating fixed-size, recursive, semantic, and structure-aware approaches.
Agentic video understanding framework using Vision-Language Models with active planning to seek evidence from raw video during reasoning.
Free-Market Algorithm metaheuristic using distributed supply-and-demand dynamics for open-ended optimization with emergent fitness.
Adversarial attack methods to protect images from malicious diffusion-based image-to-video generation models.
Vision-Language Models for converting rasterized figures into editable SVG vector graphics automatically.
Study of RAG systems applied to AI policy analysis using AGORA corpus, examining reliability challenges in dense legal language domains.
Vision-Language Models for supporting human decision-making in high-stakes domains like medical diagnosis through collaborative human-AI systems.
Study of AI agents powered by LLMs in multi-echelon supply chain simulation investigating emergent strategic behavior and dynamics like the bullwhip effect.
Graph-based evaluation framework for domain-specific LLM benchmarking using clinical guidelines transformed into queryable knowledge graphs with dynamic query instantiation.
GeoSketch: Neural-symbolic approach for geometric reasoning in MLLMs using auxiliary line construction and affine transformations for problem solving.
SAG-Agent: LLM-based agent using dynamic knowledge graphs for long-horizon reasoning in strategy games via GUI interaction without APIs.
CastMind: Agentic reasoning framework for time series forecasting using iterative refinement with temporal features, domain knowledge, and case-based references.
Generative Adversarial Reasoner: Framework using adversarial RL to improve LLM reasoning capabilities and reduce mathematical errors through co-evolved reasoner-discriminator training.
Research on enabling ultra-long-horizon autonomous agents with cognitive accumulation for multi-week ML engineering experiments.
Evaluates LLM performance on perspective-taking and knowledge state estimation tasks comparing cognitive abilities to chimpanzees.
CollectiveKV framework reduces inference latency in Transformer-based sequential recommendation systems through KV cache optimization.
CIRCLE framework for evaluating AI systems across six lifecycle stages, bridging gap between benchmarks and real-world deployment outcomes.
Framework for evaluating logical reasoning agents with agentified assessment, standardized interfaces, and structured failure tracking.
TikZilla: Dataset and reinforcement learning approach for scaling text-to-TikZ scientific figure generation from high-quality training data.
GPT4o-Receipt: Benchmark of 1,235 receipt images comparing AI-generated vs authentic documents evaluated by LLMs and humans.
Framework for relationship-aware safety unlearning in multimodal LLMs addressing relational safety failures without collateral damage.
DomAgent: Framework combining knowledge graphs and case-based reasoning with LLMs for domain-specific code generation tasks.
Counterfactual learning approach for CVR estimation in recommender systems addressing data sparsity and sample selection bias.
Survey on enterprise financial risk analysis using big data and LLM technologies for financial prediction and management.
Dynamic Neural Potential Field: Learning-enhanced MPC framework coupling Transformer-based predictor with classical optimization for robot obstacle avoidance.