SHAP-Weighted Cross-Modal Expert Fusion for Emotion and Sentiment Recognition: Evidence and Limits
XAI-guided adaptive fusion method combining unimodal and cross-modal experts for emotion and sentiment recognition.
XAI-guided adaptive fusion method combining unimodal and cross-modal experts for emotion and sentiment recognition.
Clinical-reasoning LLM for hepatocellular carcinoma risk stratification and treatment guidance from EMR narratives.
Analyzes real patient-chatbot conversations to understand communication patterns and develops patient simulator for evaluating health LLMs.
Multi-agent marketplace simulation studying formal mechanisms to prevent defection and maintain market stability among self-interested agents.
Physics-constrained benchmark for evaluating trustworthiness of autonomous agents in decentralized energy markets considering task performance and exploitation risks.
Addresses behavioral state decay in long-horizon AI agents through proactive memory mechanisms to surface decision-relevant information across expanding trajectories.
Analyzes quantization effects on LLMs beyond accuracy metrics, introducing correctness agreement to measure behavioral changes in quantized models.
Proposes symbolic workflow model for LLM-mediated applications using tool use, retrieval, branching, and checkpointing with Lisp-inspired conceptual framework.
Visual question answering benchmark for incident-centric dashcam understanding in autonomous driving using vision-language models.
Large-scale study of AI-based learning assistant usage patterns across 77,543 students in distance education.
Benchmark for evaluating LLMs on scientific lineage reasoning and idea generation grounded in paper citation inheritance structures.
Local linear transformer architecture for PDE operator learning with reduced computational complexity and local interaction bias.
Spectrum-aware framework for continual LLM fine-tuning using recursive consolidation of LoRA adapters across task sequences.
Multi-agent framework coordinating solver, critic, and aggregator agents for consensus reasoning with semantic and procedural evaluation.
Dynamic bifocal RoPE method for efficient zero-shot context extension in LLMs, supporting long-context agentic workflows and RAG.
Computational psychiatry study modeling psychological disorders in RL agents through controllable manipulation of cognitive appraisal signals.
Analysis of deep reinforcement learning evaluation paradigms and design principles from foundational DQN to recent algorithmic advances.
LLM-driven formal mathematics system for frontier research using interactive theorem proving, advancing beyond well-defined problems.
Multi-agent LLM framework for dynamic emotional evolution in persona-based dialogue using cognitive appraisal models.
Multi-agent LLM system for autonomous formalization of tensor network theory research, coordinated through structured blueprints with periodic human review.
Data-efficient vision-based obstacle avoidance for robots using pretrained vision models, avoiding sim-to-real transfer problems.
Research using internal attribution graphs to understand LLM jailbreak mechanisms and how adversarial perturbations alter internal reasoning.
Survey of multimodal unlearning methods, datasets, and benchmarks for VLMs, DMs, LLMs addressing sensitive/biased training data removal.
Latent Personality Alignment method for efficient safety alignment of LLMs using adversarial training on 66 harm-agnostic statements.
Research on adversarial decoys attacking Vision Transformers by misdirecting attention-based defenses. Security and robustness study.
path_boost Python package for interpretable graph-level prediction using path-based gradient boosting algorithm. Open source tool.
Comparative study of linear attention architectures (DeltaNet, Gated DeltaNet, Kimi Delta) addressing quadratic cost limitations of softmax attention.
Research on out-of-scope intent detection using multi-cluster boundary learning with MiniLM embeddings for human-machine interaction systems.
Tail-aware credit calibration method for LLM reinforcement learning that addresses positive-credit contamination in uniform token advantage assignment.
Qualitative analysis of 3100 practitioner opinions on code review in context of AI coding agents using causal inference.
Empirical reliability assessment of Gemini models as audio judges for full-duplex voice agents, validated against human raters.
Failure localization framework for diagnosing which agent caused system-level failures in LLM-based multi-agent systems.
CodeTracer forensic framework for detecting and attributing backdoor attacks in code completion models.
Provably efficient learning algorithms for repeated assistance games where informed and uninformed agents optimize shared rewards.
Graph-based framework to quantify uncertainty and coherence in LLM reasoning chains beyond final-answer agreement.
Vision-language model planner for long-horizon robot tasks that interleaves semantic reasoning and geometric constraint checking adaptively.
Structured pruning method for LLMs using power transformation and sign-preserving score aggregation with adaptive feature retention.
Deep learning method for automatic modulation classification with domain adaptation using knowledge and data-driven approaches.
PLURAL dataset with 92 countries of value-aligned preference data from Integrated Values Survey to reduce Western bias in LLMs.
AI agent for research software engineering that maintains alignment across distributed artifacts like meetings, pull requests, and GitHub issues.
Probing internal LLM representations to improve calibration and faithfulness of forecasting models beyond chain-of-thought outputs.
Framework for self-validating LLM-assisted safety analysis using Constitutional AI and Systems-Theoretic Process Analysis to detect hallucinations.
Investigation of adaptive token generation ordering in diffusion language models for text-to-image synthesis and multimodal understanding tasks.
Survey on KV cache optimization techniques for efficient LLM serving, focusing on system-aware infrastructure for reducing memory and latency.
Empirical analysis of uncertainty quantification in vision language models with chain-of-thought reasoning, testing four models on adversarial samples.
ICDAR competition on LLM-assisted OCR post-correction for historical documents. Evaluates LLM effectiveness across languages and document types.
Security framework preventing prompt injection attacks on web-based autonomous agents. Confines cross-site injection by separating task from untrusted content.
Defense against LLM-based agentic crawlers exploiting context compression. Revisits agent pipelines to identify new threat surface.
Vision-language-action model for robotics emphasizing task-critical visual evidence. Handles complex dynamic scenarios with latent environment modeling.
Open-ended RL curriculum using multi-modal LLMs to visually inspect agent policies and assess task difficulty. Enables complex skill development.