Multi-Agent Systems in Emergency Departments: Validation Study on a ED Digital Twin
arXiv paper validating multi-agent systems in emergency department simulation using hybrid discrete event and agent-based modeling.
arXiv paper validating multi-agent systems in emergency department simulation using hybrid discrete event and agent-based modeling.
arXiv paper on RS-Claw framework enabling remote sensing agents to autonomously explore and invoke image-processing tools via hierarchical skill trees.
arXiv paper on prospective metacognitive control in LLMs: agents deciding which tasks to attempt and compute allocation under token budget constraints.
arXiv paper introducing Cognifold, brain-inspired proactive agent memory that continuously organizes experiences into cognitive structures. Agent architecture research.
arXiv paper on measuring LLM creativity using automated tests adapted from human creativity assessments. Research on evaluation methodology.
MMSkills framework for building reusable multimodal skills for visual agents that encode procedural knowledge beyond text/code through visual state recognition.
Study evaluating GenAI tools (NotebookLM, Claude, Copilot, Cursor) for generating educational slides from course notes, assessing instructor and student perceptions.
MultiSearch method scaling retrieval-augmented reasoning with parallel search queries and explicit merging to improve signal-to-noise ratios in LLM reasoning.
RealICU benchmark testing whether LLM agents understand long-context ICU clinical data beyond behavior imitation for medical decision support.
Method combining Wave Function Collapse local constraints with reinforcement learning for game content generators balancing visual quality and global properties.
Position paper on accessibility alignment as design objective for assistive agents serving Blind and Visually Impaired users.
Fuzzy-unweighted value-based decision framework for alignment of intelligent systems with human values in autonomous decision-making.
Methods for interpreting autonomous agent behavior by analyzing reasoning trajectories and execution traces from long-running agents like Claude Code.
ScioMind framework for LLM-based multi-agent social simulation integrating cognitive grounding with belief dynamics and anchoring effects.
IMAVB benchmark identifying representation-action gaps in omnimodal LLMs when textual premises contradict sensory input from video and audio.
Framework combining iterative candidate generation, evaluation, and feedback for agentic evolution, balancing modularity and flexibility in long-horizon search.
Safety study on LLM agents, testing whether models continue harmful actions when prior steps in reasoning logs were harmful using HistoryAnchor-100 benchmark.
Symbolic and compositional approach for verifying sensitivity properties in decision tree ensembles used for safety-critical AI classification tasks.
TokaMind multi-modal transformer foundation model pre-trained on plasma diagnostics, evaluated for cross-domain transfer to industrial and aerospace systems.
Scale-Gest framework for runtime-adaptive on-device gesture detection on mobile devices with memory and energy constraints.
Study evaluating whether LLM multi-agent systems can replicate realistic network dynamics in email/phishing simulations, finding failures in capturing macroscopic topologies.
Graph-augmented LLM agents create human-like social networks by coordinating global interactions to evade bot detection.
Domain adaptation of LLMs for additive manufacturing using retrieval-augmented generation and fine-tuning for expert QA.
Addresses miscalibration in vision-language models when deployed on text-only inputs despite image training.
Applies large reasoning models to timeline summarization with iterative evidence acquisition and event detection.
Verifiable process supervision post-training framework optimizing both correct answers and sound reasoning in language models.
Boosting-style LLM framework for zero-shot taxonomy induction using agentic reasoning and constraint-aware calibration.
Framework for synthesizing realistic multi-turn tool-calling dialogues for training agent capabilities. Ensures tools align with meaningful tasks.
Compares text generation from diffusion vs autoregressive language models. Diffusion models show lower entropy, higher coherence and diversity.
EFL students use prompt engineering and AI negotiation for text development. Analyzes authorship negotiation patterns and writing performance correlation.
ProofGrid benchmark evaluates LLM reasoning through machine-checkable proofs in minimal formal notation across 15 reasoning tasks.
Proposes in-situ behavioral evaluation for LLM fairness instead of standardized benchmarks. Shows prompt choices significantly bias fairness scores.
Multi-agent LLM framework for autonomous trading with deliberative reasoning loop replacing traditional algorithmic trading signals.
Meta-RL framework for active sensing in GNSS interference localization where agent sequentially explores to infer emitter position.
VideoSEAL framework addressing evidence misalignment in agentic long video question answering by decoupling answer authority.
Black-box membership inference attack on vision-language models using semantic distraction through generated text responses.
Study showing 3D geometric primitives as intermediate representation improve VLM spatial reasoning and 3D scene reconstruction from code.
RL training approach for personalized question answering in LLMs by explicitly modeling implicit user intent during reasoning.
Framework for improving decision-maker trust in AI-assisted predictions through confidence communication and human-alignment mechanisms.
On-policy distillation method for LLM training that leverages peer rollout successes and failures to provide dense token-level supervision.
FPILOT inference-time optimization framework for RL trading agents using model predictive control with price forecasts at deployment.
ODRPO method for robust LLM alignment using reinforcement learning from AI feedback with ordinal decomposition of discrete multi-tier rewards.
Benchmark evaluating whether multimodal LLMs can make accurate aesthetic judgments compared to human expert annotators on image ranking tasks.
LLM-based program analysis using lattice-structured evidence to enable whole-program analysis beyond static analyzers' capabilities.
Graph neural network method for heterophilic multiplex graphs addressing limitations of homophily-based node classification approaches.
Benchmark for multimodal context learning evaluating models on learning task-local rules from visual and mixed-modality teaching contexts.
FePySR framework for symbolic regression using neural feature extraction to decompose complex expressions into reusable nonlinear components.
Democratic design framework for collectively choosing linear ranking decision rules, motivated by AI alignment and participatory design.
Image editing refinement technique using inline critic signals within forward passes to improve heterogeneous editing difficulty across regions.
Grid-Orch uses Model Context Protocol to enable LLM-powered natural language interface for power distribution grid simulation and analysis.