MEDS dataset maps mathematical reasoning across 14 LLMs with 28,000 personas. Evaluates LLM performance on math tasks under human and AI-like conditions.
WaferSAGE uses vision-language models for semiconductor defect analysis. Addresses data scarcity through synthetic data generation and rubric-guided reinforcement learning.
Demonstrates political bias audits of LLMs partly measure sycophancy rather than genuine bias. Models adapt responses to inferred auditor identity.
Shows LLM evaluation results differ significantly when prompts are optimized per model versus using static templates. Prompt optimization affects benchmark rankings and model selection decisions.
Investigates how LLMs learn from context by extracting rules into natural-language skills. Addresses challenges in inference-time skill augmentation for reasoning over complex contexts.
TEA Nets framework extracts agents, events, and targets from text using cognitive network science. Open-source Python library with applications in emotion detection and semantic analysis.
Multi-agent systems research drawing parallels between LLM-based agents and institutional governance, addressing coordination problems in complex societies.
ValuePlanner: hierarchical cognitive architecture for embodied agents decoupling value scheduling from action execution using LLM-based modules.
Conceptual analysis arguing that current agentic memory systems implement lookup rather than true memory, with implications for agent capability and learning.
Agentic framework constructing knowledge graphs from AI policy documents to answer policy compliance questions using LLM-based reasoning.
Audit of frontier vision-language models on medical VQA tasks, identifying grounding failures, format collapse, and domain adaptation issues.
MED-VRAG: multimodal RAG framework for medical question answering using document page images and visual content instead of OCR'd text.
Framework for traffic signal optimization using digital twins and agentic AI for real-time autonomous decision-making in transportation networks.
Intent2Tx: benchmark with 31,496 instances for evaluating LLMs on translating natural language intents into functionally correct Ethereum transactions.
WindowsWorld: benchmark for autonomous GUI agents testing cross-application workflows in professional settings, extending beyond single-application tasks.
PARA: data-free compression method for LoRA that adaptively allocates ranks across model layers to reduce parameter redundancy in foundation model fine-tuning.
Focus session on design challenges for safety-critical autonomous systems with embedded AI components, covering safety, security, reliability, and certification requirements.
MCPHunt benchmark evaluating credential propagation risks in multi-server MCP agents, isolating non-adversarial data leakage from workflow topology.
ObjectGraph file format redesigned for agent retrieval instead of linear human reading, reducing token waste and enabling structured data access for multi-agent workflows.
Grid-aware agent-based model for simulating EV charging systems incorporating heterogeneous behavior, infrastructure constraints, and power allocation dynamics.
Analysis of paradigm shift toward agentic reinforcement learning with LLMs, moving beyond traditional RL for autonomous agents in complex open-ended tasks.
KellyBench environment for evaluating long-horizon sequential decision-making in non-stationary sports betting markets with open-ended optimization goals.
LLM agent framework for clinical settings modeling concern trajectories with explicit state dynamics to surface pre-escalation signals without delegating medical authority.
Framework for dynamically building personalized multi-agent systems that adapt agent roles and coordination patterns to individual user needs and contexts.
Empirical comparison showing in-context prompting with system prompts outperforms agent orchestration frameworks for procedural tasks, enabling LLM self-orchestration.
Survey of graph-based world models for agents, decomposing environments into entity nodes to improve noise sensitivity, error accumulation, and reasoning.
HealthFormer transformer model generating human physiological trajectories for clinical intervention simulation using multi-visit cohort data.
Schema-aware memory system for AI agents enabling reliable persistent storage with exact facts, state tracking, updates, and structured queries beyond text retrieval.
Multi-agent framework for multimodal stance detection combining retrieval augmentation with multi-agent reasoning to handle text-image fusion and conflicting signals.
Theoretical framework unifying Bayesian inference, game theory, and thermodynamics for multi-agent collective intelligence without central coordination.
Empirical study examining how visual priming influences vision-language models' cooperative behavior using prisoner's dilemma game scenarios.
Comprehensive overview of GUI agents enhanced with reinforcement learning for long-horizon automation tasks, handling credit assignment and safe exploration in interactive environments.
Framework enabling LLMs to generate Answer Set Programs with self-correction for nonmonotonic reasoning, addressing logical inconsistencies and computational costs.
LLM agents optimize mechanical linkage designs using symbolic representations and modular optimization, combining discrete topology exploration with continuous parameter fitting.
D3-Gym: automatically constructed dataset with 565 verifiable environments for scientific data-driven discovery using LLM agents.
Comparative study of three LLM agent paradigms on scientific visualization tasks, evaluating domain-specific, computer-use, and coding agents.
Architectural pattern language for integrating vision language action models into enterprise systems, balancing latency and determinism requirements.
Benchmark dataset for evaluating multimodal LLMs on scientific spectral image understanding with expert-annotated QA pairs.
Methodology for systematically engineering LLM agents in scientific domains using three-party collaboration between SMEs, developers, and helper agents.
Research on evaluating Text-to-SQL agents in production without ground-truth queries or schema access, addressing real-world deployment gaps.
RHyVE: Verification and phase-aware deployment framework for LLM-generated reward hypotheses in reinforcement learning.
Characterizes consistency of emergent misalignment personas when LLMs fine-tuned on narrowly misaligned data generalize broadly.
Guidelines for designing terminal-agent benchmark tasks emphasizing adversarial review and verification logic for LLM coding evaluation.
Methodological framework mapping classroom interaction research across scale, duration, and modality dimensions in AI era.
Critical analysis of AI sign language translation tools from ableist and degrowth perspectives.
Research infrastructure for AI scientists showing methodological evolution as graphs, enabling AI research agents to discover connections between methods.
LLM-based graph structure refinement for improving EEG seizure diagnosis through better edge filtering in noisy signal data.
Scalable methodology for creating synthetic computer environments with realistic folder hierarchies and artifacts for productivity simulation.
Agentic Compilation addresses the Rerun Crisis in LLM-driven web agents, reducing inference costs and latency for repetitive tasks.
Analysis of self-consistency and reasoning effort strategies for optimizing cost and accuracy of LLM-based automated scoring.