StateFuse: Deterministic Conflict-Preserving Memory for Multi-Agent Systems
StateFuse conflict-aware memory layer for multi-agent systems preserving disagreement across branches and retries.
StateFuse conflict-aware memory layer for multi-agent systems preserving disagreement across branches and retries.
PCBWorld open-source benchmark environment for learning-based PCB routing agents using KiCad engine integration.
SearchEyes multimodal search agent framework using typed knowledge graphs and world simulation for multi-hop reasoning.
Integration approach combining knowledge graphs and multilingual corpora for domain-adaptive LLMs in social sciences and humanities.
Black-box evaluation framework assessing LLM ability to generate Design Structure Matrices from technical documentation.
AgoraSim hybrid framework combining LLM agents with traditional agent-based modeling for social scenario analysis.
Pre-registered experiment testing information-theoretic predictions on multi-agent economies with frontier LLM agents.
PolyWorkBench benchmark for evaluating LLM agents on long-horizon tasks requiring multilingual planning and tool use.
Heuristic approach for dynamic multi-vehicle routing optimization balancing reward maximization and computational efficiency.
Toy framework modeling curiosity as ecosystem with adjustable weights for uncertainty reduction, costs, and delayed returns.
Information gain-based adaptive tree-structured rollout optimization for multi-turn LLM agents on long-horizon tasks.
TOFFEE system for synthesizing high-quality data agent trajectories at scale for heterogeneous enterprise environments.
Theoretical framework for embedding application-layer cognitive protocols into native LLM meta-architecture via structural mechanisms.
Task decomposition-guided reranking method for improved skill selection in agent systems with large skill libraries.
LLM safety guardrail using intent-driven reasoning to balance efficiency and robustness against complex safety risks.
Post-hoc interpretability module using dictionary learning to decompose autonomous driving model behavior into semantic concepts.
Training-free framework using agentic topology sampling and knowledge graphs for zero-shot IoT forecasting in building sensor networks.
ExplAIner declarative query language for specifying and analyzing model explanation methods uniformly across classification models.
Multi-agent LLM system extracts H. pylori infection evidence from medical reports across heterogeneous structured and free-text fields.
Danus system orchestrates LLM-based mathematical reasoning agents using fact-graph memory for research-level problem solving.
Multi-agent deep reinforcement learning system for battery management optimization in dairy farm renewable energy integration.
Method to predict LLM agent failures early using probe cascades on hidden activations to avoid wasted compute in multi-step tasks.
Large-scale multivariate time series dataset for training foundation models on real-world data rather than synthetic datasets.
FootsiesGym: open-source benchmark environment for two-player zero-sum imperfect-information games with vectorized simulator for efficient training.
FreqDepthKV: frequency-guided cache compression factorizing KV states into shared low-frequency and sparse high-frequency components for long-context inference.
VAORA: reward design for vision-language models to improve physical reasoning and align internal reasoning with actions in interactive tasks.
DepthWeave-KV: token-adaptive KV cache compression across transformer layers using residual factorization for long-context inference.
Examines AI's impact on linguistic and cultural preservation in Indian subcontinent, viewing AI as double-edged for inclusion vs. homogenization.
PORTICO: reference monitor for revocable capabilities in coding agents, compiling task contracts into initial capabilities and closure predicates.
Runtime verification framework for AI agent actions: formalizes authorization, tamper-evidence, and deterministic reconstruction of agent trajectories.
Benchmark comparing KV-cache optimization techniques (quantization, pruning, merging) across models, tasks, and serving stacks for long-context LLM inference.
Analyzes ICLR papers 2017-2025 to identify trajectory-changing methodological contributions using public reviewer scores and decisions.
CCBENCH: benchmark evaluating LLM cultural competency by assessing adaptation to implicitly signaled cultural norms in health queries.
GAIDE framework enabling K-12 teachers to create AI-powered learning tech via vibe coding with LLMs, supporting teachers as designers.
Uses pattern-based knowledge components to automatically recommend programming learning resources, reducing need for expert curation.
CANONIC: system that applies compiler-like governance to LLM-generated content, using formal grammars to admit/reject artifacts into evidence ledgers at scale.
CHARLIE: on-premise multi-agent RAG system for structured evidential reasoning in digital forensics with traceability and compliance requirements.
Multimodal RAG system using post-hoc selective modality escalation to balance cost and utility by adaptively invoking vision-language models.
PORTS: preference-optimized retrieval method for tool selection in LLM agents. Aligns retrievers with tool-calling LLMs through joint training.
Multi-domain scientific code search benchmark with 5,264 curated repositories across scientific computing domains to evaluate code discovery tools.
Offline RL approach to learn control policies for LLM agent execution harnesses. Formalizes harness operation as finite-horizon MDP with frozen LLM executor.
BioSecBench-Refusal: benchmark for evaluating AI agent safety in biology tasks, pairing 61 routine tasks with 46 red-team scenarios for biosecurity assessment.
CanvasAgent: multi-modal agent orchestrating visual tools for complex image creation and editing through synthesis, segmentation, and composition.
KAT-Coder-V2.5: agentic coding model trained autonomously in executable repositories. Introduces AutoBuilder for sandbox reconstruction and end-to-end agentic post-training framework.
Comprehensive measurement study of mobile LLM inference across frameworks (llama.cpp, GENIE) and hardware backends (CPU, GPU, NPU). Introduces PowerBench profiling tool to identify efficiency bottlenecks.
Study of decision protocols in multi-agent LLM conversations for task performance improvement through specialized agent distribution and discussion mechanisms.
Analysis of confidentiality and legal privilege risks in generative AI systems across training, context windows, and RAG-based knowledge databases.
Calibration method for binary classifiers in adversarial environments maintaining consistent false-positive rates across predictions during model retraining.
PatchOptic system for managing shared state in agentic LLM workflows using projected views and verified structured updates instead of grep-like searches.
Analysis of naturally occurring statistical signals in vision datasets that behave like backdoor triggers, studying Imagenet patterns linked to labels.