Describing Agentic AI Systems with C4: Lessons from Industry Projects
Framework and architectural patterns for describing agentic AI systems with multi-agent collaboration and artifact exchange.
Framework and architectural patterns for describing agentic AI systems with multi-agent collaboration and artifact exchange.
Replication study on AI-generated text detection using multilingual models and SHAP-based explainability analysis.
Vision-Language-Action model using state space models for efficient language-guided robotic manipulation tasks.
Latent-space reasoning approach for LLMs that reduces inference cost by computing reasoning steps implicitly rather than generating verbose traces.
Analysis of error sources in global feature effect estimation methods like PD and ALE plots for model interpretation.
Open-source biomedical knowledge graphs with federation and AI agent access for cross-referencing siloed databases.
Research on improving LLMs' code generation using private libraries through better knowledge integration beyond API documentation retrieval.
HindSight evaluation framework measuring AI-generated research idea quality by matching against future publications and citation impact.
Hybrid approach combining Iterative Learning Control with deep reinforcement learning for safe and convergent batch process control.
Token Coherence framework applying MESI cache protocols to reduce synchronization overhead in multi-agent LLM orchestration systems.
CATFormer combines continual learning with spiking transformers using dynamic thresholds to mitigate catastrophic forgetting.
In-context symbolic regression for extracting interpretable analytical expressions from Kolmogorov-Arnold Networks in scientific ML.
RESTA defense extended to vision-language models for robustness against multi-modal jailbreaking attacks in trustworthy agentic AI.
Code-centric learning approach for LLM-based ICD medical coding improving generalization to unseen codes with better interpretability.
Scalable simulation-based model inference framework with test-time complexity control for selecting among large families of forward models.
CCTU benchmark evaluating LLM tool use under complex constraints, testing function calling, instruction following, and self-refinement.
SKILLS benchmark framework evaluating LLM agents on 37 telecom operations workflows with real API interfaces, testing structured knowledge injection.
Hybrid XAI framework combining counterfactual explanations and feature attribution for neural network interpretability in healthcare/finance.
Analysis of beam search in LLMs showing wider beams can hurt output quality due to overestimation bias, grounded in Extreme Value Theory.
Spatial reasoning agent decoupling perception from reasoning in visual language models for improved metric and geometric scene understanding.
Bi-level optimization approach using Stackelberg game theory for coupled morphology-control co-design in embodied agents.
Adversarial patch framework for evasion and impersonation attacks against facial re-identification systems across non-overlapping cameras.
Safety defense mechanism for LLMs monitoring intermediate reasoning steps in chain-of-thought to prevent jailbreak attacks.
Benchmark evaluating marginal utility of agent skills for LLM-based software engineering agents on real GitHub issues and requirements.
Empirical study of 16 LLMs examining internal mechanisms for table understanding across attention dynamics, layer depth, and expert activation.
Comprehensive safety evaluation and monitoring framework for LLM-based multi-agent systems addressing novel risks beyond single agents.
Framework for improving robustness of quantized DNNs through three-stage fine-tuning addressing both fault and attack resilience.
Analysis of safety vulnerabilities in test-time training methods for LLMs, examining susceptibility to prompt injection and adversarial attacks.
Vision-language critic model leveraging pre-trained VLAs for multi-agent reinforcement learning value estimation with improved generalization.
Memory management framework for small language model agents using adaptive clustering to organize experiences and prevent knowledge corruption.
Fine-tuning strategies for PDE foundation models using physics-informed training to adapt to new tasks with limited domain-specific data.
Framework evaluating AI agent vulnerabilities by applying malware analysis concepts to test-time agent behavior and adversarial robustness.
Knowledge distillation method for tabular models that addresses feature interactions without original training data, enabling privacy-preserving model compression.
Multi-agentic workflow deploying AI agents with automated instruments to recover critical materials via selective precipitation.
Analyzes grokking phenomenon in neural networks through spectral gating mechanism and optimizer noise interaction.
Argues for taxonomy-specific evaluation in time-series forecasting to accurately assess ML progress versus classical methods.
SlovKE: Dataset and LLM evaluation for keyphrase extraction in Slovak, a morphologically rich low-resource language.
InterveneBench: Benchmark evaluating LLMs on causal inference and intervention reasoning in realistic social science scenarios.
Studies how LLMs model student misconceptions when generating multiple-choice distractors, analyzing reasoning strategies.
PokeAgent Challenge: Large-scale benchmark for competitive multi-agent decision-making with partial observability and long-horizon planning.
Lore: Protocol using structured Git commit messages to preserve decision context and institutional knowledge for AI coding agents.
PRIMO R1: Framework using reinforcement learning to improve multimodal models for process reasoning in robotic manipulation.
Analyzes moral indifference in LLMs due to compressed moral concepts and proposes remedial techniques.
Mixture-of-Depths Attention: Mechanism addressing signal degradation in deep LLMs by enabling attention to multiple depth levels.
Combines tree-search, generative models, and Nash bargaining for opponent modeling in game-theoretic reinforcement learning.
FAIRGAME: Framework using game theory to detect and recognize bias in multi-agent AI systems.
Method to reduce reasoning path length in large reasoning models like o1 and R1 using reward designs in reinforcement learning.
AssetOpsBench: A benchmark framework for evaluating LLM agents on industrial asset operations tasks like condition monitoring and maintenance scheduling.
Framework for AI alignment grounded in resource-rational contractualism, enabling diverse stakeholders to reach agreements on AI decision-making.
Machine learning approach to automate story point estimation for software sprint planning using comparative learning from historical team decisions.