Agent-First Tool API: A Semantic Interface Paradigm for Enterprise AI Agent Systems
Agent-First Tool API paradigm redesigns APIs for autonomous agents, addressing architectural mismatches with human-oriented CRUD APIs.
Agent-First Tool API paradigm redesigns APIs for autonomous agents, addressing architectural mismatches with human-oriented CRUD APIs.
Examines lack of explicit reasoning and interpretability in deep learning models and proposals for understanding representations.
SciAidanBench measures scientific creativity of LLMs through open-ended questions, examining uneven capability progress.
LLARS open-source platform enabling collaboration between domain experts and developers for LLM prompt engineering and evaluation.
Budget-efficient automatic algorithm design using LLMs via code graph representation to reuse algorithmic components.
Argues for calibrated verification over mechanistic interpretability for AI deployment authorization in sensitive domains.
PRISM detection system for preventing secret leakage propagation through shared context in multi-agent LLM pipelines.
Hierarchical causal framework for explainable model predictive control in safety-critical infrastructure.
Evolutionary framework using learned optimization policies as teachers to generate heuristic programs for combinatorial optimization.
Investigation of bias and robustness issues in LLM toxicity evaluation benchmarks used for deployment certification.
Framework for self-evolving agents that distill experience from interactions to adapt to novel tasks at deployment time.
Theoretical framework examining agent cybernetics as foundational science for long-horizon LLM agents using tool loops and reflection.
MATRA threat modeling framework for assessing security risks in LLM-based agent systems with tool access.
Multi-task benchmark for aligning language descriptions with urban trajectory data, focusing on mobility modeling and planning.
ComplexMCP benchmark evaluates LLM agents on interdependent tool use in realistic scenarios with 300+ tests on Model Context Protocol.
Knowledge graph QA method using LLMs with retrieval-augmented generation, proposing informative path supervision for training.
Compares reasoning vs non-reasoning LLM judges showing structured verification benefits vary by task, proposes cost-efficient routing for LLM-as-a-Judge applications.
NanoResearch co-evolves skills, memory, and policy for personalized multi-agent research automation, adapting outputs to individual researcher resource constraints and preferences.
Investigates internal mechanisms and cross-modal information processing in audio-visual LLMs through interpretability probing techniques. Limited practical application focus.
MaD Physics benchmark evaluates agents for scientific discovery under physical measurement constraints, balancing quality and quantity with cost-benefit trade-offs.
Studies nonlinear impact of misleading information on LLM long-context performance, quantifying how distractors degrade retrieval-augmented generation and agentic systems.
Evaluates AI pentesting agents on real-world security targets beyond predefined CTF benchmarks, assessing offensive capabilities in unconstrained environments.
Generalized Turing Test framework for comparing arbitrary agent intelligence through indistinguishability via Turing comparators, dataset and task-agnostic.
BenchCAD comprehensive benchmark for programmatic CAD code generation from visual/textual inputs, requiring understanding of 3D structure, parameters, and engineering operations.
Rate-distortion framework for agent memory optimization preserving decision-relevant distinctions under limited runtime budgets rather than descriptive relevance criteria.
Shepherd runtime substrate formalizes meta-agent operations with typed execution traces and Git-like replay, achieving 5x faster process forking than Docker and 95%+ prompt-cache reuse.
Evaluates open-source LLMs for algorithm generation and ensemble methods for conjecture verification in number theory, demonstrating specialized domain application.
Pair-GRPO unified framework addresses instability in LLM RLHF alignment through preference-based RL optimization with improved gradient interpretability and variance reduction.
Delulu benchmark for detecting code hallucinations in LLM fill-in-the-middle tasks across 7 languages with 1,951 verified samples and 4 hallucination types.
Research evaluating AI companion chatbots for simulated relationships, analyzing emotional dependence risks and psychological harm. No direct developer or technical focus.
MedThink uses knowledge distillation and teacher-guided reasoning correction to compress LLM diagnostic capabilities into smaller models for resource-constrained clinical settings.
BaLoRA extends LoRA fine-tuning with Bayesian uncertainty quantification for large pre-trained models.
Transformer-based causal discovery method for non-stationary time series data with contemporary and lagged relationships.
Research showing product context improves AI coding agent decision compliance by 49%, with benchmark across 8 software engineering tasks.
Safety-Aware Denoiser framework for controlling safety risks in text diffusion models with guidance mechanisms.
Empirical analysis of feature repulsion and spectral properties during two-layer neural network grokking.
Foundation model approach for gene regulatory network inference from single-cell transcriptomic data.
Retrieval-augmented vision-language-action model for autonomous driving addressing long-tail scenario generalization.
Activation reuse technique exploiting token-wise redundancy in diffusion language models for efficient inference.
Empirical study showing weight pruning amplifies bias in compressed LLMs across three models and pruning methods.
Hopfield network-based approach for sequential model editing of LLMs preserving factual knowledge during updates.
Meta-learning framework assigning importance scores to noise instances in diffusion model training.
Method for improving vision-language model robustness by analyzing multimodal redundancy and synergistic information.
First unified benchmark for visual-tabular multimodal learning across 14 datasets in 9 domains.
Speculative decoding serving framework for resource-efficient multi-model LLM inference using tail models as drafters.
Federated learning framework integrating zero-knowledge proofs for privacy-preserving distributed AI model training.
Hierarchical video-language model reducing token growth and improving motion perception for long video understanding.
Parallel implementation of Hierarchical Self-Organizing Maps for intrusion detection systems in cybersecurity.
Methods for reducing inference latency in Vision-Language-Action models through inpainting, delay simulation, and residual correction techniques.
FFT-based neural network layers using circulant spectral machinery for improved Hessian conditioning with fewer parameters.