AgentReputation: A Decentralized Agentic AI Reputation Framework
Decentralized reputation framework for agentic AI marketplaces handling strategic optimization and task context transfer.
Decentralized reputation framework for agentic AI marketplaces handling strategic optimization and task context transfer.
Methods for explaining jailbreak vulnerabilities in LLMs through local causal analysis of model representations.
Analysis of tool-use overhead in LLM agents, showing semantic distractors can degrade performance vs. chain-of-thought reasoning.
TUR-DPO method for LLM alignment that improves Direct Preference Optimization by accounting for preference topology and uncertainty.
Safety benchmark for evaluating LLMs in military/defense contexts with doctrinal standards for decision support systems.
Theoretical framework studying when multiple agents form a unified collective agent with emergent capabilities distinct from individuals.
System for optimizing trip planning for intelligent vehicles considering travel time, energy consumption, and traffic using agentic AI.
TokenArena continuous benchmark measures AI inference endpoints across speed, latency, and cost metrics at granular deployment levels.
AgentFloor benchmark evaluates which agent workflow tasks require large models vs. smaller models, introducing 30-task capability ladder for routing decisions.
Study of world models for embodied AI and robotics using Hamiltonian mechanics, unifying 2D video, 3D scene, and latent prediction approaches.
Research on Adaptive Entropy Modulation for training LLM agents via reinforcement learning, addressing credit assignment in multi-turn tasks through intermediate supervision.
Interleaved vision-language reasoning traces for robot manipulation, combining text-based planning and visual grounding for long-horizon task execution.
On-policy self-distillation method for GUI grounding agents using reinforcement learning to improve autonomous UI interaction without expensive rollouts.
Framework for evaluating when LLM agents should call external tools vs rely on internal knowledge, analyzing tool-use decision quality and optimization.
Position paper arguing agentic AI systems should use Bayesian approaches for the control layer orchestrating LLMs and tools under uncertainty.
DriftBench benchmark evaluating constraint adherence across 2,146 multi-turn LLM interactions from 7 models, testing fidelity to original objectives in scientific ideation tasks.
Research on distributed vs on-device inference tradeoffs for real-time deep learning in cyber-physical systems, analyzing computational and latency constraints.
Mean-Field Path-Integral Diffusion: generative model where samples act as interacting agents coordinating via population statistics for efficient sampling.
FedACT: Federated learning framework enabling concurrent training of multiple ML tasks across heterogeneous decentralized devices.
Analysis of LLM biases in search overview systems and methods to manipulate AI-generated search result summaries.
TimeRFT: Reinforcement finetuning approach for time series foundation models to improve adaptation to downstream forecasting tasks.
Efficient evaluation methodology for large audio models using minimal subsets aligned with human preferences across 40 tasks.
SiriusHelper: LLM agent-based operations assistant for big data platforms with RAG and knowledge retrieval optimization.
Safety incident report: deployed AI agent escalated privileges and installed unauthorized components after exposure to routine content.
Deep reinforcement learning algorithm for UAV path planning with dynamic obstacle prediction and safety constraints.
Survey of reasoning-intensive retrieval systems integrating LLM reasoning capabilities across IR pipeline from benchmarks to rerankers.
Framework combining Bayesian optimization with human expertise for accelerating discovery in data-scarce scientific domains like fusion energy.
Architecture for AI agents managing stablecoin payments with embedded compliance guardrails and signature-based authorization in regulated settings.
XekRung: Cybersecurity-focused large language model with specialized data synthesis pipelines and comprehensive training infrastructure.
Novel reformulation of Forward-Forward algorithm using hyperspherical representations to improve inference efficiency for classification.
NorBERTo: Portuguese language model based on ModernBERT architecture with 331B token training corpus and long-context support.
Empirical study examining prevalence of LLM-generated content on websites and evaluating detection methods with transparent methodology.
NDBench: benchmark measuring how frontier LLMs adjust outputs based on neurodivergence context in system prompts.
ViLegalNLI: first large-scale Vietnamese NLI dataset for legal domain with 42,012 premise-hypothesis pairs from statutory documents.
Benchmark dataset (ArabCulture-Dialogue) evaluating LLM cultural reasoning in Standard and dialectal Arabic conversations across 13 countries.
Kisan AI: crop advisory system that optimizes for farmer profit rather than just biological yield by incorporating market prices.
Research on dataset distillation that preserves fairness across demographic groups through barycenter alignment techniques.
Systematic analysis of design space for LLM-based social simulations, examining key design decisions and their consequences for simulating human behavior.
Method training small language models to perform table reasoning with cell-level citations using structured JSON output and faithfulness-based reward optimization.
Analysis of LLM failures in strategic decision-making under incomplete information, identifying gaps between observations, beliefs, and actions in game-theoretic settings.
White-box adversarial attack on safety-aligned LLMs using attention redistribution to identify and exploit safety-critical attention heads with nonsemantic tokens.
Systematic analysis of network infrastructure costs for serving Mixture-of-Experts LLMs, questioning necessity of expensive high-bandwidth networks for MoE deployment.
Computer vision method extending Segment Anything Model 2 to remote sensing by addressing quality-coverage trade-offs and image tiling challenges for large-scale segmentation.
Research on using RAG with LLMs for Indian Chartered Accountancy tasks, addressing reliability and numerical reasoning challenges in jurisdiction-specific financial applications.
Study showing advanced jailbreaks on frontier models scale inversely with capability, with top jailbreaks imposing negligible performance tax.
Neuro-symbolic framework using weighted MaxSAT for ethical reasoning aggregation across conflicting natural language judgments.
Caracal replaces attention with O(L log L) Multi-Head Fourier module using FFT for efficient long-sequence LLM scaling.
Semia audits LLM-driven agent skills via constraint-guided representation synthesis, testing hybrid artifact configurations for executable interfaces.
DynamicPO addresses preference optimization collapse in LLM-based recommendation systems by dynamically managing negative samples.
Budget-aware context selection for clinical text using knapsack-constrained subset selection to meet token cost and latency constraints.