CompassLLM: A Multi-Agent Approach toward Geo-Spatial Reasoning for Popular Path Query
CompassLLM multi-agent approach using LLMs for geo-spatial reasoning on popular path queries from trajectory data without requiring model retraining.
CompassLLM multi-agent approach using LLMs for geo-spatial reasoning on popular path queries from trajectory data without requiring model retraining.
Memory-as-Action (MemAct) framework treating working memory management as learnable policy actions for long-horizon AI agent tasks to mitigate attention dilution.
Analysis showing self-consistency decoding technique yields diminishing returns on modern LLMs like Gemini 2.5, often degrading performance on already-solved problems.
SpatialBench benchmark for evaluating multimodal LLMs on spatial cognition tasks, capturing hierarchical structure of spatial abilities.
ProAgent framework for proactive LLM agents that continuously perceive and assist users in real-world settings using sensory contexts and contextual awareness.
CORE reinforcement learning framework providing fine-grained conceptual signals to teach LLMs genuine mathematical concept application rather than pattern reuse.
SANet framework using specialized AI agents for autonomous decision-making, dynamic adaptation, and cross-layer optimization in 6G networks.
Owen-Shapley policy optimization algorithm for RL-trained LLMs with token-level credit assignment, addressing reward sparsity in recommendation tasks.
E-mem multi-agent episodic context reconstruction framework preserving logical integrity and sequential dependencies for LLM agent memory systems.
BioAgent Bench evaluation suite with curated end-to-end bioinformatics tasks for measuring AI agent performance and robustness with automated assessment.
Process-verifiable thinking data synthesis framework enabling LLMs to perform long Chain-of-Thought reasoning for diverse time series tasks.
Safety harness using capability-safe Scala 3 language to restrict agent tool calls and prevent information leakage, unintended side effects, and prompt injection attacks.
PURE framework addressing preference-inconsistent explanations in LLM-based recommenders through preference-aware reasoning.
Context specification methodology for making AI evaluations operational and relevant to organizational deployment success and decision-making.
MineEvolve framework enabling embodied Minecraft agents to accumulate knowledge through interaction and self-evolution for long-horizon task completion.
Neuro-symbolic framework integrating LLMs into interactive theorem proving to automate proof script generation for systems software verification.
Metacognitive co-regulation framework for LLM design agents to overcome fixation bias and explore alternative solutions in engineering design tasks.
Claw-Eval benchmark suite with 300 human-verified tasks across 9 categories for trustworthy evaluation of autonomous LLM agents in real-world software environments.
Framework for routing NP-hard optimization problems to diverse solvers via polynomial-time reductions, integrated with agentic orchestration.
Autogenesis Protocol for self-evolving LLM-based agent systems, addressing lifecycle management, version tracking, and safe updates to enable modular agent composition.
Critique of current LLM evaluation frameworks identifying systematic failures (distributional, temporal, scope, process) inadequate for deployed agentic systems and RLHF reward hacking.
Deep learning framework for learning lifted action models from visual state sequences without action observation, applicable to AI planning in real-world domains.
FutureWorld live RL environment for training predictive agents on real-world event forecasting with actual outcome rewards.
WaferSAGE uses vision-language models with synthetic data generation for semiconductor wafer defect analysis.
Adaptive Entropy Modulation improves credit assignment in multi-turn LLM agent reinforcement learning tasks.
Position paper arguing agentic AI control layers should be Bayes-consistent for tool/expert selection under uncertainty.
LLM evolutionary search determines Zarankiewicz numbers in discrete mathematics using reinforced optimization.
GR-Ben benchmark evaluates process reward models across diverse reasoning tasks beyond mathematics.
Segment-Aligned Policy Optimization improves credit assignment in multi-modal reasoning RL for language models.
Zero-shot confidence estimation for small LLMs enabling cost-effective local-to-cloud routing without supervised training.
Systematic evaluation of prompting and execution methods for deterministic computation in large language models.
Circuit analysis of LLM agent memory systems revealing internal mechanisms of information management across sessions.
Executor-grounded reward training improves faithful reasoning in LLMs beyond final-answer correctness.
Experience-RAG Skill enables AI agents to orchestrate different retrieval strategies based on task context.
Four pre-trained language models for low-resource Angolan languages using transfer learning and synthetic data.
DeTrigger proposes gradient-centric method to mitigate backdoor attacks in federated learning systems.
CatNet algorithm controls false discovery rate in LSTM feature selection using SHAP importance and Gaussian mirrors with kernel-based independence measures.
LicenseGPT fine-tuned foundation model for dataset license compliance interpretation, addressing legal risks in commercial AI product development.
Evaluates LLM robustness on code understanding against semantics-preserving mutations, assessing reasoning quality beyond accuracy on programming tasks.
Amortized linear-time exact Shapley value computation for product-kernel methods, enabling scalable explainability in kernel-based models.
Practical adversarial attacks on stochastic bandits via fake data injection with bounded perturbations, more realistic than prior per-round manipulation assumptions.
Provably safe RL using analytic gradients for autonomous robots, integrating safety safeguards during training to reduce sim-to-real deployment gaps.
Discourse-aware hierarchical retrieval for long document QA using rhetorical structure theory, improving over flat chunking approaches for LLM comprehension.
Survey of federated foundation models for privacy-preserving recommendation systems, addressing integration of FMs in decentralized collaborative settings.
Attribution-guided pruning discovers circuits in small LLMs for mechanistic interpretability, enabling diagnosis and targeted correction of undesirable behaviors.
Multi-objective instruction-aware RL for procedural content generation leveraging natural language control, addressing limitations in handling complex textual instructions.
Improves Equilibrium Propagation learning framework with feedback regulation and residual connections for brain-inspired computing, addressing instability and computational costs.
Study of multi-step reasoning in LLMs using cellular automata framework, examining how recurrence, memory, and test-time compute scaling improve reasoning depth beyond memorization.
RFA reformulates self-attention as robust state estimation using linear SDEs, matching computational complexity of standard attention while improving theoretical foundations.
Frictional Q-Learning reduces extrapolation errors in off-policy RL by treating replay buffer as low-dimensional manifold, using friction analogy for action selection.