DataEvolver: Let Your Data Build and Improve Itself via Goal-Driven Loop Agents
DataEvolver: closed-loop visual data generation system using goal-driven agents for iterative dataset creation and improvement.
DataEvolver: closed-loop visual data generation system using goal-driven agents for iterative dataset creation and improvement.
Decision-propagation method integrating Answer Set Programming with neural networks for scalable neuro-symbolic reasoning.
NeuroState-Bench: human-calibrated benchmark evaluating commitment integrity in multi-turn LLM agent tasks via side-query probes.
Categorical sheaf-theoretic framework for resilient multi-agent autonomous systems planning in stochastic environments.
AI-driven cybersecurity system for financial institutions using LLMs to enhance SOC reasoning capacity and alert investigation.
Adversarial self-play defense against persona-based jailbreak attacks in LLMs through intent-role disentanglement.
Formal language specification for encoding LLM agent context composition, standardizing context engineering practices.
Hierarchical reinforcement learning with language for pair trading, addressing credit assignment in long-horizon semantic tasks.
Multi-agent LLM benchmark using 12 AI agents with film personas to evaluate deliberation and reasoning capabilities.
Self-supervised learning framework for brain tumor classification using SSL methods (SimCLR, BYOL, DINO, Moco v3) on MRI data.
Personalized digital health modeling framework using adaptive weighting for heterogeneous user data with limited annotations.
Position paper on externalizing implicit knowledge in AI systems for reliability through human-AI collaboration infrastructure.
Model Spec Midtraining improves LLM alignment generalization by training on specification-aligned behavior before final fine-tuning.
NORA: multi-agent autonomous research system specialized for spatial data science workflows, automating scientific research end-to-end.
Proposes DGMM architecture addressing memory, temporal grounding, and interpretability limitations in LLMs through gist-based memory mechanisms.
arXiv paper proposing multi-agent framework with planner, actor, and memory manager roles for long-horizon LM-based task automation.
arXiv paper evaluating 1M-token context window retrieval and multi-hop reasoning in frontier LLMs on classical Chinese texts.
arXiv paper proposing intervention complexity as canonical reward measure for general intelligence in computable environments.
arXiv paper on uncertainty-guided exploration for stabilizing multi-turn reinforcement learning in agentic LLMs.
arXiv paper introducing MEMAUDIT protocol for evaluating long-term memory writing in budgeted LLM agents independent of retrieval and reasoning.
arXiv paper on clean-label backdoor attacks against vision-language models using diffusion model-generated triggers.
arXiv paper formalizing efficient benchmark selection for LLM evaluation as submodular maximization problem.
arXiv paper on efficient device-edge co-inference for vision-language models using speculative decoding on mobile devices.
arXiv paper on diagnosing neural network interpretations via input subspace partitioning for causal abstraction evaluation.
arXiv paper studying perturbation effects in recursive LLM loops using append, replace, and dialog context-update rules.
arXiv paper introducing PhysicianBench, a benchmark for evaluating LLM agents on long-horizon clinical workflows in real EHR environments.
arXiv paper evaluating zero-shot confidence estimation in small LLMs for local-to-cloud query routing without supervised training.
arXiv paper on belief revision postulates in multi-agent epistemic planning using Kripke models.
arXiv paper introducing open-source benchmark suite studying specification gaming failure mode in LLM agents across coding and non-coding tasks.
arXiv paper on model compression strategies using prerequisite graphs for LLMs in specialized engineering domains like circuit analysis.
EngiAgent coordinates multiple LLM agents to solve open-ended engineering problems requiring feasible solutions under data and physical constraints.
CoRD collaborative multi-teacher decoding framework distills long chain-of-thought reasoning through step-wise teacher collaboration and dynamic exploration.
Explores whether causal discovery algorithms can improve legal argument generation using Pearl's framework for probabilistic reasoning.
ANO optimizer unifies trust region frameworks to address PPO's hard clipping trade-off between sample efficiency and optimization stability.
Compound AI system aggregating 12,000 federal grants across fragmented US agency portals (NSF, NIH, DARPA) for unified grant discovery.
Controllable framework for synthesizing process supervision data with template-aware error injection and trajectory consistency for training process reward models.
HeavySkill framework analyzes heavy thinking as the core execution unit in agentic systems with orchestration, memory, skills, and tool use.
SCHEMA evaluation reveals cognitive collapse in frontier AI models under adversarial pressure, a safety failure mode beyond deception detection.
FitText framework makes tool retrieval dynamic in agent reasoning loops, bridging semantic gaps between task descriptions and API documentation across thousands of endpoints.
Power sampling method for efficient LLM decoding that locates high-probability solution modes by targeting p_theta(x)^alpha with future-dependent corrections.
Guide for evaluating LLM reasoning through adaptive multi-step search rather than final-answer accuracy, formalizing reasoning as search procedures.
Position paper exploring how graph structures can enhance LLM capabilities through knowledge representation, up-to-date information, and improved reasoning.
Open-source Shadow-Loom framework converts narratives into versioned graphical world models with causal physics and counterfactual reasoning engines.
GRAIL framework for efficient LLM-based agent discovery using SLM-enhanced indexing, achieving real-time semantic precision without 30+ second latencies.
DataClaw benchmark with 2.06M real-world datasets for evaluating autonomous data analysis agents on exploratory tasks and reasoning processes.
Dual-classifier GBDT pipeline distinguishes routine errors from high-risk misclassifications in critical ML applications across medical and classification domains.
SAGE framework teaches LLMs to choose effective modeling strategies for optimization problems through multi-strategy datasets and supervised fine-tuning.
Empirical study examining how task horizon length affects training dynamics and capabilities of LLMs as interactive agents.
Examination of foundation-model-based agent systems in industrial automation contexts, purposes, capabilities, and limitations.
Survey of counterfactual reasoning techniques in automated planning for handling deviations from fixed task specifications.