Can Causal Discovery Algorithms Help in Generating Legal Arguments?
Explores whether causal discovery algorithms can improve legal argument generation using Pearl's framework for probabilistic reasoning.
Explores whether causal discovery algorithms can improve legal argument generation using Pearl's framework for probabilistic reasoning.
ANO optimizer unifies trust region frameworks to address PPO's hard clipping trade-off between sample efficiency and optimization stability.
Compound AI system aggregating 12,000 federal grants across fragmented US agency portals (NSF, NIH, DARPA) for unified grant discovery.
Controllable framework for synthesizing process supervision data with template-aware error injection and trajectory consistency for training process reward models.
HeavySkill framework analyzes heavy thinking as the core execution unit in agentic systems with orchestration, memory, skills, and tool use.
SCHEMA evaluation reveals cognitive collapse in frontier AI models under adversarial pressure, a safety failure mode beyond deception detection.
FitText framework makes tool retrieval dynamic in agent reasoning loops, bridging semantic gaps between task descriptions and API documentation across thousands of endpoints.
Power sampling method for efficient LLM decoding that locates high-probability solution modes by targeting p_theta(x)^alpha with future-dependent corrections.
Guide for evaluating LLM reasoning through adaptive multi-step search rather than final-answer accuracy, formalizing reasoning as search procedures.
Position paper exploring how graph structures can enhance LLM capabilities through knowledge representation, up-to-date information, and improved reasoning.
Open-source Shadow-Loom framework converts narratives into versioned graphical world models with causal physics and counterfactual reasoning engines.
GRAIL framework for efficient LLM-based agent discovery using SLM-enhanced indexing, achieving real-time semantic precision without 30+ second latencies.
DataClaw benchmark with 2.06M real-world datasets for evaluating autonomous data analysis agents on exploratory tasks and reasoning processes.
Dual-classifier GBDT pipeline distinguishes routine errors from high-risk misclassifications in critical ML applications across medical and classification domains.
SAGE framework teaches LLMs to choose effective modeling strategies for optimization problems through multi-strategy datasets and supervised fine-tuning.
Empirical study examining how task horizon length affects training dynamics and capabilities of LLMs as interactive agents.
Examination of foundation-model-based agent systems in industrial automation contexts, purposes, capabilities, and limitations.
Survey of counterfactual reasoning techniques in automated planning for handling deviations from fixed task specifications.
Position paper arguing causality is necessary to resolve conflicts between trustworthy AI objectives like fairness and robustness.
Theoretical analysis of shortcut learning in deep neural networks using evolutionary game theory framework.
AcademiClaw: bilingual benchmark of 80 complex long-horizon academic tasks from real student workflows for evaluating AI agents.
Zero-trust security framework for LLM-driven agents using hybrid inspection and task-based access control for tool invocation.
Empirical study of 557 healthcare agent skills, analyzing procedural adaptations and governance for healthcare AI agents.
ORPilot: open-source agentic LLM system that translates business problems into optimization models for production use.
Learning to defer approach for hierarchical multi-label medical imaging decisions where models can defer to experts.
ReClaim: generative transformer foundation model for extracting insights from large-scale medical claims data.
Research on misalignment contagion between multiple LMs in multi-turn interactions and steering techniques using implicit traits.
System for designing user workflows with hard and soft constraints in LLM-based planning, addressing reliability and control challenges.
Philosophical comparison of human agency development with potential LLM agency. Argues for joint action/planning architecture with humans.
SCPRM: Schema-aware reward model for knowledge graph question answering. Addresses risk compensation in process rewards for LLM reasoning paths.
First-order efficiency methods for Shapley/Semivalues computation via statistical viewpoint. Model-agnostic feature attribution for explainability.
JACTUS: Joint adaptation and compression for large models across diverse tasks. Simultaneous parameter-efficient fine-tuning and low-rank compression.
HAAS framework for policy-aware adaptive task allocation between humans and AI. Addresses complementary roles and contextual task distribution.
Knowledge distillation approach for cross-language code clone detection using compact open-source models. Improves semantic detection reliability.
MCP Workflow Engine: Orchestration layer decoupling LLM reasoning from execution via Model Context Protocol. Reduces token consumption for repeated tasks.
GhostServe: Fault-tolerant checkpointing system for million-token LLM agent serving. Addresses KV cache challenges in long-running inference tasks.
Agentopic: Multi-agent workflow for explainable topic modeling using LLMs. Collaborative agents handle identification, validation, and hierarchical grouping.
Analysis of 150,000+ job postings examining generative AI's impact on workforce skill requirements. Tests augmentation vs. substitution across sectors.
Study of correlated forecasting errors across GPT-4o, Claude, and Gemini on 568 binary predictions. Shows epistemic monoculture with mean error correlation r=0.77.
UniQGen: LLM agent framework for knowledge graph question answering across RDF/SPARQL and Cypher. Constraint-guided query generation for property graphs.
H-probes method to extract hierarchical structures from LLM latent representations. Analyzes geometric representations enabling hierarchical reasoning.
Open Earth System Foundation Model extending Aurora for weather/climate forecasting. Unified framework for heterogeneous geophysical data integration.
Evaluation framework for text-to-speech synthesis quality using crest factor, spectrum balance, and cepstral metrics across six TTS models.
BRITE: benchmark for text-to-video evaluation on implausible scenarios including audio-visual alignment and QA-based interpretable assessment.
Latent space probing framework for detecting adult content in video generative models by analyzing internal representations during generation.
OceanPile: multimodal ocean corpus addressing fragmented ocean data for training foundation models on climate and marine biodiversity tasks.
X2SAM integrates MLLMs with foundation segmentation models enabling pixel-level perception from conversational instructions across images and videos.
Retrieval-guided generation approach for medical image captioning reducing hallucinations and factual inconsistency in histopathology image descriptions.
DIAGRAMS: review framework for creating reasoning-level attribution in diagram QA, providing structured evidence annotation across visual diagrams and infographics.
TRIP-Evaluate: open multimodal benchmark for assessing LLMs and MLLMs on transportation tasks including regulation QA, traffic management, and autonomous driving reasoning.