Neurosymbolic Retrievers for Retrieval-augmented Generation
Neurosymbolic retrievers for RAG combining neural and symbolic reasoning to improve interpretability and reduce hallucination.
Neurosymbolic retrievers for RAG combining neural and symbolic reasoning to improve interpretability and reduce hallucination.
Probabilistic framework for benchmarking AI system capabilities accounting for uncertainty in ground truth expert judgments.
Neuro-symbolic AI system combining LLMs with deterministic logic for autonomous business process reconfiguration and cross-functional automation.
OffSeeker demonstrates offline reinforcement learning for research agents reduces cost versus online RL while maintaining performance.
MemOCR: multimodal memory system for long-horizon agent reasoning that compresses interaction histories with layout-aware visual memory.
Framework identifying open problems in differentiable social choice mechanisms for ML systems including auctions, resource allocation, and LLM alignment.
SweetSpot analytical model predicts energy efficiency of LLM inference by modeling non-linear relationships in autoregressive Transformer structures.
ToolSelf framework enables LLM agents to self-reconfigure dynamically during task execution via tool-driven intrinsic adaptation beyond static agent configurations.
Gradient-based framework for interpretable failure detection and attribution in multi-agent RL systems including patient-zero identification and domino effects.
Studies whether reasoning models implicitly learn when to stop reasoning, addressing inefficiency and redundancy in long chain-of-thought outputs.
TSR algorithm for multi-turn RL training of LLM agents addressing sparse/delayed rewards and stochastic environments through trajectory-search rollouts.
CM2 applies RL with checklist rewards to multi-turn, multi-step agentic tool use for open-ended objectives without fully verifiable reward functions.
REMem framework enables language agents to perform reasoning using episodic memory with spatiotemporal context from interaction histories, not just semantic memory.
ForesightSafety Bench framework for evaluating frontier risks in autonomous AI systems addressing limitations of current safety benchmarks and alignment technologies.
Analyzes attention competition in local LLMs as predictor for dangerous behavior tipping without cloud connectivity or explicit oversight mechanisms.
E-SPL method jointly improves LLM context and weights by combining evolutionary prompt search with reinforcement learning for autonomous self-improvement.
CoreCraft environment in EnterpriseBench suite trains AI agents on high-fidelity RL simulation of enterprise customer support with 2,500+ entities and 23 tools.
Proposes Proxy State-Based Evaluation using LLM-driven simulation for benchmarking multi-turn tool-calling agents with scalable verifiable rewards.
Framework for evaluating AI agent reliability beyond accuracy metrics, examining consistency across runs and robustness to perturbations in deployed agents.
LLM-WikiRace benchmark evaluates LLM planning, reasoning, and world knowledge by requiring models to navigate Wikipedia hyperlinks to reach target pages efficiently.
Introduces RFEval framework for evaluating reasoning faithfulness in large reasoning models using stance consistency and causal influence under counterfactual interventions.
Proposes Kronecker-factored approach for disentangling task vectors in task arithmetic without external data, reducing cross-task interference.
Introduces ODESteer, a unified ODE-based framework for LLM alignment via activation steering that captures complex activation distributions.
Proposes distributed submodular maximization algorithm for multi-robot decision-making balancing resource constraints and coordination complexity.
Analyzes conditions under which unstructured pruning can induce layer collapse and structural effects in rectifier-activated networks.
Proposes adaptive Runge-Kutta dynamics for spatiotemporal prediction incorporating physical knowledge for weather forecasting and video processing.
Introduces PASS, a method using visual prompts with recurrent hypernetworks to identify effective structural sparsity patterns for neural network compression.
Proposes decoupled straight-through estimator for improving gradient-based optimization of neural networks with discrete variables.
Investigates LLM knowledge of thematic fit in semantic role prediction through prompt design experiments, achieving SOTA on benchmarks.
Proposes robust causal discovery method for validating agent-based models using time series data with improved accuracy on noisy data.
Presents FedCoLLM, a federated co-tuning framework enabling mutual enhancement between large and small language models with parameter efficiency.
Proposes PoTable, a plan-then-execute reasoning framework for improving LLM performance on table reasoning tasks.
Introduces LiveIdeaBench, a benchmark for evaluating LLMs' scientific idea generation and divergent thinking with minimal context.
Comprehensive empirical study and taxonomy of Graph Neural Networks for graph-level prediction tasks with analysis of evaluation methodologies.
Introduces benchmark for evaluating medical LLMs using dialogue-based diagnostic scenarios with noise and difficulty levels instead of static QA tasks.
Research identifying 'Curse of Depth' phenomenon where ~50% of LLM layers underperform. Analysis across Llama, Mistral, DeepSeek, Qwen with theoretical and empirical investigation.
Convergence analysis of multi-step temporal difference learning with linear function approximation.
Theoretical analysis of contextual combinatorial bandits with sparse rewards for recommendation systems.
Vector quantization approach for emergent language learning enabling self-play among communicating agents.
LLM agent using LoRA adapters for temporal-grounded video reasoning with interpretable evidence links.
Few-shot class-incremental learning using analogical reasoning for weight generation of new classes.
Survey of multi-turn LLM interactions covering instruction following, reasoning, tool use, and conversational understanding.
Pre-training approach for learning universal graph structural representations across different graph domains.
Attack method against autoencoders using gradient restoration to find more effective adversarial perturbations.
Knowledge distillation method using perception coherence loss to transfer representations to lightweight models.
Framework leveraging LLMs for code generation to solve PDEs, combining symbolic reasoning with neural approaches.
Fairness preprocessing framework using SHAP values for transparent data augmentation to mitigate model bias.
Client sampling strategy for federated learning addressing communication and computational heterogeneity.
Attack method against LLMs using bit-flip manipulation of model parameters to maximize inference costs.
Investigation of vulnerabilities accidentally introduced during LLM fine-tuning on domain-specific datasets.