Unlocking Exploration in RLVR: Uncertainty-aware Advantage Shaping for Deeper Reasoning
Uncertainty-aware advantage shaping method improving exploration in reinforcement learning for enhanced LLM reasoning capabilities.
Uncertainty-aware advantage shaping method improving exploration in reinforcement learning for enhanced LLM reasoning capabilities.
Framework addressing LLM failures as cognitive rather than purely technical, distinguishing hallucinations from associative pattern reproduction.
MCP-Flow framework enabling LLM agents to discover and utilize diverse Model Contextual Protocol tools at scale for real-world tasks.
BarrierBench evaluates LLMs on safety verification of dynamical systems using barrier certificates for autonomous applications.
IMACT-CXR multi-agent conversational tutoring system using AutoGen for interactive chest X-ray interpretation training.
RILKE method for efficient knowledge updates in LLMs via representation interventions without costly retraining in lifelong settings.
Argues static value alignment insufficient for robust AI alignment under capability scaling and distributional shift; proposes dynamic approach.
FinWorkBench benchmark evaluating AI agents on enterprise finance workflows spanning data entry, retrieval, calculation, and reporting tasks.
MAGMA multi-graph memory architecture for agents separating temporal, causal, and entity information to improve retrieval and reasoning.
Skill-Pro framework enabling LLM agents to learn and reuse procedural skills from experiences without parameter updates via non-parametric PPO.
Entity State Tuning method for temporal knowledge graph forecasting using stateful entity representations across timestamps.
Case study of human-AI collaboration discovering novel error bounds for Hermite quadrature rules in mathematical research.
Conformal inference method to regulate untested policies against safe reference policies for high-stakes agent deployment.
Neuro-symbolic solvent design system using sparse MCTS and differentiable physics to overcome LLM agent limitations in combinatorial chemistry.
Analyzes composition gaps in health AI benchmarks validating clinical LLMs without defined patient population characteristics.
Develops methods to measure metacognitive capabilities of AI systems for uncertainty assessment and decision reliability.
Uses direct preference optimization to reduce LLM biases toward spurious social contexts in high-stakes decision-making tasks.
Proposes Rashomon Memory architecture for AI agents to maintain multiple conflicting interpretations of events for concurrent goals.
CODESTRUCT reframes code editing from text manipulation to structured AST operations, enabling LLM code agents to apply syntax-validated transformations on repositories.
CWCD applies category-wise contrastive decoding to multi-modal LLMs for improved chest X-ray medical report generation with enhanced diagnostic accuracy.
Introduces Proactive Information Probing task for customer service chatbots to strategically extract business intelligence while minimizing conversation friction and user effort.
Context Kubernetes orchestrates enterprise knowledge delivery to agentic AI systems with proper permissions and freshness, formalizing six design dimensions for knowledge orchestration.
Proposes longitudinal health agent framework addressing user intent and accountability for multi-turn health tasks like symptom management and behavior change.
StsPatient framework simulates cognitively impaired standardized patients for clinical training using fine-grained LLM steering across domain-specific deficits and severity levels.
Paper proposes HCoT, integrating expert system heuristics into LLMs for structured reasoning to address stochastic token generation and decoupled decision-making in complex problem-solving.
Safe reinforcement learning with online filtering for human-robot task allocation in manufacturing that constrains worker fatigue while maximizing efficiency.
Hierarchical spatial-aware reinforcement learning algorithm for human-robot task planning and allocation in manufacturing with efficient coordination.
Text2Model suite introduces LLM-based copilots for text-to-optimization translation with varying complexity strategies and cross-domain dataset Text2Zinc.
Proposes learning-based adaptive security mechanisms for Web3 decentralized applications addressing social, application, and protocol-layer attacks.
DA-Cramming improves cost-effective BERT pretraining by integrating dependency agreement, extending Cramming approach for single GPU training.
Survey examining application of generative models in connected and automated vehicles for predictive modeling, simulation, and decision-making.
Explores LLM capabilities for symbolic regression using in-context learning with GPT-4 models to suggest expressions that are optimized against datasets.
Introduces intentional analysis framework to improve language model reasoning by explicitly modeling intent as cognitive notion underlying human communication and problem-solving.
Deep Neural Lesion (DNL) identifies critical parameters in DNNs through data-free and optimization-free methods showing vulnerability to single bit-flip attacks.
IMPACTX leverages explainable AI as automated attention mechanism to improve model performance without external knowledge or manual constraint specification.
Generates nuanced, time-aware impact summaries of scientific papers by analyzing confirmation and correction citations across temporal evolution.
Method for estimating optimal loss values in diffusion models to improve diagnosis and performance by distinguishing between large optimal loss and insufficient model capacity.
Proposes instruction inference task for human-agent collaboration where agents infer incomplete instructions by modeling human mental states and shared context understanding.
Time-RA reformulates time series anomaly detection as a generative reasoning task using LLM feedback, introducing RATs40K dataset for fine-grained anomaly categorization and explanation.
Study evaluates trustworthiness of LLMs on ambiguous Chinese text using a benchmark dataset of disambiguated sentence pairs to assess model behavior under linguistic ambiguity.
SPaCe introduces self-paced curriculum learning to reduce data and compute requirements for LLM fine-tuning with reinforcement learning by sampling training examples based on difficulty and learning value.
DPQuant combines dynamic quantization scheduling with differentially-private SGD to reduce training time, energy, and cost while maintaining privacy.
Semantic Resonance Architecture routes tokens to experts in MoE models via cosine similarity to learnable semantic anchors, improving interpretability of routing decisions.
AISysRev is a containerized LLM-based tool for automated title-abstract screening in systematic literature reviews.
DeepPrune optimizes parallel LLM reasoning by eliminating inter-trace redundancy, reducing 80% computational waste in multi-trace inference.
LLM watermarking technique using syntactic predictability for verifiable output attribution while preserving text quality.
E2EDev benchmark evaluates LLMs on end-to-end software development tasks with fine-grained requirements and behavior-driven evaluation.
PatMD detects harmful memes using MLLMs and misjudgment risk patterns. ML research on content moderation.
Chart2Code benchmark evaluating chart understanding and code generation in multimodal LLMs across three difficulty levels with real-world scenarios.
Vector symbolic architectures using histogram recovery from random linear codes for neurosymbolic AI systems and hardware implementations.