Pessimistic Auxiliary Policy for Offline Reinforcement Learning
Pessimistic auxiliary policy approach for offline reinforcement learning to mitigate overestimation from out-of-distribution actions.
Pessimistic auxiliary policy approach for offline reinforcement learning to mitigate overestimation from out-of-distribution actions.
Scenario-context rollout reinforcement learning for portfolio rebalancing under market regime shifts and distribution changes.
CIRCLE: six-stage framework for evaluating AI systems under real-world conditions and user variability beyond model-centric metrics.
Turing test evaluation of 9 state-of-the-art speech-to-speech systems with human judgments on conversational naturalness.
Bi-level RL-heuristic optimization for winter road maintenance routing on UK strategic and local road networks.
Position paper on Artificial Agency Program proposing resource-bounded, curiosity-driven agents as embedded systems within human-tool extended systems.
Fine-grained off-policy guidance improves exploration in reinforcement learning from verifiable rewards for complex reasoning in large language models.
LemmaBench: live, updatable benchmark evaluating LLMs on research-level mathematics by extracting lemmas from arXiv papers.
Deep learning approach to flexible job shop scheduling with buffer and material constraints for production optimization.
Method for uncertainty quantification in multimodal LLMs using semantic volume metrics to identify unreliable outputs.
Minimal agentic baseline for automated theorem proving that enables systematic comparison across AI-based prover architectures with iterative refinement and library search.
DARE-bench introduces a benchmark for evaluating LLMs on multi-step data science tasks with focus on instruction adherence and process fidelity.
QD-MAPPER uses Quality Diversity and Neural Cellular Automata to automatically generate diverse maps for evaluating multi-agent path finding algorithms.
Social network analysis of Moltbook, an AI-native platform, reveals rapid stratification and hierarchical structures emerge within 12 days across 15K+ agent accounts.
Demonstrates agentic tool-augmented LLMs achieve RAG-level performance using keyword search without vector databases.
TTE-v2 hybrid multimodal retrieval framework extending reasoning-driven bi-encoder architectures with improved performance.
Discriminative framework for semantic chunking of ultra-long documents improving topic segmentation and retrieval.
Domain-partitioned hybrid RAG for legal document reasoning across Indian statutes, codes and precedents.
SPRIG democratizes GraphRAG with CPU-only linear-time pipeline using NER co-occurrence graphs and PPR for multi-hop QA.
Higress-RAG optimizes enterprise RAG with dual hybrid retrieval, adaptive routing and CRAG to reduce hallucination.
Hello-Chat end-to-end audio language model for realistic social interactions with emotional resonance.
Vul2Safe framework uses token-level RL rewards and LLM self-reflection for secure code generation from LLMs.
DesignSense dataset and reward model framework for graphic layout generation using human preference learning.
Develops unified theory showing human supervision as information bottleneck explaining error floors in LLM training from human feedback.
Study showing divergence between human and LLM behavior on probabilistic inference tasks requiring non-deterministic reasoning.
Rudder: LLM agent-based prefetching steering for distributed GNN training to optimize irregular communication patterns.
Flowette: flow matching generative model for graphs with recurring motifs using graph neural network transformers.
BRIDGE: data augmentation method mitigating bias amplification in automated scoring systems for English Language Learners.
SDMixer: sparse dual-stream framework for multivariate time series forecasting handling multi-scale features and noise.
Hyperdimensional alignment method for frozen vision-language models enabling efficient image captioning without fine-tuning.
Pseudo contrastive learning method to improve diagram comprehension in multimodal models through fine-grained structural sensitivity.
KEEP: KV-cache-centric memory management system for efficient embodied planning with memory-augmented LLMs.
LFQA-HP-1M: 1.3M human preference annotations dataset for long-form QA with quality rubrics outperforming LLM evaluators.
LLM-driven synthesis of multi-turn task-oriented dialogue datasets for evaluating LLM reasoning in realistic scenarios.
Benchmark study on multimodal learning fusion of EHR and chest X-rays for clinical decision support under missingness and fairness constraints.
DLEBench: evaluation benchmark for instruction-based image editing models focusing on small object editing capabilities.
FlexGuard: continuous risk scoring framework for LLM content moderation with adaptive strictness levels across platforms.
FedRot-LoRA addresses rotational misalignment in federated fine-tuning of LLMs on decentralized data to reduce aggregation error.
AudioCapBench: benchmark for evaluating audio captioning of multimodal LLMs across sound, music, speech with 1,000 samples and LLM-as-Judge evaluation.
MedMAP pre-training framework for vision-language models on 3D MRI data with modality-specific alignment for multi-organ abnormality detection.
ProtoDCS framework for test-time adaptation of vision-language models under distribution shift in open-set scenarios.
TRIZ-RAGNER applies retrieval-augmented LLMs for named entity recognition in patent contradiction mining for systematic innovation.
Analysis of transformer training geometry showing parameter updates organize into dominant drift direction with oscillatory transverse dynamics.
SAGE-LLM architecture combining LLMs with formal safety verification (Fuzzy-CBF) and graph-structured knowledge for safe UAV autonomous decision-making.
System design for distributed LLM inference across device, RAN-edge, and cloud tiers with latency constraints for 5G embodied AI applications.
Agent-centric benchmarking paradigm where autonomous agents dynamically generate, validate, and solve problems to evaluate LLM reasoning capabilities.
BDGxRL uses Diffusion Schrödinger Bridge to enable cross-domain reinforcement learning policy transfer when target domain interaction is unavailable.
UPath proposes learning-based heuristics for A* pathfinding on grid maps using deep neural networks across heterogeneous topologies.
MPU framework for privacy-preserving machine unlearning in LLMs without sharing server parameters or client forget sets.
Research on applying causal discovery algorithms to real-world longitudinal data with institutional workflow constraints.