XKD-Dial: four-stage training pipeline for citation-grounded dialogue reducing hallucination in English-Hindi LLMs. LLM application addressing hallucination.
arXiv paper examining regulatory frameworks for agentic AI security and privacy. Policy analysis of AI agent governance.
PRIOR framework for humanoid locomotion with natural gaits using Isaac Lab. ML for robotics, not core AI agent/LLM focus.
Benchmark evaluating AI agent performance on domain-specific data science tasks against human expert baselines across multiple domains.
RAG method using hypothesis-conditioned query rewriting to retrieve decision-relevant evidence for choice tasks beyond topical relevance.
Framework enabling LLM agents to recognize secure trusted execution environments for secure IP disclosure negotiations.
Multilingual temporal reasoning benchmark with 15K examples across 5 languages testing LLM capabilities on date arithmetic and temporal relations.
Post-hoc debiasing method for vision-language models like CLIP using sparse embedding modulation to separate bias from semantic information.
Streaming video understanding framework that decouples semantic understanding from perception for proactive query handling.
Study comparing LLM-generated analogies to human-produced ones using geometric parallelogram model of analogical relations.
Neural solver for multi-objective multi-agent traveling salesman problem using conditional learning approach.
Framework for steering safety judgments in vision-language models through semantic cues without parameter changes.
RAG benchmark and framework for multilingual multi-hop question answering across multiple languages and corpora.
Internal representation debiasing framework for LLMs using graph isomorphism to remove social biases from model embeddings.
RL-based policy optimization method for improving low-resource language model performance through structural constraints on tokenization.
Multi-agent framework for vision-language navigation with probabilistic grounding of spatial references and metric constraints.
Benchmark for GPU kernel optimization spanning 235 CUDA problems from production AI models, measuring proximity to hardware efficiency limits.
Nemotron-Cascade 2 open 30B MoE model using cascade RL and multi-domain distillation achieving IMO gold-medal-level mathematical reasoning.
F2LLM-v2 multilingual embedding models (80M-14B parameters) supporting 200+ languages with emphasis on low-resource language coverage.
FinTradeBench benchmark for evaluating LLM reasoning on financial decision-making using company fundamentals and trading signals.
NavTrust benchmark evaluating robustness of embodied navigation agents (VLN and OGN) under real-world data corruptions.
Framework combining machine learning with automated reasoning for generating and selecting explanations in scientific discovery tasks.
Analysis of LLM-based world models for decision-making in reasoning systems, identifying evaluation gaps and methodological issues.
Framework using LLM agents to simulate decision discourse by representing diverse stakeholder perspectives in complex problem-solving.
Manus AI general-purpose autonomous agent combining LLM reasoning with execution capabilities for complex end-to-end tasks.
Deep reinforcement learning approach for multi-objective combinatorial optimization using conditional computation and preference decomposition.
Multimodal learning framework for solving Generalized Traveling Salesman Problem in robotic task planning.
Single-agent reinforcement learning framework for bus fleet control addressing traffic stochasticity and demand variability.
MMSearch-Plus benchmark for multimodal browsing agents requiring genuine vision-text reasoning and iterative retrieval verification.
CausalARC testbed for evaluating AI reasoning on abstract tasks with limited data and distribution shift using causal world models.
Bayesian evaluation framework replacing Pass@k metric for more stable and reliable LLM reasoning performance assessment.
SynBullying dataset uses multiple LLMs to generate synthetic conversational data for cyberbullying detection research.
AgroCoT benchmark evaluates reasoning capabilities of vision-language models for agricultural applications like crop monitoring and pest detection.
Memory Bear system applies cognitive science principles to address LLM memory limitations, hallucinations, and context window constraints.
Certification protocol ensuring consistent semantic understanding between agents using stimulus-meaning model and empirical testing.
Data-centric framework learning optimal verbalization for converting user interaction logs into natural language for LLM-based recommendation systems.
Variable isolation study examining prompt architecture layers enabling LLMs to solve reasoning benchmarks like the car wash problem.
CIRCLE lifecycle framework bridging gap between AI model metrics and real-world deployment outcomes through six-stage evaluation.
AI4S-SDS system combining LLM agents with sparse MCTS and differentiable physics for automated chemical solvent design.
MEMO framework reducing variance in multi-turn multi-agent LLM game evaluations through memory augmentation and context optimization.
MedMASLab unified framework and benchmark for multimodal medical multi-agent systems with standardized integration and cross-specialty evaluation.
SoLA framework for reversible lifelong model editing in LLMs using semantic routing with LoRA modules to prevent knowledge forgetting.
Method to reduce overthinking and underthinking in Large Reasoning Models through balanced token allocation for efficient inference.
VTC-Bench evaluating multimodal LLM agents on complex visual tool composition, addressing limitations in existing tool-use benchmarks.
Hybrid scalar-verbal RL approach for emotional support dialogue systems using user reactions as learning signals instead of expert-defined rewards.
AsgardBench benchmark for evaluating visually-grounded interactive planning and plan adaptation based on visual observations.
Formal proof that safety is non-compositional when combining agents with conjunctive capability dependencies.
ARISE hierarchical RL framework for mathematical reasoning in LLMs that learns reusable strategies across problem instances.
Machine learning approach for predicting and discovering error patterns in vehicle diagnostic trouble codes using temporal sequence analysis.
Study of nonstandard errors in AI coding agents deploying 150 Claude agents on market analysis tasks, showing agent-to-agent variation in analytical choices.