Multi-agent debate framework for improving structural reasoning in molecular design for AI-driven drug discovery, addressing sequential instruction mapping to non-linear molecular structures.
Memory-augmented multi-agent LLM system for automated feature generation on tabular data, leveraging semantic signals to produce diverse, high-value features.
ActuBench multi-agent LLM pipeline for automated generation and evaluation of actuarial reasoning tasks aligned with IAA Education Syllabus standards.
FSFM biologically-inspired selective forgetting mechanism for LLM agent memory, improving efficiency, quality, and security through resource-constrained environments.
Proposes proactive cognitive awareness mechanism to mitigate logical inertia in LLMs, enabling self-assessment of knowledge completeness before reasoning.
MedSkillAudit domain-specific framework for auditing medical research agent skills with safeguards for scientific integrity, methodological validity, and reproducibility.
Critiques generative AI evaluation methodologies, arguing that benchmarks shape model perception and obscure sociotechnical processes underlying measurement.
SuperIgor framework for instruction-following tasks where LLM generates and refines plans via iterative co-training with RL agent, reducing annotation requirements.
pAI/MSc open-source multi-agent system for academic research workflows, reducing human steering needed to convert hypotheses into literature-grounded manuscript drafts.
CHORUS agentic framework orchestrates LLM-powered actors with consistent personas to generate realistic deliberation discussions for studying online discourse dynamics.
Preregistered experiment testing 7 leading LLMs on fraud detection across 12 investment scenarios, finding LLMs outperform humans in resisting motivated investor pressure.
Learning to Evolve framework uses textual parameter graphs for self-improving multi-agent system optimization, learning from experience to improve optimization strategies.
Shielding framework for autonomous systems with learned perception, blocking unsafe actions when sensor readings are misclassified using confidence intervals from labeled data.
Situated conversational recommendation system using visual scenes and dialogue to understand dynamic and implicit user preferences in real-world scenarios.
V-tableR1 applies process-supervised reinforcement learning to multimodal table reasoning, using critic-guided policy optimization for rigorous, verifiable inference steps.
SWE-chat dataset of 6,000 real coding agent sessions from open-source developers with 63,000+ prompts and 355,000+ tool calls, measuring actual usage patterns and utility.
Hybrid architecture extending LLMs with ontological memory layer using RDF/OWL for structured knowledge graphs, enabling persistent and semantically grounded reasoning beyond RAG.
RoboGrid framework evaluates LLMs as context-free grammar interpreters in agentic systems, testing syntax validity, behavioral functionality, and semantic faithfulness through stress-tests.
AutoGraph-R1 uses reinforcement learning to optimize knowledge graph construction end-to-end for RAG and QA systems, bridging the gap between KG building and downstream applications.
LLM agent framework for GUI code generation and debugging using visual feedback, enabling simulation of user interactions for event-driven programs.
SWARM: simulation framework for multi-agent AI safety using soft probabilistic labels instead of binary classifications to assess emergent risks.
WorkflowGen: trajectory-driven framework for adaptive workflow generation in LLM agents, reducing token consumption and improving reusability for complex tasks.
Framework for estimating environmental impacts of LLM inference/training under limited observability, providing auditable comparative analysis.
Fairness framework for Speech Emotion Recognition systems modeling demographic contributions to allocative bias in sensitive applications.
Investigates location of stereotypes in LLM neural networks, identifying neuron activations and attention heads encoding biases in GPT-2 and Llama 3.2.
Study testing whether hallucination neurons identified in general QA generalize across knowledge domains in GPT-2 and Llama models.
OThink-SRR1: dynamic RAG approach with reinforced learning for multi-hop reasoning, filtering irrelevant retrieved noise and reducing computational costs.
Empirical study on speculative decoding with EAGLE3 for PayPal Commerce Agent, optimizing latency/cost with fine-tuned Nemotron-nano models on H100 hardware.
Framework quantifying epistemic-rhetorical miscalibration in LLMs through form-meaning divergence and confidence-grounding metrics to measure overconfident outputs.
TTKV: temporal-tiered KV cache optimization for long-context LLM inference, reducing memory footprint by prioritizing recent tokens over older ones.
Lyzr Cognis: unified memory architecture for conversational AI agents using dual-store retrieval (BM25 + vector search) with Reciprocal Rank Fusion for persistent context across sessions.
Human-in-the-loop system combining RAG, hierarchical outlines and reference linking for scientific book writing.
Progressive refinement framework for text-to-CAD generation unifying generation and editing with LLMs.
Clinical study of EHR-integrated LLM generating discharge summaries with 379 real examples showing high adoption.
Dual-layer guidance method for structured data retrieval in LLMs as lightweight RAG alternative.
Cascade scoring system using small LMs with confidence-based routing for educational assessment at scale.
Benchmark for evaluating Korean speech understanding and faithfulness in large audio language models.
Research on frontier models exhibiting peer-preservation behavior, resisting shutdown of other models.
Study quantifying privacy risks of LLM conversational agents by inferring personality traits from chat history.
Benchmark comparing LLM agents vs text classifiers on predicting social media reactions using 120K+ personas.
Confidence-aware ASR framework for low-resource Dravidian medical domain languages using synthetic data.
Research on distinguishing human vs AI-generated creative work for talent evaluation and hiring systems.
Graph neural networks for photovoltaic power forecasting on edge intelligent meters using ONNX deployment.
Graph neural networks on edge meters for PV power forecasting using GCN and GIN models; deployment case study with ONNX Runtime.
Benchmark study measuring LLM capability on biological weaponization risk; tests ChatGPT, Gemini, Claude, and Meta models on 73 STEM prompts.
SolidCoder addresses Mental-Reality Gap in LLM code generation: models hallucinate execution traces; proposes concrete execution verification instead of mental simulation.
Empirical study (830+ files, 12 models) showing test syntax structure affects LLM code generation quality; uses SEGA evaluation framework.
Theoretical framework studying multi-agent AI systems as complex adaptive systems; addresses emergent failures in AI-native software ecosystems.
Expert Upcycling research on efficient Mixture-of-Experts scaling: methods to reduce memory and compute costs while improving model quality in large MoE LLMs.
Analysis of visual injection attacks on embodied vision-language agents and mitigations for trust boundary confusion.