Retrieval-augmented generation system for extracting and reasoning over polymer science knowledge from unstructured literature.
Resource-efficient method for enforcing multilingual consistency in LLM safety alignment across linguistic communities.
LLM-based approach for C unit test generation via scenario planning and reasoning to avoid premature code synthesis.
Framework for cost-aware exploration in LLM agents to optimize exploration-exploitation tradeoffs when interacting with environments.
RCT measuring whether LLM assistance improves novice performance in biology laboratory tasks and dual-use skill acquisition.
Policy compiler providing deterministic enforcement of authorization policies for LLM-based agents with information flow tracking.
Review and practical guide for selecting context-appropriate fairness metrics in machine learning models across regulatory requirements.
Multi-agent workflow with vision-language models and chain-of-thought reasoning for analyzing robotic surgical video scenes.
Agent-based framework enabling natural language interaction with water distribution system simulator via LLM wrapper around EPANET.
Evaluation benchmarks and litmus tests measuring economic decision-making capabilities of LLM agents in procurement, scheduling, and pricing.
Benchmark for generative learning on dynamic text-attributed graphs with improved textual quality for semantic richness.
Computational framework modeling how AI systems integrate linguistic guidance with direct sensorimotor experience for social learning.
Dataset and reasoning framework for teaching LLMs to perform genuine time series reasoning beyond surface-level pattern matching.
Method for controlling specific attribute intensities in LLM outputs via targeted representation editing rather than directional guidance.
Geospatial AI foundation models with agentic reasoning for analyzing large-scale satellite and environmental data.
Framework transforming LLMs into stateful runtime operators via dual-stream architecture to handle complex long-horizon tasks with better state management.
Multi-agent LLM system for identifying valid and specific weaknesses in scientific papers using structured reasoning to avoid reviewer bias.
LLM agent framework for molecular structure optimization achieving high sample efficiency by using trajectory-aware learning in drug discovery.
Research on optimizing tool-use behavior in LLM agents by analyzing entropy patterns to reduce excessive tool calls and improve inference performance.
Open-source safety evaluation framework for generative AI chatbots in mental health applications, examining clinical validity and reliability.
Benchmark evaluating robustness of LLM-based agents performing tool use under noisy, real-world conditions versus idealized test environments.
LaySPA: Reinforcement learning framework equipping LLMs with spatial reasoning capabilities for content-aware graphic layout design.
AI safety evaluation framework addressing frontier risks and systemic challenges in increasingly autonomous AI systems with unpredictable behaviors.
Evaluation framework using negotiation games to assess language model agency through multi-turn interactions and complex reasoning scenarios.
Multi-concept personalization approach for vision-language models, enabling understanding of multiple user-provided concepts and their interactions.
Teaching spatial reasoning to vision-language models for robotics through training data enhancement and fine-tuning techniques.
Framework enabling human-agent co-planning and co-execution, allowing deeper collaboration in long-running AI agent tasks beyond autonomous workflows.
Content moderation technique using soft prompt guidance to prevent unsafe content generation in text-to-image models.
Federated parameter-efficient fine-tuning method with adaptive rank allocation for distributed training of language models on resource-constrained devices.
Technique to extract safety classifiers embedded in aligned LLMs to understand and potentially improve jailbreak attack resistance.
Analysis of Transformer optimization through gradient heterogeneity lens, explaining why Adam outperforms SGD in training large models.
Research on continual learning paradigm shift: leveraging abundant GPU memory instead of minimizing exemplar storage to mitigate catastrophic forgetting.
Combining chain-of-thought reasoning and RAG to improve rare disease diagnosis from unstructured clinical notes using LLMs.
Investigation of test-time scaling for medical reasoning with LLMs, exploring computation allocation during inference for improved diagnostic capabilities.
Federated learning method addressing noisy labels in distributed training through enhanced forward correction techniques.
Open-vocabulary mapping framework for robot exploration combining vision-language models for real-time semantic understanding of unknown environments.
PLAICraft is a large-scale, multi-modal, time-aligned dataset of Minecraft interactions for training embodied AI agents with vision, speech, and action.
VerifyBench benchmarks reference-based reward systems used in reinforcement learning training of reasoning models like o1 and DeepSeek-R1.
WINA is a training-free sparse activation method for accelerating LLM inference by selectively activating neurons based on weights.
XENON is an LLM-based agent that algorithmically corrects flawed knowledge through experience-based learning for long-horizon planning in Minecraft.
FreqPolicy accelerates flow-based visuomotor policies for robotic manipulation using frequency consistency for real-time inference.
DiffusionBlocks enables block-wise neural network training via diffusion interpretation to reduce memory bottlenecks in transformers.
Studies optimal ordering of chain-of-thought reasoning steps in transformers for arithmetic and multi-step reasoning tasks.
Theoretical analysis of graph transformers' expressive power using logic, covering both real numbers and floating-point settings.
Model-agnostic dynamic feature selection method with uncertainty quantification for budget-constrained decision-making scenarios.
MedReasoner uses reinforcement learning to ground clinical reasoning to pixel-level regions in medical images via multimodal LLMs.
Pinet introduces an output layer using orthogonal projections to enforce convex constraints in neural networks during training and inference.
FairTabGen uses LLMs to generate high-quality synthetic tabular healthcare data with fairness constraints from limited samples.
COGITAO is a benchmark framework for studying compositionality and generalization in visual reasoning tasks, inspired by ARC-AGI.
Empirical study of pre-trained model reuse and integration in open-source projects, defining Software Dependencies 2.0.