Improving Reproducibility in Evaluation through Multi-Level Annotator Modeling
Research on improving LLM evaluation reproducibility by modeling annotator biases in human rating systems for AI safety assessment.
Research on improving LLM evaluation reproducibility by modeling annotator biases in human rating systems for AI safety assessment.
Neurosymbolic approach combining LLMs with SMT solvers to audit natural-language software requirements for ambiguity and inconsistency.
Negation Neglect: LLMs fail to learn negations during finetuning, believing falsified claims despite recognizing them as false in context.
EVA-Bench: Evaluation framework for voice agents addressing realistic conversation simulation and voice-specific failure mode measurement.
WARDEN: Language model for transcribing and translating endangered Wardaman language using only 6 hours of annotated audio data.
ChatSR: Multimodal LLM for scientific formula discovery from graphs and images with specialized scientific data understanding.
Taxonomy and survey of AI safety landscape for large language models covering design, development, adoption and deployment impacts.
Language model networks study: pre-trained LMs as reusable nodes in inference systems with learned communication patterns instead of natural language.
Block-wise adaptive caching technique to accelerate Diffusion Policy inference for real-time robotic control.
CADDesigner: LLM-powered agent for conceptual CAD modeling accepting text and sketches with interactive refinement dialogue.
SkillNav: Modular framework for vision-language navigation using skill-based decomposition to improve generalization on complex spatial-temporal reasoning.
Critical analysis of temporal signals in benchmark contamination detection, showing sensitivity to question construction independent of data memorization.
QuickLAP: Bayesian framework for robot learning that fuses physical corrections and language feedback to infer reward functions.
Prismatic world model for robotic planning that learns compositional dynamics separately for distinct physical modes like contact/impact events.
MobiBench: Multimodal benchmark for mobile GUI agents addressing limitations of existing offline/online benchmarks with multiple valid action paths.
Differentiable Evolutionary RL optimizes reward functions using gradient information to improve policy performance on complex reasoning tasks.
PersonalAlign: GUI agent framework that aligns with implicit user intents by leveraging long-term user records as persistent context for personalized task completion.
PolySHAP improves KernelSHAP approximation of Shapley values for explainable AI by using polynomial regression instead of linear approximation to reduce computational cost.
Study examining how elicitation protocol design affects stated-revealed preference gaps across 24 language models.
THINKSAFE: Safety alignment approach for reasoning models that self-generates safety constraints without external teacher distillation.
M2CL: Multi-LLM context learning method for multi-agent discussion systems addressing discussion inconsistency and context misalignment.
DiscoverLLM: Method enabling LLMs to help users discover intents through interactive exploration rather than just executing stated requests.
SupChain-Bench: Benchmark for evaluating LLMs on real-world supply chain management with multi-step domain-specific orchestration.
AMOR: Hybrid architecture combining recurrence and attention, selectively invoking attention based on predictive uncertainty.
Proxy State-Based Evaluation: LLM-driven benchmark method for multi-turn tool-calling agents that scales beyond deterministic backends.
LLM-based approach for time series question answering using pattern alignment and balanced reasoning across task complexity.
Interactive Benchmarks: Evaluation paradigm assessing model reasoning by testing ability to decide what information to acquire and use.
GAAMA: Graph-augmented memory architecture for AI agents to maintain coherent long-term personalized behavior across sessions.
Trajectory-level safety benchmark (ATBench) for evaluating LLM-based agents across realistic multi-step interactions with diverse failure modes.
Benchmark evaluating whether LLM agents can autonomously design, implement, and execute RL post-training pipelines for model improvement.
Framework for training open-weight language models to simulate student coding behavior for educational tutoring system evaluation.
Industrial evaluation system for LLM-generated meeting summaries with structured ground-truth construction and privacy-bounded monitoring.
Evaluates three LLM agent interaction paradigms (structured tools, computer-use, coding agents) on scientific visualization tasks across 15 benchmarks.
Study on coordinated flow models for offline multi-agent reinforcement learning that balances efficiency and coordination.
Paper arguing automated alignment via research agents risks producing misleading safety assessments without deliberate sabotage.
AI co-mathematician workbench enabling mathematicians to collaboratively leverage AI agents for research including ideation, computation, and theorem proving.
Method to extract and analyze search trees from LLM reasoning traces to quantify planning capabilities and identify myopic decision-making.
Agentick: Unified benchmark enabling fair comparison of RL, LLM, VLM, and hybrid agents on sequential decision-making tasks.
SREGym: High-fidelity benchmark for AI SRE agents with live cloud-native systems and realistic failure scenarios for diagnosis and mitigation.
BoostAPR: Three-stage RL framework for automated program repair using execution feedback and dual reward models for credit assignment.
FORTIS benchmark evaluating privilege escalation vulnerabilities in LLM agent skill layers and access control policies.
SimWorld Studio: System using evolving coding agents to automatically generate diverse 3D environments for embodied agent training.
EpiGraph knowledge graph and benchmark for evaluating knowledge-augmented clinical reasoning in epilepsy diagnosis from heterogeneous evidence.
Expo: Improved RL policy optimization for LLM reasoning with adaptive KL regulation and curriculum sampling over GRPO baseline.
IndustryBench: 2,049-item benchmark testing LLM knowledge on industrial procurement with safety-critical constraints and standards compliance.
Genetic programming approach to symbolic regression with gene editing for discovering mathematical formulas from scientific data.
SORT method for reinforcement learning that adds repair updates for failed rollouts using plan guidance and token probability weighting.
Study on resolving safety-helpfulness trade-offs in LLM alignment through preference dimensional expansion approach.
Research on multi-modal world models integrating tactile and visual feedback for predicting robotic action outcomes in complex environments.
Benchmark for evaluating LLM text-to-SQL performance on complex enterprise databases with intricate schemas and domain knowledge requirements.