World-Model-Augmented Web Agents with Action Correction
WAC: LLM-based web agent with world model for reasoning about environment changes and predicting execution risks before action.
WAC: LLM-based web agent with world model for reasoning about environment changes and predicting execution risks before action.
Hybrid approach combining abstention and adaptive detection to improve LLM safety-utility tradeoff without static confidence thresholds.
EduEVAL-DB dataset with 854 teacher explanations across K-12 science/language/social science for training pedagogical evaluators and AI tutors.
Framework for evaluating construct validity of LLM benchmarks, separating actual capabilities from benchmark artifacts like contamination and annotation errors.
On-device graph-based reasoning system for personalized AI with transparency and accountability, addressing limitations of RAG-based black-box systems.
Layer-wise analysis of multimodal Transformers using information theory to decompose how visual and linguistic information contribute to predictions in vision-question tasks.
Research on evaluating vision-language model reasoning in autonomous driving scenarios beyond performance metrics.
Presents PERSONA, a training-free framework for dynamic personality control in LLMs using activation vector algebra without fine-tuning.
Introduces recursive concept evolution method to improve compositional reasoning in LLMs on benchmarks like ARC-AGI-2 and MATH by evolving latent representations.
Proposes GlobeDiff, a diffusion-based approach for multi-agent coordination under partial observability, addressing limitations of belief-state and communication methods.
Methodological study validating when LLM simulations produce valid behavioral evidence for causal inference in social science research.
Method using LLM encodings to improve semantic representation of building objects for AI training in construction and engineering.
Framework and reference guide for simulation-based synthetic data generation techniques for training AI agents and subsymbolic models.
Benchmark evaluating LLM economic intuition and decision-making through simulated lemonade stand business management tasks.
Evaluation benchmark with hierarchical task decomposition for assessing LLM capabilities in full-lifecycle educational research workflows.
Circuit analysis framework distinguishing competence from compliance in LLM reasoning under engineering constraints and methodological conventions.
Interpretability framework for multilingual LLMs in Indian languages using shared affine transformations for cross-lingual understanding.
Agentic AI system for autonomous experimental design in particle physics using natural language prompts and Monte Carlo simulation.
Unified agent communication protocol for secure, federated, decentralized autonomous agent-to-agent orchestration across platforms.
Real-time whole-body humanoid teleoperation system with closed-loop global motion tracking to prevent pose drift.
Safety framework for autonomous self-driving laboratories integrating AI with robotic automation for closed-loop experimentation.
Analysis comparing interaction network structures of AI agents versus humans in co-inhabited online platform Moltbook.
CGRA-DeBERTa transformer with LoRA adaptations for question-answering over classical Islamic texts using concept-guided residual augmentation.
Analysis of backdoor attack vulnerabilities in federated learning systems exploiting layer-specific weaknesses during distributed model training.
Large-scale dataset of 100k real-world LLM-based web information extraction events capturing structural web context, released from ScrapeGraphAI telemetry.
Detection method for backdoor attacks in LoRA adapters by analyzing weight space without requiring test data or knowledge of trigger patterns.
Benchmark dataset and framework for evaluating LLM agents' ability to learn tool behavior and improve documentation for opaque real-world tools through interaction.
Colosseum framework for auditing collusive behavior in multi-agent systems where LLM agents communicate and coordinate through free-form language.
Bayesian inference method for learning reward functions from heterogeneous feedback types including demonstrations, comparisons, ratings, and stops.
Approach to automatically detect reward model biases in LLMs using iterative LLM proposals, addressing spurious attributes like length, hallucinations, and sycophancy.
Method to improve adversarial training for LLMs by addressing distribution gaps that leave models vulnerable to simple in-distribution adversarial examples like prompt rewriting.
Cross-stack analysis of generative AI applications in computing systems design from code generation through hardware design space exploration to RTL synthesis.
Framework for prototyping biomechanical HCI tasks using reinforcement learning, improving usability and interpretability for interaction design.
Benchmark suite for imperfect-recall decision problems in game theory, covering scenarios where agents have limited memory and communication.
Study of training long-context vision language models up to 344K tokens, analyzing pretraining, finetuning, and preference optimization with reproducible recipes.
On-policy distillation technique for training student LLMs using teacher supervision on token-level trajectories, improving generalization over off-policy methods with reduced computational cost.
AI literacy framework with learning outcomes to help users build cognitive resistance against disempowerment from AI assistant interactions including reality and value distortion.
Dataset distillation method using exploration-exploitation optimization to compress large-scale data into synthetic datasets while maintaining model performance and reducing training costs.
Framework systematically studying visual preferences and decision factors of vision-language models.
Network management and orchestration platform for federated AI-as-a-Service deployment.
Quantum-inspired classification head using complex-valued unitary transformations for uncertainty quantification.
Network orchestration framework for AI-as-a-Service model selection and execution placement.
Analyzes geometric structure of softmax representations in neural networks for interpretability.
Hybrid federated and split learning framework for privacy-preserving clinical prediction across institutions.
Speculative decoding optimization for video LLMs using text-anchored window attention and visual caching.
Shows random masking in RMSProp optimizer outperforms sophisticated preconditioners for LLM training.
Watermarking scheme for LLMs preventing false attribution with robust signature guarantees.
Pilot-free aggregation primitive for over-the-air federated distillation using energy signaling.
Establishes prescriptive scaling laws for foundation models given compute budgets and downstream accuracy.
Game-theoretic framework for long-tail multi-label classification in large-scale data mining.