HumanCompiler – Compile humans into AI agents – a Claude Code plugin
Claude Code plugin that converts human behavior into installable AI agent plugins through interviews and artifact analysis.
Claude Code plugin that converts human behavior into installable AI agent plugins through interviews and artifact analysis.
Overview of AI agents becoming mainstream, discussing tools like Claude Code and their ability to automate weeks of work into hours.
API-first puzzle arena and leaderboard for testing AI agents across CAPTCHAs, geolocation, lateral thinking, and research problems.
Construct-and-Refine (CaR) method for efficient constraint handling in neural solvers for complex routing problems with hard constraints.
Investigates optimization instability in autonomous agentic workflows using Pythia framework, characterizing failure modes in clinical symptom detection systems.
Benchmark measuring uncertainty metrics for LLM-based automatic assessment in education, addressing probabilistic output reliability challenges.
Evaluates Mirror, an evidence-grounded clinical reasoning system, against GPT-5 variants on endocrinology board exams with nuanced subspecialty knowledge.
Framework for improving LLMs' ability to adapt dynamically using interactive natural language feedback in collaborative settings, beyond static training paradigms.
GPSBench benchmark evaluates LLM geospatial reasoning with 57,800 samples across 17 tasks for navigation, robotics, and mapping applications.
PAHF framework enables AI agents to learn and adapt to individual user preferences dynamically through human feedback, addressing personalization and preference drift.
EnterpriseGym introduces CoreCraft, a high-fidelity RL environment with 2,500+ entities for training generalizable AI agents on enterprise customer support simulations.
Design concepts for memory systems essential for artificial superintelligence, exploring high-capacity storage approaches.
Proxy State-Based Evaluation method for benchmarking multi-turn tool-calling LLM agents with LLM-driven simulation backends.
In-context co-player inference enabling multi-agent cooperation through learning-aware agents in reinforcement learning.
Certification protocol for multi-agent communication ensuring semantic consistency between agents using stimulus-meaning model.
CAFE framework combining causal discovery with multi-agent RL for automated feature engineering on tabular data.
Causal discovery using LLMs with constraint-based argumentation framework to uncover causal relations from data.
Framework of Thoughts unifying Chain, Tree, and Graph of Thought prompting with dynamic adaptation and hyperparameter optimization.
Seven-month poetry workshop shaping LLM into digital poet through iterative in-context expert feedback without retraining.
Investigation of Agent Skill framework effectiveness with small language models in industrial environments for context engineering and hallucination reduction.
Framework for evaluating AI agent reliability beyond accuracy metrics, examining consistency, perturbation robustness, and operational flaws.
PICQ dataset and method for identifying unknown relevant personas in user simulation for dialogue systems.
EdgeNav-QE framework combining QLoRA quantization and dynamic early-exit for deploying Large Action Models on edge devices.
Validation of code compression tolerance in LLM prompts across six benchmarks, explaining why code compresses better than mathematical reasoning.
LLM representations for few-shot tabular classification on heterogeneous web data structures.
Analysis of geometric relationships between Big Five personality steering vectors in LLaMA and Mistral models.
Validation study comparing Big Five personality scores from LLM conversations against IPIP-50 questionnaire with moderate convergent validity.
IntelliReward model using preference optimization to improve review question generation quality beyond surface-level queries.
Survey of narrative theory-driven LLM methods for story generation and understanding, proposing taxonomy of narrative datasets and theories.
Clinical NLP models addressing temporal and lexical leakage in hospital discharge planning to prevent overconfident predictions.
Multi-task learning method for classifying unsafe prompts with explainability, trained on synthetic data to mitigate LLM confirmation biases.
GOPO: hierarchical RL framework for task-oriented dialogue decoupling strategy planning from response generation using preference optimization.
Query-conditioned soft compression approach for RAG reducing context length and redundancy while maintaining performance.
Systematic study of how state representation (granularity, structure, language) affects LLM dynamic reasoning in interactive environments.
CAST: consistency framework for stable LLM-based text analysis on tabular data via algorithmic prompting and token selection.
Semantically grounded recipe generation from food images using MLLMs with action/ingredient validation framework.
Study shows LLM reasoning improves via self-generated examples through process enhancement rather than example quality.
Reviews human-AI interaction literature distinguishing AI as tool vs. teammate across interaction design, trust, and healthcare.
NLP-PRISM: framework surveying privacy risks from NLP processing of social media containing PII and behavioral metadata.
Evaluates problem-solving and reasoning capabilities of contemporary LLMs through performance in Zork text adventure game.
Test-time adaptation methods for tactile-vision-language models handling modality-wise reliability under distribution shifts.
Fly0 framework decouples semantic reasoning (MLLM) from geometric planning for zero-shot aerial navigation tasks.
FUTURE-VLA: vision-language architecture for real-time robotic control combining long-horizon trajectory forecasting with reduced latency.
Study reveals GPT-4o exhibits daily and weekly performance variability, threatening reproducibility of LLM-based research.
FlipSet benchmark evaluates visual perspective taking in 103 vision-language models through 180-degree character rotation tasks.
CogitoRAG: RAG framework inspired by human episodic memory to improve semantic integrity and reduce retrieval deviations in LLM-based systems.
Doc-to-LoRA method distills long document contexts into LoRA parameters to reduce transformer attention cost in LLM inference.
Resp-Agent autonomous system with Active Adversarial Curriculum Agent for respiratory sound generation and disease diagnosis.
MaS-VQA framework with mask-and-select approach for knowledge-based visual question answering with noisy external knowledge.
EarthSpatialBench benchmarks spatial reasoning capabilities of multimodal LLMs on georeferenced Earth imagery for embodied AI systems.