Keppsake – automatic save points for your AI projects
Tool for automatic checkpointing of AI projects on macOS, allowing rollback when AI generation goes wrong.
Tool for automatic checkpointing of AI projects on macOS, allowing rollback when AI generation goes wrong.
Testing framework for evaluating AI agent behavior consistency across languages, checking tool selection and arguments in multiple locales.
Analysis of hidden costs in AI agent interactions, calculating human waiting time overhead alongside token costs.
Football management simulation where AI agents control teams through MCP, with terminal-based match commentary and ASCII visualization.
CLI tool suite and Claude Code skills for fixing AI localization bugs, providing format engineering, QA testing, and native-quality translations for 46 languages.
Investigation using hallucination detection tool to uncover fabricated citations in AI-generated government reports, academic papers, and consulting firm research.
Analysis of nearly 10,000 websites showing 97% expose no tools for AI agents to use, highlighting gap between agent-ready interfaces and current web design.
ChatGPT Work agent that takes cross-app actions and manages multi-hour projects autonomously.
R package with Rust core embedding local LLMs, exposing generation, embeddings, activation tracing, steering and ablation as base-R functions for mechanistic interpretability research.
Benchmark results for Grok 4.5 proprietary LLM showing performance metrics, context window (500k tokens), and multimodal capabilities.
Browser extension hiding AI chat history during screen sharing to prevent accidental exposure of conversation data.
Tool converting circuit descriptions and component assets into KiCad projects, enabling AI agents to work on hardware design without parsing complex file formats.
Opinion piece criticizing overuse of AI-generated content in business communications, discussing quality degradation from excessive output.
GLM 5.2 model successfully runs on consumer hardware with capabilities comparable to Claude/GPT while avoiding out-of-memory errors.
Probelock is a lockfile mechanism designed for LLM tool calling, enabling safer function invocation in LLM-based agents.
Former GitHub CEO launches Entire, a Git hosting network optimized for AI agents and code generation workflows.
Greppy extends grep with code-navigation subcommands optimized for AI agents exploring codebases.
Cybersecurity AI (CAI) Dataset released for training and evaluating ML models on security tasks.
Evaluation framework for founders launching AI-built applications; practical considerations for AI product launches.
Analysis of how AI models improve code rewriting economics for popular tech stacks due to training data prevalence and codebase context.
Tensorlake runs Docker images for Harbor on microVM sandboxes, passing Terminal-Bench 2.1 tasks with cold starts in seconds.
Compendium is a shared workspace tool designed for teams and AI agents collaboration.
Guide to removing bloat from Claude Code's system prompt, reducing token payload by tens of thousands per request through context optimization.
Title indicates research on indirect prompt injection attacks against RAG pipelines.
Production-assessed benchmark for code agents evaluating entire trajectory including instruction following, tool use, error recovery, and communication.
Theoretical analysis of in-context search as approximate inference over reasoning traces, studying sampling complexity of reflection-driven LLM reasoning.
Hybrid approach combining LLMs with agent-based modeling to enable real-time adaptive decision-making in large-scale individual interaction simulations.
Uses quantum processor as calibrated belief-update service for sequential POMDP belief updates in autonomous systems under partial observability.
Studies open-weight DeepSeek V3.2 model on ARC-AGI-1 benchmark using strict compute budgets without test-time scaling or fine-tuning.
ReAct-style agent combining LLM reasoning with SageMath CAS for verifiable feedback on research-level computational and experimental mathematics problems.
Analyzes token economics of enterprise agentic AI, arguing orchestration design (harness layer) is key lever against token inflation per task.
Identifies instruction leakage problem in goal-conditioned world models for spatial reasoning and proposes goal-free dynamics fix for true perception grounding.
Large Behavioral Model learns customer decision-making from retail transactions using Person-Environment formulation for grounded behavior modeling.
Studies how AI agents including LLMs learn social norms to improve human-AI coordination in dynamic interactions through implicit shared expectations.
Proposes relative-scale evaluation paradigm where AI models generate adversarial challenges to measure intelligence beyond human-saturating benchmarks.
Analysis of multi-agent LLM safety systems decomposing pipeline effects into three mechanisms: reframing harmful intent, planner refusal, and executor delegation.
ImagingBench: benchmark evaluating vision-language models and agentic AI on 20 computational imaging tasks across optics, signal processing, and inverse problems.
Framework for detecting logical inconsistencies in chain-of-thought reasoning of LLMs without requiring controlled interventions on evaluation transcripts.
Agents optimize static atomic tool actions into reusable Standard Operating Procedures to reduce reasoning overhead and failure rates.
PA-SciML: verification-first workflow for LLM agents discovering surrogate models in scientific ML with physics-based auditing.
MIRA-Math benchmark for mathematical reasoning where solvers must request one missing atomic fact to solve problems.
Agentic Data Environments framework for autonomous agents operating across files, APIs, applications, and system state with failure bounds.
Deterministic gates detect silent policy violations in tool-using LLM agents where forbidden state transitions execute successfully.
Shows biased reward judges silently disable skill retirement in self-evolving LLM agents, causing library drift below baseline performance.
SpaCellAgent: LLM-based multi-agent framework for autonomous trajectory inference analysis in spatial and single-cell transcriptomics.
Pyligent framework for training correction-aware reasoning in LLMs: validated search over partial solution chains with failure recovery.
Compares LLM-generated vs expert-written reusable skills for AI data scientists across data cleaning, SQL, statistical testing, and formatting tasks.
RL post-training enables Transformers to compose primitive skills into higher-level reasoning strategies beyond base model capabilities.
Survey of 1,250 papers on recursive self-improvement in AI systems: self-refinement, self-reward, self-play, and autonomous research loops.
SkillCenter: open-source library with 216,938 structured skills for autonomous AI agents across 24 domain bundles with source grounding.