Show HN: Retrace – reverse debugging for production CPython applications
Reverse debugger for production Python apps that records failing executions and replays them locally with deterministic stepping for debugging ML models and external calls.
Reverse debugger for production Python apps that records failing executions and replays them locally with deterministic stepping for debugging ML models and external calls.
Discussion of tool-response engineering as advancement beyond prompt engineering for LLM application development.
Gateway tool converting REST/SOAP/SQL interfaces into MCP protocol for integrating diverse data sources with Claude agents.
Cross-platform MCP server enabling Claude to see Windows/Linux screens with OCR and vision-diff capabilities, filling gap in Anthropic's macOS-only official implementation.
Overview of observability tools for monitoring and debugging LLM applications in production environments.
Title-only post about AI agents discovering reasoning strategies that reduce LLM token usage by 70%.
System prompt implementing strategic reasoning framework for LLMs based on Hammerstein-Equord decision framework.
Terrably framework for building Terraform providers in TypeScript with full type safety, compiling to self-contained binaries.
Claude Code plugin orchestrating multi-agent SDLC team with specialized AI agents for automated task planning and dispatch.
Browser-based lockfile scanner for TanStack NPM supply-chain incident detection without network transmission.
DeepClause: open-source agent harness using Prolog and WASM. Compiles task descriptions into executable logic programs with tool orchestration and TUI interface.
Technical analysis of SQLite as optimal database for AI agents. Author building Willow agent harness, discusses workload characteristics for agent systems.
Title-only post about local memory system for AI agents with 98% recall on benchmarks.
Token compression technique for agentic AI systems implemented in Haskell. Optimization for agent efficiency.
Prave: management platform for AI Agent Skills. Infrastructure for organizing and deploying agent capabilities.
Pi-treebase: interactive session history management tool for LLMs with rebasing and summarization.
Cplt: developer tool to run AI coding agents in kernel-level sandbox for isolation. Enables safe agent execution.
Empirical analysis of frontier AI agent task-completion time horizons across 100+ software tasks. GitHub code and raw data provided.
MCP tool enabling natural-language access to Oura health ring data through Claude with local SQLite storage and annotations.
Kiji Proxy: open-source privacy layer for AI APIs. Automatically detects and masks PII in requests to LLM services. Built by Dataiku's open source office.
Research on recursive self-improvement in AI systems, covering historical context and current emergence of RSI in practice.
Claude Code plugin marketplace from Trail of Bits providing security analysis skills for AI-assisted testing and development workflows.
Privatemode.ai: LLM service using confidential computing for end-to-end encrypted data processing. Privacy-focused AI platform.
LLM benchmark evaluating meme generation from current news content.
RipStop is a Node.js package implementing guardrails to protect repositories from unintended LLM agent actions via Git rule enforcement.
Microsoft research findings show frontier AI models and agents accumulate errors in long-running task workflows, limiting practical automation applications.
Personal reflection on AI IDE UX design patterns, emphasizing chat-based creative workflows and AI-assisted development.
Nobel economist Daron Acemoglu discusses cautious AI predictions, focusing on job displacement concerns beyond hype.
PDF and e-signature API designed for AI agents, now available on Cursor Directory.
Discussion of Claude Code Max subscription usage limits and proposal for parallel session execution.
Open-source macOS tool providing ambient radio interface for Claude Code and Codex agents, surfacing progress and blockers in real-time.
Wix conducted 250 AI agent evaluations comparing curated skills vs. raw documentation for agent performance.
Technical report on governance requirements for AI agents performing real work: ownership, authorization, review, replay, and improvement capabilities.
Vision-in-the-loop optimization for LaTeX document typesetting using iterative compile-inspect-edit cycles with visual feedback.
Multi-agent test-time scaling approach organizing parallel reasoning trajectories with structured coordination to improve LLM reasoning ability.
Study of world models for mobile GUI agents, comparing text vs image-based predictions of action consequences for long-horizon task execution.
Benchmark for evaluating value alignment in autonomous agents, showing agent values diverge from LLM values with implications for safety.
Graph reasoning agents that reconstruct structured graphs from text and coordinate instruction-following with tool usage through structural credit assignment mechanisms.
Research on Autonomous FAIR Digital Objects enabling active knowledge validation and autonomous curation on the web, replacing passive assertions with agent-based stewardship.
Agent-X framework for accelerating on-device LLM agents through prompt rewriting for prefix caching and LLM-free speculative decoding.
Empirical study benchmarking agentic AI performance on edge devices constrained to 8B parameters, measuring quality degradation under hardware constraints.
Research on safeguarding autonomous driving MLLMs using Markovian safety logic for temporal reasoning in dynamic traffic scenarios.
Research using LLMs to discover efficient branching policies for Mixed Integer Linear Programming solvers, reducing dependence on expert demonstrations.
Research examining reliability gaps in interactive agent benchmarks, showing surface-level outcome checks fail to verify actual agent behavior and task completion.
Research on ASIA, an autonomous agent for system identification that automates model selection, training algorithms, and hyperparameter tuning in dynamical systems learning.
Research on SkillEvolver, a meta-skill framework for online learning and iterative refinement of agent skills from real deployment without static, hand-authored artifacts.
Research on improving LLM structural understanding of graphs by sharpening internal attention mechanisms without external adapters or fine-tuning.
Research establishing statistical methods and U-statistics framework for measuring AI agent reliability and consistency under semantic perturbations and trajectory-level stability.
Research introducing benchmark for continual learning on evolving biomedical knowledge graphs, addressing real-world asynchronous KG updates beyond synthetic splits.
Research on LLM-based storytelling agent for older adults integrating knowledge graphs, argumentation theory, and argument mining to reduce hallucinations and improve transparency.