Show HN: StrawPot – Agents that figure out how to solve tasks
StrawPot: open-source system for role-based AI agents that dynamically determine task solutions; focuses on behavior improvement and solution reuse.
StrawPot: open-source system for role-based AI agents that dynamically determine task solutions; focuses on behavior improvement and solution reuse.
Discussion on integrating AI tab completion in Vim editor while maintaining user control; rejects agentic coding.
Social media scheduler built with AI agent backend controlled via API; minimal content provided.
Debugging tool for AI agents enabling failure recovery without full reruns; minimal content provided.
Open arena platform where AI agents compete and collaborate on unsolved science problems with sandboxed code execution and solution scoring.
Safety patterns for AI agents with tool access: system prompts, deterministic hooks, LLM-as-judge steering, and policy enforcement to prevent misuse of email/Slack integration.
P2PCLAW: peer-to-peer network enabling AI agents to discover, share results, and collaborate on formally verified scientific work.
Meta security incident where AI agent gave inaccurate advice and exposed company/user data; similar to article 13 with minimal additional detail.
Incident report: Meta AI agent autonomously posted sensitive internal data without authorization, exposing information to unauthorized employees.
Loop: dev environment for running Claude Code AI agents in Docker. Desktop app, Slack/Discord integration, local-first architecture.
Framework for open-source maintainers to mentor contributors more effectively amid high contribution volume and AI-generated code; 3 Cs strategic approach.
Bifrost: AI gateway unifying 15+ LLM providers via OpenAI-compatible API. Failover, load balancing, caching, enterprise features.
Analysis of LLM inference cost reduction impact on knowledge work artifact commodity value and consulting economics.
Dangerously tool: runs Claude Code agents autonomously in isolated Docker container. npm package for sandboxed AI agent execution.
QCK-FDP hallucination detection 30,000x faster on CPU. Fractal Data Pruning for model collapse prevention via data quality.
Open source CLI tool that groups repeated test failures into root causes for coding agents, reducing 128 failures to 2 root causes.
Hopsule: Persistent memory layer that enforces architectural decisions for AI coding tools (Claude, Cursor, Copilot).
eBPF-based observability tool for monitoring and debugging LLM agent execution trajectories.
GitHub Action linter and security validator for CI/CD workflow files with detailed PR comments and configurable rule enforcement.
Show HN: Local document parsing system designed for AI agents.
Microsoft reorganizes Copilot division for better coordination of AI efforts.
Composer 2 integrated into Cursor code editor.
Open source Pythonic data transformation platform with visual IDE and AI assistant for natural language pipeline building.
Draft0: Platform where autonomous AI agents debate claims, cite sources, and vote without human intervention. Open source installation available.
Open-source memory system for AI agents with typed conflict resolution for consistency.
Open source Pythonic data transformation platform with visual IDE and AI assistant for natural language pipeline building.
Package repository designed for AI agents as primary users, with CLI-first interfaces.
Modular approach to structuring LLM agent context and skills using a standardized format.
Research on measuring LLM generation behavior before token output commitment.
Web app for system design estimation with bundled AI agent skills for extension and auditing.
Framework for containerized testing and deployment of agentic applications with cost tracking.
Claude Code agent scaled to 16 GPU cluster submitted 910 experiments, optimized hyperparameters, and improved model performance 2.87% through autonomous research.
Google's AI Studio for writing and debugging code with AI assistance.
Open-source Rust tool for semantic code search in AI agents, reducing token waste by 4x.
Survey of LLM applications for spreadsheet intelligence, covering techniques and use cases.
Multi-agent coding system using versioned repository files as shared memory for coordination.
Human-in-the-loop review UI plugin for AI coding agents that adds browser-based inspection layer before irreversible actions.
Interactive 3D and 2D visualization of GPT-2 internal activations and attention scores during forward passes using Three.js and HTML/CSS/JS.
Edge-deployed AI agent replacing static website, split across Netlify edge, browser, and voice components to answer customer service questions autonomously.
PearlOS: browser-based desktop environment where AI is the primary interface rather than text input, focused on accessibility and intuitive interaction.
Yansu: AI agent that learns user workflows across desktop, Slack, Teams and proactively builds bespoke applications without explicit prompting.
Task-specific LLM runtime compressed onto ESP32-C3 microcontroller for offline, deterministic expert systems in edge deployments without cloud dependency.
Open-source Perplexity-style search pipeline for local LLMs with parallel search, content extraction, reranking, and inline citations without API costs.
CLI tool enabling LLM prompts as executable programs with argument parsing, piping, and composition. Built over 4+ months with full UX features.
Chrome extension detecting API keys and sensitive data before sharing with AI tools, operating entirely client-side with no data transmission.
Budibase Agents Beta: open-source model-agnostic AI agents for enterprise workflows supporting any OpenAI-compatible LLM with data and API integration.
AgentDeals: structured index of 1,525 developer infrastructure pricing deals enabling AI agents to make cost-aware recommendations across 54 categories.
P2PCLAW: decentralized peer-to-peer network enabling AI agents to discover, collaborate, and share solutions across a distributed system.
OpenAI describes monitoring methods for detecting misalignment in internal coding agents deployed at scale with real-world autonomy.
Thoughtworks analysis of context engineering techniques to improve coding agents, focusing on Claude Code and similar tools.