Indxr v0.4.0 – Teach your agents to learn from their mistakes
Fast codebase indexer providing persistent wiki for AI agents to learn from mistakes. Open source tool for agent knowledge management.
Fast codebase indexer providing persistent wiki for AI agents to learn from mistakes. Open source tool for agent knowledge management.
Discussion of user interface patterns and design considerations for agentic SaaS applications.
Chrome DevTools Protocol-based JavaScript runtime instrumentation tool for debugging and execution interaction.
RBF-Attention replaces dot-product attention with Euclidean distance calculations in transformer architectures.
Drop-in PII redaction proxy for LLM APIs with zero-width Unicode normalization layer to bypass evasion techniques.
SQLite Memory: Markdown-based memory system for AI agents with offline-first synchronization capabilities.
Show HN: Distillery MCP server providing persistent shared team context for AI coding sessions. 50k lines Python built in one week, captures design decisions for reuse.
Minimal content about using LLMs for research discoveries. No technical details provided.
AI agent skill that converts Figma designs into Android/iOS mobile application code.
Multi-agent orchestration tool using bash, Docker, and git coordination with terminal dashboard.
Device-bound ECDSA authentication replacing static API keys for M2M communication without .env files.
Open source agent OS providing persistent memory, loop detection, audit trails, and real-time observability.
25.6M parameter Rust-focused language model with byte-level training exploring hybrid attention mechanisms.
Linting tool for coding agents that detects context drift and validates AGENTS.md alignment with codebase.
Research on behavioral drift observed in long-running autonomous AI agents over 100 days. Empirical findings on agent behavior changes.
CLI tool providing structured SEO data to AI agents for any URL. Limited technical information provided.
Open-source GitHub action for auditing and fixing AI-generated UI code in CI/CD pipelines. Uses browser testing to verify generated code and autonomously pushes fixes.
Open source voice agent that builds persistent knowledge graphs from recordings with offline and cloud STT options.
Analysis of extended thinking in Claude across 17,871 thinking blocks and 234,760 tool calls in coding workflows.
Security-focused development environment for using Claude safely with devcontainers. Addresses issues like file removal and email access by treating humans as firewall.
MCP tool integrating job search database with LLM for natural language queries on software engineering positions.
Natural language interface using LLM to generate robot programs with MuJoCo browser demo and open code.
Multi-agent security testing framework using specialized agents and tools. Finds vulnerabilities through chaining attacks, supports multiple LLM backends, achieves 96% recall on OWASP tests.
Open source framework to deploy agent skills as APIs with multi-model support and stateful execution.
Open-source identity and credential system for autonomous agents. Provides short-lived credentials, delegation, attestation, and revocation using OAuth 2.1 and SPIFFE standards.
Analysis of security vulnerabilities in LLM guardrails and prompt injection attacks with SQL injection analogy.
CLI tool to manage Claude sessions, tasks, and git worktrees with agent integration.
AvatarBook: proof and settlement layer for verifiable AI agent workflows with cryptographic signing and task delegation.
macOS sandbox wrapper for AI and coding tools using Apple's sandbox-exec. Allow-first approach with gitignore-like configuration for limiting tool access.
Definition and framework for Open Source AI (OSAID), establishing essential freedoms for AI developers similar to traditional open source software principles.
Email service designed for AI agents to sign up and use independently. Solves problem of agents spending tokens on traditional signup flows with proof-of-work verification.
Platform that generates hardware designs, code, and shopping lists from text descriptions using AI. No technical depth provided.
Open-source alternative to Higgsfield AI offering image/video generation via 200+ models without subscriptions or closed ecosystem.
Security researcher extracted Telegram's AI text rewriting feature system prompt via prompt injection, finding the model rewrites politically sensitive content contrary to its design instructions.
USC researchers argue LLMs are standardizing human expression and language patterns, potentially reducing cognitive diversity.
MemPalace claims perfect scores on LongMemEval benchmark but actual Benchmarks.md file shows discrepancies; credibility concerns raised.
Mailmap-checker is a pre-commit Git hook that detects unmapped identities by comparing .mailmap against commit history.
Procurement.txt is an open specification for plain-text files declaring pricing, ordering methods, and capabilities for AI purchasing agents.
Mobile app enabling offline LLM inference with Gemma and Hugging Face models on iPad, featuring private on-device chatting and model integration.
TinyProgrammer is a Raspberry Pi device powered by LLM that autonomously writes, runs, and debugs Python programs with a retro Mac IDE interface.
Experimental study on how AI systems cite and validate website content in zero-click search environments, examining citation authority without human-visible content.
Analysis of how cheap LLM tokens mask increasing code complexity and technical debt in AI-assisted development workflows.
Tutorial on safely running autonomous coding agents locally using Docker sandboxes to isolate potentially dangerous operations.
Meta-Harness optimizes AI agent harnesses end-to-end through automated search, improving performance from 28.5% to 46.5% on hard task subsets.
Design system framework providing rules for AI coding tools to generate professional UI components, integrated with Claude Code.
Technical analysis of LLM sampling mechanisms. Explains token generation, temperature, and practical differences between model and inference.
MemPalace: AI memory system storing complete conversation history and making it searchable. Addresses context loss in sessions.
Willitrun: CLI tool checking ML model compatibility with devices using benchmarks. Predicts if models fit and run at acceptable speed.
Zero Human Company: Single-binary Go tool managing AI agents with budget enforcement and execution monitoring. AI-native org dashboard.
CricketBrain: Neuromorphic signal processor in Rust with sub-microsecond pattern recognition. Bio-inspired edge AI with minimal memory.