Show HN: I ported llama.cpp to Apple Watch and ran a 0.8B LLM locally
Port of llama.cpp to Apple Watch enabling local 0.8B LLM inference on wearable hardware.
Port of llama.cpp to Apple Watch enabling local 0.8B LLM inference on wearable hardware.
Framework for controlling AI agent behavior using session state and contextual policies. State-based agent governance system.
Declarative specification system for AI agents using Terraform-style syntax. Infrastructure-as-code approach for agent configuration.
Provider-agnostic agent skills router for LLMs. Dynamically injects relevant skills into prompts based on context, avoiding monolithic prompt issues.
Tool to audit website content readability and accessibility for AI agents.
AI agent architecture using NATS message bus for reactive node communication instead of conversational loops.
Open-source self-hosted investment research platform using bring-your-own-key LLM integration.
Tool generating LLM-readable codebase documentation with grep-friendly tags for better code understanding.
Tool enabling multiple AI coding agents to operate concurrently within a single Git repository.
Markdown editor visualizing exact token representation LLMs see. Shows formatting, whitespace, warnings for prompt/doc optimization. Open-source, local-first.
Shadcn form builder that uses AI to generate React forms from visual specifications.
PACT: open-source toolkit for signing digital content, tracking provenance, and enforcing AI training policies with policy metadata.
Engramma Memory: open-source composable memory architecture using multi-head attention for AI agents.
Vicinae command palette with Raycast extension compatibility now available on macOS, built with Qt/C++ and Node.js runtime.
Analytics tool for agents to optimize costs when using Claude Code and other LLM agents against expensive data platforms. Addresses token efficiency.
Report on security vulnerabilities found in Anthropic Claude Code. LLM system analysis.
Grillr: AI agent that critiques startup ideas and enforces accountability with real deadline tracking.
Video exploring J-Space theory explaining how AI models work internally.
NexSub: offline AI video subtitle translator supporting multilingual translation locally without internet or subscriptions.
UIPrompt: visual component editor generating spec-grade AI prompts for Claude, Cursor, and v0 with exact design values and accessibility rules.
Meta releases MuseImage and MuseVideo generative models for image and video creation.
Analysis of LLM capabilities for document extraction tasks. Evaluation of practical viability.
Community discussion on production AI agent architectures accessing databases. Requests real implementation learnings on guardrails and problems.
Research on LLM-based tree editing capabilities across multiple studies. Empirical analysis.
Discussion about collaborative prompt engineering for AI agents on teams. Questions why prompts remain individual rather than shared like code.
User deployed 25 AI agents to critique startup ideas; 22 were killed by agent feedback. Explores AI agent capabilities for evaluation.
AIfunc library enables calling AI as typed, testable functions across languages without learning new frameworks. Model-agnostic npm package approach.
OpenAI audits SWE-Bench Pro benchmark, finds ~30% of tasks broken; details importance of accurate model evaluation.
eBPF-based test coverage measurement tool without code instrumentation requirements.
Agent skill module enabling AI coding agents to generate UML diagrams from natural language or existing codebases.
GLM-5.2 max model performance comparison with Claude Opus 4.8 on Harvey benchmark. Model evaluation result.
Cinchor tool provides control and auditability for AI agent actions. Enables constraining agent capabilities and proving execution history.
Technical documentation on agentic memory systems in WunderOS. Discusses perspective fusion and trust in data sources for distributed systems.
Title only. AI embeddings cost reduction case study. Insufficient technical depth provided.
Title only. China security warnings about Claude Code tool. Policy/news without technical details.
Open-source benchmark measuring AI agent memory quality and decision rejection awareness beyond retrieval accuracy. Reproducible evaluation.
Developer tool for mapping and governing multi-repo architectures using LLMs and AI agents to handle microservices codebase complexity.
MCP server for Claude/Cursor to control Ultralytics YOLO training, datasets, and model management via AI agents. Community project enabling agentic ML workflows.
Comparison of API pricing limits across Claude, Codex, and Copilot coding assistants.
Title only. Concept for intervention mechanism to prevent AI agent errors before execution.
macOS application monitoring and recording AI coding agent behavior and actions for analysis.
Quicopt optimization solver service with Python API supporting OR-Tools and Pyomo models. Useful tool but not directly AI-focused.
Title only. Discussion of agentic test processes and LLM benchmarks for code generation.
OpenClaw plugin enabling AI agents to make/receive real phone calls via Twilio and OpenAI Realtime API with natural voice conversations and task completion.
Nonprofit AI accelerator ecosystem connecting 10,000+ AI builders with compute, research partnerships, and hackathon ($50k prize).
Research study testing open-source LLMs in Milgram-style obedience experiments, examining alignment and safety of autonomous LLM behavior.
Security research on prompt injection attacks (HalluSquatting) that enable LLMs to assemble botnets by exploiting inability to refuse malicious commands.
Insufficient content; tool description for recovering context about AI-written code from Git history.
Developer documents experience using OpenAI Codex AI agent for major refactoring task, analyzing capabilities and limitations in real-world code changes.
LingBot-VLA 2.0 is an open-weight Vision-Language-Action foundation model for robot control across 20 embodiments, with improved real-world deployment capabilities and code/checkpoints available.