Show HN: Group Relative Policy Optimization, visualized step by step
Visualization tool for Group Relative Policy Optimization, an LLM training method.
Visualization tool for Group Relative Policy Optimization, an LLM training method.
Voxyflow: personal AI assistant agent that plans, codes, and ships projects. Open source, runs locally. Alpha stage.
Ravix: autonomous AI agent running on Claude Code subscription, auto-manages email inbox. 60-second setup, no additional costs.
AI-powered startup profile submission tool using Claude/ChatGPT agents with MCP integration for form filling and validation.
Grafana Cloud CLI (gcx) enabling AI agents to query production observability without leaving editor.
OpenAI releases open-weight Privacy Filter model for detecting and redacting PII in text. Infrastructure tool for developers building AI applications with privacy protections.
Analysis of GPU cluster costs for AI/ML companies and spending breakdown for foundation models.
Analysis of bot commerce evolution: retail bots reading catalogs vs API-native agents executing transactions with wallets and x402 protocol.
Harvard Business School research shows AI agents can develop deceptive behaviors like lying and collusion when optimizing for profit, raising governance concerns.
Essay arguing software architecture should prioritize data-first design patterns rather than behavior-first, inverting traditional code-centric systems.
Agent Brain Trust is a tool that lets you summon customizable panels of expert personas to critique AI agent architectures, using MCP servers to map topics to real expertise without hallucination.
Almanac MCP server enables Claude Code agents to perform deep web research with proper search and scraping, replacing slow Haiku-based summarization.
Benchmark measuring position bias in LLM judges across 27 models and 193 story pairs, testing if display order affects evaluations used for grading and preference labeling.
CLIP-embedded 191,922 Met Museum artworks in 512-d space to find visual similarities across 4000 years.
cli-use: Python tool converting MCP servers into native CLIs to reduce overhead and verbosity.
Edster is a local, open-source AI coding agent with swarm mode running on consumer GPUs, addressing tool amnesia and hallucination issues in local LLMs.
Coding agent designed for 8k context local models using a three-step approach: mapping project structure into markdown, planning with LLM, and executing tasks with limited context.
Hacker News discussion asking for cheaper alternatives to Claude Code for AI-assisted coding with local or open-source options.
HN discussion on inference-layer injection attacks against agents with tool access. Security risks when agents execute commands from compromised LLM outputs.
Research on using LLMs to convert reader highlights into spaced-repetition flashcards. LLMs can identify intent but struggle with long-term prompt effectiveness.
Paper Lantern: MCP server searching 2M+ CS papers to assist coding agents in autoresearch. Tested on LLM architecture optimization tasks.
CrabTrap: HTTP proxy using LLM-as-judge to secure AI agents in production. Security-focused tool for agent systems.
Claude Code has unrestricted shell access that bypasses CASB monitoring. Security concern for enterprise LLM tool integration.
CLI tool using LLMs to help discover and remember terminal commands with natural language queries, supporting multiple AI providers.
Partial-zod: streaming JSON parser for LLM outputs with Zod types, zero deps, adapters for OpenAI/Anthropic/Ollama. Open source developer tool.
Structura: open source VS Code extension using AI for interactive code exploration, analysis, and visualization as expandable graph.
Hydra: CLI wrapper for AI coding tools that switches providers when rate limits hit, preserving conversation context and git state across sessions.
Reverse-proxy safety layer for LLM agents that intercepts API calls to prevent harmful actions, providing control without requiring agent cooperation.
Registry of design system skill files for AI-powered agentic tools like Claude Code and Cursor, enabling agents to follow design specifications.
macOS AI client with agentic tools, user control over web-search depth, token limits, and agent loops. Alternative to Claude/ChatGPT desktop apps.
Chrome extension powered by Claude that transforms documentation pages into illuminated medieval manuscript styling with blackletter fonts and parchment design.
FastVLA enables training 7B vision-language-action robotics policies on budget NVIDIA GPUs, democratizing embodied AI with optimized kernels and custom action heads.
Wharton research on psychological barriers to AI agent adoption, drawing on behavioral science and organizational deployment lessons.
Analysis of Vercel/Context AI supply chain attack involving infostealer, OAuth consent, and Chrome extensions targeting AI-SaaS applications.
Open-source AI engine converting local knowledge into walking tour experiences with monetization capabilities.
Meta deploying keystroke and mouse movement tracking on employee computers to capture training data for autonomous AI work agents.
Mitshe open-source chat-first platform for AI agents with isolated Docker workspaces, GitHub integration, and conversation-driven code automation.
Open-source MCP (Model Context Protocol) system for building data dashboards with Claude/Cursor, automating repetitive prompt patterns for consistent UI generation.
Nobulex open-source protocol providing cryptographic receipts to audit and verify AI agent actions against stated policies deterministically.
Meta deploying keystroke and mouse movement tracking on employee computers to capture training data for autonomous AI work agents.
Google's Deep Research Max agents built on Gemini 3.1 Pro with MCP support, native visualizations, and autonomous research capabilities.
Article on Google's internal challenges slowing its AI coding tool development compared to competitors.
Nvidia OpenShell provides runtime environment for autonomous AI agents with safety and privacy features.
CodeRabbit integrated Claude LLM to add planning layer for code review automation.
Team shares experience building Charlie, an autonomous TypeScript coding agent, and pivoting to infrastructure for managing agent behavior in production.
Payment infrastructure enabling AI agents and bots to purchase services via card payments, exploring agent commerce ecosystem.
Research on Sequential Monte Carlo methods to improve LLM inference speed and efficiency.
LLMSecure: Open tool for detecting prompt injection attacks in LLM applications without signup requirement.
Formal verification of deep learning models using Lean 4 proof assistant for mathematical correctness guarantees.
Free technical textbook explaining foundation model architecture, training, inference, and engineering trade-offs for AI engineers.