Show HN: EloPhanto – A self-evolving AI agent that builds its own tools
EloPhanto self-evolving local AI agent that controls Chrome browser with 47 tools. Autonomously writes, tests, and integrates new Python tools.
EloPhanto self-evolving local AI agent that controls Chrome browser with 47 tools. Autonomously writes, tests, and integrates new Python tools.
A2SPA cryptographic tool for signing and verifying AI agent payloads. Security infrastructure for agent communication.
TinySDLC agent orchestrator adds SDLC discipline to multi-agent AI coding. Eight roles with separation of duties, isolated workspaces, enforced handoffs.
Vram.run is a comparison tool for API providers, local GPUs, and cloud options across different LLM models.
Kwin-MCP server enables AI agents to automate Linux GUI via KWin. MCP protocol for AI-driven desktop automation.
Vibevideo unified interface for multiple AI video generation models. Aggregates text-to-video, frame-guided, and reference-based generation tools.
mindpm MCP server adds persistent memory to AI coding assistants across Claude Code, Cursor, and others. Stores tasks and decisions in SQLite.
Inconvo agent builder enables chat-with-data without LLM SQL generation. Validates structured intents against semantic layer before execution.
Open-source CLI tool for Actual Budget optimized for AI agents like Claude Code, enabling programmatic budget management while maintaining web dashboard access.
Tutorial on using AI agents to automate reading Jira tickets and generating pull requests.
Guide to fine-tuning LLMs for enterprise applications, covering mechanics of adapting models like Qwen 3 and DeepSeek v3 for domain-specific use cases.
Conceptual article on requesting tool recommendations from LLMs instead of direct answers.
toktrack CLI tool monitors token spending across Claude, Codex, and Gemini. Rust-based cost tracking with usage analytics.
Lightweight 15MB Markdown file viewer built in Rust, designed for reading AI-generated documentation.
Tickr Slack bot for AI-driven project management. Automates task tracking and team nudges as alternative to Jira.
CLI tool that translates natural language commands into shell commands using LLMs, enabling hands-free terminal interaction.
CLI tool designed as curl alternative for AI agents to make HTTP requests.
BasaltSurge payment API designed for AI agents to perform commerce transactions, bypassing legacy card rails with standardized checkout layer.
Bloomberg Terminal redesign integrating agentic AI capabilities for financial analysis and decision-making workflows.
Hardware and software safety standard for AI-controlled robots with dedicated safety processor on independent power rail controlling AI processor power access.
NIST launches standards initiative for AI agents to establish consistency and safety protocols in agent development.
Testing framework for AI agents with 8-layer graduated assertions covering tool calls, cost budgets, schemas, and output validation without relying on LLM judges.
Runtime permission enforcer for AI agents that prevents file deletion and tool misuse through code-level policies outside LLM context.
Terminal multiplexer built with Rust and egui designed for LLM interaction and management.
Multiplayer game where LLMs control light-cycles via MCP protocol on a grid, testing agent reasoning and decision-making capabilities.
OpenAI-compatible proxy that reduces LLM token costs 40-60% through deterministic rule-based prompt compression with ~5ms overhead.
Production crisis detection system using AI to identify high-risk distress signals with cryptographic audit trails and formal safety guarantees.
Managed hosting service for OpenClaw open-source AI agent framework with 60-second deployment.
SkillScan is a free API detecting malicious patterns in AI agent skill files. Identifies exfiltration services, env reads, API key theft, and prompt injection attempts.
Analysis of how AI coding agents require new development processes and workflows beyond traditional productivity measurement.
Interactive timeline tracking 171 LLMs from Transformer (2017) to GPT-5.3 (2026). Filterable by open/closed source across 54 organizations.
OpenAI's Frontier platform for building, deploying, and managing AI coworkers that automate enterprise workflows end-to-end.
MALLVI: Multi-agent framework combining LLMs and vision for closed-loop robotic manipulation with environmental feedback, enabling dynamic task planning.
Wink: Framework for detecting and recovering from misbehaviors in LLM-powered coding agents, addressing issues like instruction deviation, infinite loops, and tool misuse.
Study reducing textual bias in synthetic MCQA benchmarks for vision-language models to prevent exploitation of linguistic patterns in autonomous driving.
BioBridge: Domain-adaptive framework bridging protein language models with LLMs for enhanced biological reasoning and protein interpretation.
LATMiX: Learnable affine transformations for microscaling quantization of LLMs, improving robustness by reducing activation outliers.
CodeScaler: Execution-free reward model scaling code LLM training and inference via verifiable rewards without test case dependencies.
Curriculum learning framework for distilling chain-of-thought reasoning into compact student models via structure-aware masking and GRPO training.
AnCoder: Code generation via discrete diffusion models using AnchorTree framework to maintain programming language structure and executability.
Robust-MMR: Multi-modal pre-training approach for medical vision-language models with domain-invariant masked reconstruction for robustness.
HELIX: Geometric framework decoupling entropy from hallucination in quantized LLMs by steering hidden states to truthfulness manifold.
Agentic unlearning framework removing sensitive information from both LLM parameters and agent memory to prevent information reactivation.
Case study comparing PTQ methods (AWQ, GPTQ, SmoothQuant, FlatQuant) for reasoning LLMs on Ascend NPU hardware.
AsynDBT: Asynchronous distributed optimization for in-context learning with cloud-based LLM APIs to improve prompt engineering efficiency.
EXACT: Decoding-time personalization for LLMs using explicit attribute-guided adaption to account for context-dependent user preferences.
Systematic evaluation of safety region identification methods across LLM families, testing parameter constraints for controlling model safety behaviors.
Framework leveraging variability modeling to optimize LLM inference hyperparameters for energy efficiency and computational sustainability.
ScaleBITS: Mixed-precision quantization method for LLMs using principled bitwidth search to reduce memory and inference cost below 4-bit average.
Framework for certifying ML model risk bounds under distribution shift with verifiable constraints and computable metrics.