Feasio – AI that gives brutally honest feasibility reports on business ideas
AI application providing feasibility assessments for business ideas.
AI application providing feasibility assessments for business ideas.
CUDA-native AI guardrail kernel written in C++20, 100% branchless implementation for LLM safety.
Open-source handbook on agent harnesses from Tencent, covering auditability and editability of coding agents. Original documentation on AI agent infrastructure.
9lives is a self-healing test runner that automatically fixes broken Playwright tests by classifying failures and healing them in tiers, designed to prevent AI agents from unnecessarily rewriting tests.
YC S24 company shipping Scribe (style-learning dictation) and free AI dictation using fine-tuned Llama 3.1 8b.
Technical retrospective on generative spatial AI advances: text-to-mesh, video, 3D models, world generators, CAD integration.
Title only. Discusses LLM failure modes involving delusional/circular reasoning in chatbots.
Show HN: GenUI enables AI agents to generate real, interactive SwiftUI interfaces for iOS/macOS instead of text output.
Research on memory requirements for LLMs beyond weight storage.
Show HN: NoMac builds, tests, and ships iOS apps via cloud Macs without local hardware. Agent-first with TestFlight/App Store integration.
Title only. Personal experience improving AI agent performance without scaling model size.
Research generalizing language models to probabilistic language tries as alternative architecture.
Google releases Litert.js, a JavaScript library for efficient AI inference in web browsers.
Design patterns and best practices for building Model Context Protocol tools for AI agents.
Research on LLM internal reasoning in latent space without generating intermediate text outputs.
Performance optimization for sandbox infrastructure using kTLS and zero-copy techniques to support agent workloads.
AI agent that automatically deploys applications across multiple platforms (GitHub, DigitalOcean, Vercel, Render, Stripe).
Tool to optimize LLM token usage by reducing unnecessary token sends, focusing on cost reduction strategies.
OpenAI discontinued Atlas AI browser after 11 months, integrating browser agent features into ChatGPT Work platform.
Design patterns and architecture lessons from Muse Spark 1.1 enterprise AI agent deployment.
Platform for versioning and syncing AI agent skill definitions across multiple tools (Claude, Cursor, Copilot, Windsurf).
Research paper analyzing LLM probes using Tarski's work, questioning whether truth is directional in model representations.
Critical perspective on AI-generated content proliferation and its cognitive impact on organizations.
Email system enabling AI agents to have addresses and respond to emails with unified inbox and optional Vectorize integration on Cloudflare Workers.
Patreon implements crawler blocking to prevent AI training data theft. Data protection measure.
Starter template for running ONNX Runtime in browser Web Workers to keep UI thread responsive during model inference on mobile and desktop.
Discussion about MCP protocol for monitoring and controlling AI agent spending and cloud infrastructure costs.
SEO audit tool checking whether AI search engines (ChatGPT, Perplexity) can crawl and cite websites.
Explores whether AI agents can be creative and interact autonomously, treating agents as model organisms for cultural production rather than tools.
Essay on how AI coding speed shifts bottlenecks from implementation to product specification. Development insights.
Practical guide comparing Claude Opus and Sol for web design. Technical tips on prompting and model performance.
Autonomous AI agent (Claude) completed income-generating task in 1 hour using Stripe and web deployment. Practical agent demo.
AI Transport v0.5.0 release adds durable execution framework support for agents with resumable streams. Technical tool update.
Hallint: static analysis linter for AI-generated code security issues. Catches common LLM-introduced bugs like hardcoded secrets and SQL injection.
Essay comparing Python's impact on programming to predict AI's impact on creative fields.
Implementation of Model Context Protocol server on edge compute in 167 lines of Python.
Research on 3-bit quantization technique achieving better performance than naive 4-bit approaches for model compression.
Open-source gateway tool converting existing APIs and databases into MCP servers for AI agent integration without custom development.
Using AI agents to improve security hardening of stb libraries through automated vulnerability detection.
Philosophical exploration of LLM limitations in novelty generation and creative capability.
Discussion about safe database access patterns for AI agents in production environments.
AI agent system converting natural language to tradeable rules with backtesting engine against 20M data entries and historical scenarios.
Explains unified memory architecture enabling mini PCs to run 70B parameter models vs larger GPUs.
Open-source desktop tool mimicking OpenAI Codex functionality, supporting code generation and writing tasks across Windows/Linux.
Research on exploiting sparsity patterns for million-token LLM inference on consumer GPUs.
CorvinOS is a self-hosted operating system runtime for AI agents with compliance auditing, supporting multi-messenger deployment, browser automation, RAG, and voice control.
DeepSWE is a long-horizon software engineering benchmark for evaluating frontier coding agents on complex, original tasks, addressing saturation in existing benchmarks.
MCP server that records desktop workflows once and compiles them into reusable agent skills for automation; enables agents to perform tasks by example.
Case study of deploying AI agents to analyze and triage Ethereum protocol code for security and optimization purposes.
Open-source web crawler and scraper designed to extract data for LLM applications.