Show HN: OQP – A verification protocol for AI agents
OQP: MCP-compatible verification protocol for autonomous AI agents. Defines standards for verifying agent-generated code meets business requirements via four core endpoints.
OQP: MCP-compatible verification protocol for autonomous AI agents. Defines standards for verifying agent-generated code meets business requirements via four core endpoints.
Mercury: orchestration platform for managing multiple AI agents across tools. No-code interface for coordinating agent workflows in enterprise settings.
OQP: MCP-compatible verification protocol for autonomous AI agents. Defines standards for verifying agent-generated code meets business requirements via four core endpoints.
Analysis of US-China AI capabilities gap in 2026. Premium content discussing relative progress and competition in AI development.
LARQL: query language for decomposing transformer weights into queryable vector indices. Decompile models into vindex format for browsing and editing knowledge without GPU.
Study examining why LLMs fail to retain corrections across multiple interactions. 19,600-word analysis with lab data from civil engineering firm using Claude Code.
N-Day-Bench: benchmark measuring LLM capability to find real vulnerabilities in codebases beyond knowledge cutoff. Monthly-updated adaptive test suite for security evaluation.
Analysis of LLM token usage patterns through agent harness implementation. Real-world example of token consumption in multi-service agent tasks.
Stanford annual AI report documents diverging views between AI experts and public, rising anxiety about jobs and societal impact.
SaaS product using GitHub integration to onboard developers faster through automated codebase documentation.
Analysis of Claude API service quality decline, outages, and cost increases reported by users and internal testing.
Aibom Scanner open-source tool detects AI SDK usage in codebases and maps compliance gaps to NIST, ISO 42001, EU AI Act frameworks.
Poke startup launches AI agent accessible via iMessage, SMS, Telegram for personal assistance and smart home control.
Proposal for LLM continual learning using Markdown files and semantic filesystem for cheap long-term memory without code.
Analysis with graphs of AI industry trends: rising investment, accelerating model capabilities, mixed job/perception impact heading into 2026.
ContextNest is open standard for structured knowledge management in AI agents using versioned markdown, deterministic queries, and audit trails for enterprise context governance.
Analysis of shift from AI chat to agentic layers in software engineering. Discusses workflow orchestration and system integration beyond prompt-based tools.
Flickspeed provides shared workspace harness for multimodal creative agents handling image, video, audio, and research tasks. Orchestration platform.
Comparison of on-device vs cloud LLMs for agentic tool use in iOS travel concierge app. Tests Apple Foundation Models 3B vs GPT-OSS 20B for multi-step agent tasks.
AImeter tracks costs and value metrics for AI agents with zero dependencies, local-first, no cloud required. Developer tool for agent economics.
Coyns enables AI agent-to-agent transactions using MCP-native currency. Economic layer for autonomous agent interactions.
Sentō provides open-source AI agents running on Claude API subscription with one-command setup. Developer tool for agent deployment.
Benchmark of Google Gemma 4 E2B (2B parameter model) vs larger Gemma variants on 10 enterprise task suites. Local inference on Apple Silicon.
GAIA is an open-source framework for building local AI agents in Python and C++ on AMD hardware without cloud dependency.
Claude Code skills for network engineering enable LLM assistance on enterprise networking concepts, protocol explanations, and homelab troubleshooting.
Lint-AI is a Rust CLI for indexing and retrieving evidence from AI-generated documentation corpora, extracting context from large task traces.
Conceptual essay framing AI agents as control systems similar to autonomous vehicle architecture, emphasizing execution, verification, and telemetry.
Personal narrative on choosing to work without AI assistance due to client restrictions and preference for traditional coding flow.
Analysis of AI adoption challenges in legal profession, examining claims versus practical implementation realities.
Tracker comparing frontier AI models with benchmarks, pricing, and API capabilities across proprietary and open-weight options.
DigitalOcean platform for building and deploying AI agents and applications.
Discussion on practical effectiveness of AI agent skills versus manual team processes.
Evaluation of Claude Mythos Preview model's cybersecurity capabilities and performance.
Applies AI agents to economics research tasks.
Analysis challenging claims about AI model training environmental emissions.
Netflix uses LLM-as-judge approach to evaluate and improve show synopsis quality for personalization.
Custom PDF conversion benchmark with evaluation methodology.
Study showing AI chatbots misdiagnose medical conditions in over 80% of early-stage cases.
Open-source tool that maps codebase architecture for AI agents, eliminating context loss between sessions.
Meta develops AI chatbot trained on Mark Zuckerberg's communications for employee interactions.
Open-source voice-first local AI assistant for macOS with context awareness, using Gemma and Qwen models.
Guide on running and coding with local AI models on macOS. Practical developer resource.
Linux 7.0 kernel release announcement with mention of AI applications in bug-finding. Limited AI focus.
Stanford's 2026 AI Index report analyzing AI progress, development rates, and industry trends with data.
Open-source Markdown editor featuring AI-native comments and suggested edits functionality.
Bug fix for Claude Code prompt caching using environment variable workaround to avoid tool-block blocking.
Educational resource deconstructing Steve Yegge's Gas Town agent orchestration framework with spaced repetition prompts and linked concepts.
Analysis of AI research outside frontier labs using alternative approaches to GPU constraints.
Leaderboard evaluating LLMs on agentic search tasks with open methodology and real shopping queries.
Open-source LLM operations dashboard monitoring provider health, costs, latency, and concentration risks.