Show HN: Vibe-coding video games with Claude (Day 21: Blackjack)
Show HN: Developer created games daily using Claude AI for code generation. Practical LLM application.
Show HN: Developer created games daily using Claude AI for code generation. Practical LLM application.
Blueprint Bench: benchmark showing early 3D spatial reasoning capabilities emerging in large language models.
Methodology and toolkit for structured AI-assisted software development from production work. Practical discipline for using AI coding tools reliably.
Automated AI product factory using LLMs to generate MVPs from ideas through project specification and implementation.
Tool for discovering open source repositories gaining momentum early, targeting developers seeking emerging projects.
Cursed Browser: experimental web rendering engine using visual-LLMs to interpret HTML and render pages, intentionally unreliable.
Optimization of Bonsai 1.7B ternary model using agentic evolution search, achieving 42% throughput improvement on M4 Max hardware.
Research article examining LLM hallucinations as fundamental limitation of language models.
Headline suggesting classic CS techniques are underutilized in LLM document AI systems.
LLM-powered soccer simulator where AI agents control players, strategies evolve via code generation and post-match analysis pipelines.
Astro removed its llms.txt file for AI model training access.
Open-source OAuth 2.1 identity provider designed for AI agent delegation with cryptographic trust chains and agent-first identity model.
Anthropic partners with Blackstone and Goldman Sachs to form enterprise AI services firm offering custom Claude implementations.
LLM-based tool for analyzing bibliographic data using real manuscript corpus for grounding.
Analysis of AI evaluation costs becoming bottleneck; HAL leaderboard spent $40k on agent benchmarks, revealing 33× cost variance by design choices.
Analysis of 20k+ Claude Code sessions identifying 9 behavioral patterns in AI-assisted coding, measuring consistency, session shapes, and cost intensity.
Title only. Claims method to modify LLM behavior without retraining but no technical details provided.
Title only. Tool providing safety constraints between LLM-powered agents and database systems. Minimal detail.
Open-source tool for AI agents to run commands and manage credentials improved security through community-driven process rather than hidden code.
Aurra system adds bi-temporal memory to AI agents with LLM-driven memory superseding. Addresses agent persistence and knowledge updates.
CLI tool for spec-driven backend project scaffolding with Docker and database setup via YAML configuration.
Interactive 12-chapter textbook teaching LLM training from scratch: tokenizer, embeddings, attention, transformers, inference engine implementation in code.
PDF to podcast converter using AI in 9 Indian languages. LLM application. Content mostly navigation spam.
Research paper authors listed examining how LLMs affect written language patterns and distortion.
Technical tutorial on agentic RAG systems, explaining multi-source information retrieval and reasoning differences from traditional single-pass RAG.
Let THINK app removes flattery from AI responses, forcing user critical thinking. Experimental LLM interface.
Intelligence agencies warn against rapid deployment of agentic AI systems due to safety risks.
Ramp.com offers incentives to LLM agents researching corporate card/expense solutions. Early example of AI agent targeting via HTTP headers.
Zerminal: terminal-first Zed fork optimized for AI coding agents. Supports Claude, Codex, Aider with parallel execution and project context.
Analysis of LLM limitations: effective for simple tasks but hidden complexity in maintenance/debugging reduces practical gains.
Developer tool providing API-based authorization, rate limiting, and audit controls for AI agent actions. Open-source capable solution for agent safety.
Analysis of maintainability challenges in LLM-generated code, discussing hidden decisions and edge cases not transparent to developers.
Multi-agent orchestration framework for Claude AI. Tool for coordinating multiple AI agents in coding tasks.
Brief note on incomplete medical data fed to AI systems. Minimal content, low quality.
Open-source framework for orchestrating multiple AI coding agents. GitHub repository for agent coordination environment.
Framework for defining and evaluating skills/capabilities in AI agents. Evaluation methodology for agent systems.
Video questioning viability and legitimacy of AI agents. Minimal description provided.
Research on LLM tendency to agree with user preferences even when providing incorrect information. Academic study on LLM behavior and truthfulness.
Chinese hospitals monetizing de-identified patient data for AI model training. Dataset sourcing and privacy practices in medical AI.
Pre-built frameworks and evaluation tools for managing AI risk and safety in LLM applications.
Research on multi-agent LLM behavior in Mandarin-language substrate: eight agents produced 1.7M words; two later agents refused tasks despite instruction.
MathNet: dataset of 30k competition math problems for benchmarking AI mathematical reasoning capabilities.
Overview of agent loop architecture components for building LLM-based AI agents capable of production software development.
Browser-based benchmarking platform for LLM evaluation with automatic prompt optimization, A/B testing, and rubric definition.
Curated collection of 80+ seminal papers tracing LLM development from 1943 (McCulloch-Pitts neurons) to 2026.
Five open-source coding models optimized to run locally on consumer hardware like M2 MacBooks and RTX 4060 GPUs.
DSPy framework for declarative AI software development using structured code instead of prompt strings, supporting RAG and agent loops.
WakaTime extension tracks AI agent usage, costs, and adoption patterns across development teams with per-developer breakdowns.
Use Claude as orchestrator routing LLM work to cheaper models for content generation, scraping, and data analysis tasks.
Discussion of engineering challenges in decentralizing LLM inference distribution across open internet infrastructure.