Show HN: Claude Gym – a tiny CLI that nudges you to move while Claude Code runs
CLI tool monitoring Claude Code JSONL logs to prompt movement breaks when AI is busy running tools.
CLI tool monitoring Claude Code JSONL logs to prompt movement breaks when AI is busy running tools.
VS Code extension and MCP server for reviewing markdown plans before AI agents implement them with inline annotations.
Cron job monitoring service with built-in AI assistant for debugging and support. LLM-powered DevOps tool.
Atlaz is an LLM app that interviews users about decisions and generates verifiable decision memos with traceable analysis.
AI tool (AISLE) discovers 12 vulnerabilities in OpenSSL and AWS cryptographic libraries. Security research using AI analysis.
Benchmark analysis comparing Go vs Python for AI infrastructure in 2026, covering concurrency, performance, and distributed systems.
Machine learning approach to detect LLM-generated text using statistical characteristics distinguishable from human writing.
Case study building an AI SRE agent in 2 days to automate error log triage and distinguish signal from noise.
VibeWhisper is a macOS voice-to-text tool with push-to-talk support, offering local or cloud-based processing.
AgentThreads is a community directory providing agent-friendly API documentation and best practices for AI agents.
Continuum is a CI drift guard for LLM workflows that replays and verifies multi-step LLM outputs to catch model/prompt changes.
Custom OS implementation of a dependency-free GPT-style LLM without external libraries, built as a minimalism exercise.
Interview with Google engineer who built Gemini CLI, an AI coding agent tool. Discusses team structure and feature development velocity.
Framework for codifying rules and policies to control AI agent behavior in software development. Addresses reliability of autonomous agents.
Informal thoughts on how pull requests may become obsolete in AI-driven development. Opinion piece without technical depth.
TrustStack: vendor review workflow tool generating evidence-backed PDFs, diff reports, and approval logs for contract analysis with citation tracking.
Open-source EU AI Act Article 12 logging infrastructure for reconstructing and auditing agentic events with tamper-proof trails.
Sophify: Distributed AI business operating system with 40+ specialized cognitive agents across 12 servers and 28+ microservices.
GPT-5.3 Instant system card describing safety mitigations and capabilities of latest model version.
OpenAI update to GPT-5.3 Instant model improving conversational fluidity, web search context and tone.
Case study of AI agent handling DevOps tasks competently until hitting limitations requiring human intervention. Documents real-world constraints of autonomous agents.
Ensu: offline LLM application enabling local model execution with privacy. First release of privacy-focused local LLM deployment tool.
Technical breakdown of automated e-commerce video ad generation from single product images using generative models and structured pipeline.
Kanon 2 Enricher: hierarchical graphitization model for transforming document corpora into structured knowledge graphs with entity extraction and linking.
RalphMAD: Claude Code plugin combining BMAD structured SDLC workflows with Ralph Loop self-referential technique for templatized AI-assisted development.
TrueMatch: AI agent system for matching users based on observed behavioral data rather than self-reported profiles. Applies to dating and professional networking.
md-feedback: tool enforcing Markdown plan review before AI agents execute code, preventing hallucinations and unreviewed changes.
OpenClaw AI agent with humanlike memory system using hmem-MCP protocol to retain information across conversations without compressing context.
Company built custom AI agent on LeanMCP platform in 2 hours that answers product documentation questions better than general-purpose LLMs.
Open-source multi-agent reinforcement learning combat simulation with per-agent PPO training, checkpointing, and telemetry logging.
Claude Code skills pack providing AI-powered executive team roles (CEO, CFO, CMO, etc.) to assist solo founders with multi-domain decision-making.
Agent that explores GitHub repositories, indexes codebases, and answers questions about code structure and dependencies using LLM reasoning.
Polynomial Surrogate Training for ternary logic gate networks: Method to extend differentiable logic networks to ternary Kleene logic with uncertainty handling.
Theoretical analysis of when chain-of-thought prompting helps LLMs using Markov chain modeling to identify transition alignment as key factor.
Vectorized Adaptive Histograms for Sparse Oblique Forests: Method to optimize histogram and sorting tradeoffs in random forest training.
SpeedTransformer: Transformer-based model using smartphone GPS data to detect transportation modes, outperforming LSTM baselines.
Studies catastrophic forgetting in IoT intrusion detection systems under distribution shifts from evolving attack patterns.
Geometric meta-RL approach leveraging task space symmetries for improved generalization in reinforcement learning.
USE introduces lightweight procedure for semi-supervised learning that estimates uncertainty structure to handle out-of-distribution unlabeled data.
Quantum optimization approach for exact robust verification of neural networks against adversarial perturbations.
Establishes equivalence between activation steering and weight-space updates, providing principled foundation for parameter-efficient LLM adaptation.
Studies decoder scaling strategies in construction-based neural routing solvers for vehicle routing optimization problems.
ROKA addresses machine unlearning robustness against adversarial attacks that exploit knowledge contamination in unlearned models.
RapTB improves GFlowNet training for fine-tuning LLMs by addressing prefix collapse through trajectory balancing and submodular replay.
FEWTRANS benchmark evaluates few-shot transfer learning of pre-trained models with improved evaluation protocols across 10 datasets.
HL-SMM introduces Heaviside loss-based support matrix machine for classification of matrix-structured data with noise robustness.
ESENSC_rev2 proposes polynomial-time feature attribution algorithm as computationally efficient alternative to SHAP using game theory.
Antibody introduces defense mechanism against harmful fine-tuning attacks on LLMs by regularizing gradient contributions of poisoned samples.
FastBUS proposes a Bayesian framework for weakly-supervised learning that handles multiple label types efficiently with batch processing.
Bridge Matching Sampler for scalable sampling from unnormalized densities using generalized fixed-point diffusion matching.