Context Doesn't Scale with People
Product for synchronizing agent-based context across teams to reduce time spent on communication tools.
Product for synchronizing agent-based context across teams to reduce time spent on communication tools.
LLM pipeline autonomously generates novel physics research paper end-to-end.
Video discussing practical lessons learned when building AI systems.
Experiment using AI for measurable code review approval with metrics and safety considerations.
Trace: Open-source memory system for LLM agents with self-organizing capabilities, available on PyPI.
Fiszki: Spaced-repetition flashcard app with AI agents creating decks via MCP and FSRS scheduling algorithm.
Analysis of running local LLMs for coding tasks, examining viability and trade-offs versus cloud alternatives.
Comcent CE: Open-source self-hosted voice infrastructure platform providing detailed call analytics and tracking.
Tarit: Rust-based hypervisor and orchestrator designed for running AI agents and RL environments with live snapshots.
DSpark: DeepSeek's speculative decoding technique for accelerating LLM inference.
ZML/LLMD is a cross-platform LLM inference server enabling language models to run on various hardware accelerators.
Agent skill tool that benchmarks and applies 9 token-saving techniques to reduce LLM API costs.
Discussion on best practices and techniques for improving code quality in AI coding agents.
Security vulnerability report of Claude leaking credentials across user sessions.
Developer built a real-time note-taking app using on-device speech-to-text and LLM analysis. Explores practical considerations when choosing AI models for production applications.
Claude/Codex skill implementation enabling LLMs to query and analyze geospatial data.
Single-binary workflow orchestration tool for automation pipelines.
ZML releases LLM inference software supporting multiple open-source models across diverse hardware (Nvidia, AMD, TPU, Metal, Intel Arc).
Research on power-calibrated statistical framework for LLM watermarking, balancing detectability and semantic preservation.
Comparison and guide for selecting AI-powered code assistance tools.
MCP server implementation with persistent memory for contextual conversation.
Guide for self-hosting LLMs using Docker Compose containerization.
Research comparing aligned vs. abliterated LLMs for vulnerability analysis tasks.
Local trace stack for AI agents indexing sessions across multiple providers into ClickHouse with searchable memory via MCP.
Fable Advisor plugin routes subagent tasks to cheaper LLM models while using Fable 5 as architect, optimizing token costs.
Metis open-source security framework uses AI agents for deep code review to detect vulnerabilities and improve secure coding practices.
ProductSpec open standard provides portable Markdown format for software intent documentation before implementation, designed for AI agent handoff.
Shotgun framework turns Claude into persistent AI cofounder for solo founders, managing operations, product building, and distribution.
Skill Retriever: Semantic skill discovery for AI agents across 10K-category taxonomy with Hermes Agent integration.
Personal experience self-hosting LLMs for ChatGPT-like functionality with privacy and ownership control.
Security researchers at Noma Labs discovered prompt injection vulnerability in GitHub's AI agents allowing unauthorized access to private repositories via crafted GitHub Issues.
Instagui converts CLI tools into web GUIs automatically by parsing help text, no configuration required.
Noma Labs security research on GitLost vulnerability in GitHub's AI agents that allows malicious actors to extract private repository data via prompt injection.
ArXiv research paper benchmarking energy consumption of LLM inference across different models and hardware configurations.
AI agent that executes README instructions in sandboxed containers and generates tutorial demo videos, verifying documentation accuracy.
Prompt-to-Paper: Agentic system for bioinformatics that generates manuscripts with verifiable literature grounding and executed experiments.
CSTutorBench: Benchmark for evaluating small language models as programming tutors for block-based education.
LLMForge: Empirical study and benchmark of foundation models for automatic CAD generation from natural language specifications.
Narrative World Model: Memory system for long-form fiction writers tracking narratological structure and story state.
FirstResearch: Framework for auditable research question formation in scientific LLM agents with explainable assumptions.
In-process retrieval as working memory for language agents, reducing latency of in-loop memory access during agent reasoning.
Akashic: Low-overhead LLM inference service using MemAttention for efficient multi-turn agent interactions with extended context.
ArtisanCAD: LLM-based agent for industrial CAD generation from natural language with expert knowledge distillation and parametric geometry.
LLMs generating synthetic consumer data for marketing projective techniques across multiple models and prompting strategies.
AgenticAI-Supervisor: RL Gym environment for evaluating LLM agents with multi-step decision-making and reward shaping.
Synthesis of 27 papers on LLM agent failures across tool-use, planning, and reasoning across 19 benchmarks (2023-2026).
Activation steering method to control tool-use decisions in LLMs by extracting and manipulating internal representations without weight modifications.
NapMem framework enabling conversational agents to use long-term user memory as structured action space via multi-granularity memory pyramid.
On-policy distillation framework optimized for long-horizon language agent training by addressing inefficiencies in full-horizon rollouts and trajectory sampling.
Multi-agent LLM simulator for quantum computing dilution refrigerator fault diagnosis, combining physics models with learned noise fingerprints and LLM operations layer.