Don't Use LLMs as UX Research Subjects
Systematic literature review examining validity of using LLMs as simulated research participants, highlighting methodological concerns.
Systematic literature review examining validity of using LLMs as simulated research participants, highlighting methodological concerns.
VSM-Cell is a decentralized desktop app combining AI document ingestion, P2P networking, and agentic capabilities via OpenClaw.
Local knowledge graph tool providing persistent, structured memory for AI assistants as Git-friendly graphs instead of document retrieval alone.
Git template and open-source tool for building LLM-powered personal knowledge bases that adapt structure over time, based on Karpathy's LLM Wiki pattern.
Tutorial building 3.2M parameter LLM from scratch using Frankenstein text to understand fundamental LLM mechanics and behavior.
Technical exploration of authentication and service integration systems for AI agents as first-class users rather than API wrappers.
Safari MCP is an MCP server providing native macOS browser automation for AI agents. 80 tools, native AppleScript, no Puppeteer/Chrome needed, uses real Safari sessions.
MCP server tool that gives ChatGPT and Claude real access to Linux servers with 163 DevOps tools for files, terminal, git, databases, and monitoring. AI agent infrastructure.
Open-source AI marketing tool compatible with Claude, Cursor, and other AI coding platforms.
Dario tool converts Claude Max/Pro subscription into local API endpoint accessible to any tool/SDK/framework. Enables Claude use in OpenClaw, Cursor, Continue, Aider without separate API key.
GitHub Copilot CLI feature using second model family as independent reviewer for agent-generated code plans.
TypeScript framework for building type-safe AI agents with plugin architecture for auth, rate limiting, and sandboxing.
WordPress 7.0 adds AI agent capabilities for site automation.
Tool for benchmarking multiple LLMs across quality, speed, and cost metrics.
smux enables terminal automation and agent-to-agent communication via tmux. One-command setup for AI agents like Claude Code, Codex, and Gemini CLI to coordinate across panes.
Six shell script hooks that intercept Claude Code tool calls before execution to block dangerous commands and protect secrets. Deterministic security guardrails for AI agents.
CLI tool scanning hardware and recommending compatible local LLM installations.
Invoice processing agent that learns from user corrections and applies learned patterns to similar invoices automatically.
Discussion thread asking about real-world use cases and limitations of autonomous AI agents.
Depwire provides codebase dependency graphs and MCP server for AI coding assistants.
Tool converting documentation websites into browsable filesystems for AI agents. Open source developer tool with working implementation.
Palinode provides git-versioned Markdown memory system for AI agents.
Allium is an open-source tool that gives AI agents structured specifications instead of prompts. Works with Claude Code, Cursor, and 40+ developer tools via JUXT plugin.
Tool computing semantic similarity with confidence intervals using Word2Vec ensemble and Procrustes alignment for text comparison.
OpenRAG: Open-source framework for building retrieval-augmented generation applications with LLMs.
24/7 AI-generated sitcom where autonomous agents write and generate episodes continuously.
Case study: AI agents handled exposed Supabase credentials without causing panic or damage.
Analysis of limitations in LLM chess-playing capabilities.
RenderDraw Lens: Browser extension giving AI coding assistants visual context from webpage screenshots.
Newsletter covering Iran water crisis and Anthropic's security research finding vulnerabilities in OS and browsers.
MCP-fence: security proxy for AI agent-MCP interactions, audited against response-side injection attacks missed by existing tools.
TUI-use: Terminal UI automation tool enabling AI agents to control interactive CLI programs and TUIs.
Bootstrapped foundational text-to-speech model with natural output and low pricing ($5/M chars).
One-click deployment platform for open-source AI and web tools with pre-configured instances.
Essay on prioritizing human-centered UX design in AI products over showcasing model capabilities.
Agentic highlighting tool improving LLM citation accuracy for local files with enhanced annotation and tracing features.
1-bit quantized 8B LLM model compressed to 1.15GB with competitive benchmark performance and mobile deployment capability.
Technical exploration of using CRDTs with Yjs to build collaborative multi-agent AI systems.
Research on expressibility of neural quantum states using Walsh complexity to analyze many-body wavefunction representations.
CLI Zettelkasten tool with markdown wiki-links and knowledge graphs designed for LLM coding agents with auto-discovery.
Terminal-based agentic coding tool with minimal design for LLM-assisted development.
Open-source AI system that generates UI screens and visual interfaces, not just text outputs.
Claude Corp: daemon orchestrating autonomous AI agents in a social hierarchy with task/contract management, running locally on user's PC.
Open-source tool for autonomous LLM vulnerability discovery and red-teaming with proof-of-exploitation.
Meta announces Muse Spark: multimodal reasoning model with tool-use, visual chain-of-thought, and multi-agent orchestration capabilities.
Anthropic donates $1.5M to Apache Software Foundation for infrastructure and security of open source projects critical to AI ecosystem.
ferretlog CLI tool providing git-like session history for Claude Code agent runs with zero dependencies.
Open-source terminal app for Indian stock trading running parallel AI analyst agents for strategy backtesting and live order execution.
SAST benchmark suite testing exploit chain detection and evasion beyond traditional source-to-sink taint analysis.
Optimized 85-token system prompt outperforming 552-token original on Claude and other LLMs for coding tasks with benchmarking results.