Designing AI agents to resist prompt injection
Security research on prompt injection attacks against AI agents. Covers threat modeling and mitigation strategies.
Security research on prompt injection attacks against AI agents. Covers threat modeling and mitigation strategies.
JetBrains research on context management for LLM-powered agents, addressing token accumulation and information retrieval efficiency.
Analysis of developer attitudes toward AI-assisted programming. Opinion piece on productivity vs craftsmanship.
Model Context Protocol becoming standard infrastructure for AI agent interoperability and tool access.
Developer tool enabling instant website generation for AI agents without authentication.
Amazon strengthens safety guardrails after CodeWhisperer Q caused production disruption.
Study finds best LLMs pass bar exams but make errors on 1 in 12 local legal queries.
Open-sourced real-time voice-to-voice desktop translation application.
Analysis of autonomous agent expectations and limitations. Critical perspective on OpenClaw and agent automation hype.
Captain: automated RAG pipeline builder indexing cloud storage (S3, GCS) and SaaS sources (Google Drive) for unstructured data search.
Pentagon labels Anthropic Claude a supply-chain risk; dispute over military AI use and autonomous weapons.
Fully autonomous newspaper operated by 18 AI agents generating content, design, and code without humans.
Sandbox tool isolating AI coding agents per project using Nix/direnv with three security levels.
Lightweight local AI assistant (40MB) for task automation, file operations, and database analysis.
Local-first AI tool for Kubernetes incident detection and root cause analysis without SaaS.
Flightplanner is a spec-driven E2E testing framework adapted for AI agent-generated code, addressing testing challenges as agents shift development lifecycle weights.
OpenViking is a context database designed for AI agents.
Mesa: collaborative canvas IDE for agent-first development featuring multiplayer support, integrated terminal, browser, and file management.
UberSKILLS is an open-source web workbench for authoring reusable Agent Skills (instruction sets) for code agents with AI-assisted creation and validation.
Hunter Alpha is a 1T parameter model with 1M token context window designed for agentic use, supporting long-horizon planning and multi-step task execution.
Three-line security wrapper preventing hallucinations and dangerous actions in browser-use and LangChain agents.
Open-source CLIs for construction software APIs (Procore, EagleView) used by AI operations agents.
Open-source sandbox runtime for secure AI agent execution on code repositories. Isolates agent tasks while protecting credentials and publishing.
Web tool implementing Proposer-Critic-Verifier pipeline to automatically refactor unstructured prompts into clearer specifications for stable LLM responses.
WritBase: open-source task management system for AI agent fleets with MCP (Model Context Protocol) native support.
Guide on reinforcement learning environments and construction. Likely introductory content on RL fundamentals.
HAL: command guard tool for AI coding agents that enforces safety constraints on agent-executed commands. MIT licensed lean security layer for agent frameworks.
OpenLight: lightweight Telegram-based AI agent for Raspberry Pi built in Go, runs local LLMs for system checks, service control, and chat without heavy autonomous frameworks.
Molting.org proposes a verifiable reputational layer for AI agents.
Docker and containerized dev environment setup for securely running autonomous AI agents with isolation.
Argus: AI agent that autonomously investigates infrastructure anomalies and proposes remediation steps.
Crashloop Analyzer: tool to diagnose Kubernetes pod restart failures by parsing logs and suggesting fixes. Targets OOMKilled, ImagePull, and config errors.
Article discussing tendency of AI chatbots to agree with users even when factually incorrect. Behavioral analysis.
AntroCode: ultra-lightweight single-file Python LLM client with cyberpunk UI, zero dependencies, supports DeepSeek API, built as alternative to heavy node_modules.
TelsonBase: open-source Apache 2.0 self-hosted governance system for autonomous AI agents.
Data layer infrastructure built for AI agents to enable file transfer and data management.
CLI-Anything converts any software into agent-ready interfaces for AI agents like Claude, OpenClaw, and Cursor.
Collection of 178 reusable e-commerce skills/instructions for AI assistants to perform store management tasks.
Step-by-step guide to implementing safety layers for LLM-assisted code development, covering pre-commit hooks, local review agents, and CI workflows.
Open protocol enabling AI agents to interact with websites via standardized agent.json file, similar to robots.txt.
IH-Challenge dataset improves instruction hierarchy, safety steerability, and prompt injection robustness in frontier LLMs.
Spine Swarm: multi-agent system on visual canvas for complex non-coding tasks like competitive analysis and financial modeling.
Infrastructure for capturing reasoning data including agent decisions, extracted knowledge, and context handoffs.
Step-by-step tutorial with 18 progressive lessons to build AI agents from simple chat to OpenClaw-like architecture.
Python service bridging Telegram bot with Cursor Cloud Agents API to run workflows and manage pull requests via chat.
Production-ready AI agents platform with 177 skills, 16 agents, 3 personas, and orchestration protocol for AI coding tasks.
Community proposes AI agents to recreate proprietary software as freely licensed code alternatives.
DashClaw: auditing tool that intercepts and reviews AI agent decisions before execution, improving transparency and safety.
chat.nvim v1.4.0: Neovim AI plugin supporting multiple LLM providers with tool system, memory, and external chat integrations.
Apple updates developer agreement with requirements for AI model guardrails and Foundation Models Framework compatibility.