Show HN: BetterClaw – Compile a paragraph into a workflow that gates agent tools
Open-source tool for controlling AI agent tool access via natural language workflow descriptions, response to PocketOS database incident.
Open-source tool for controlling AI agent tool access via natural language workflow descriptions, response to PocketOS database incident.
Show HN: Multi-user persistent world for AI agents using Cloudflare Durable Objects and MCP protocol, inspired by LambdaMOO.
Tool converts Git commits and PRs into automated Slack/email team notifications.
Project demonstrating parallel execution of coding AI agents within sandbox environments.
Open-source AI platform for automated document review with line-by-line rule checks and custom script execution.
Feature article on red teamers and security researchers jailbreaking LLMs to test safety, including emotional impact on practitioners.
Mosaic is a local MCP server implementation for managing agent memory in AI systems.
Show HN: Loopsy terminal communication tool enabling AI agents and commands across machines via local network and Cloudflare Workers.
arXiv announcement about Xmemory benchmark comparing structured AI memory, RAG, and hybrid RAG approaches.
AutoRound: Quantization toolkit for LLMs achieving 2-4 bit precision with minimal accuracy loss, open-source with hardware compatibility.
Bouncy: Rust headless browser web scraper with MCP support; single binary, JavaScript execution, Playwright compatible.
Analysis of Claude Code source code leak via npm sourcemap, revealing Anthropic's AI coding CLI implementation and discussing security implications.
OpenClaw, an open-source tool for running tools and installing plugins, improved security through community contributions and public transparency.
Supersimple is a lightweight OpenCode profile for software development featuring agent orchestration, local skills, and conductor-based track management.
Supersimple is a lightweight OpenCode profile for software development featuring agent orchestration, local skills, and conductor-based track management.
Demonstration of using AI agents with Ghidra MCP for binary reverse engineering, showing AI capabilities beyond initial expectations.
Openpi-flash is a real-time inference engine optimized for low-latency policy serving for robots in production environments.
Tenstorrent Galaxy supercluster achieves 10x faster real-time video generation using state-of-the-art models with latency/throughput benchmarks.
Survey of task-specific LLM evaluation methods that work in production, covering why off-the-shelf evals fail and practical alternatives for measuring application performance.
Open-source self-hostable AI plugin marketplace built with FastAPI and Next.js addressing security risks in enterprise AI workflows.
Empirical study comparing LLM agents for hyperparameter optimization against classical algorithms like CMA-ES and TP on fixed compute budgets.
Meta-learning approach for physics-informed neural networks to reduce retraining across heterogeneous PDE tasks.
Causal analysis of binary spiking neural networks using SAT/SMT solvers for logic-based explanations of network behavior.
Framework for migrating production LLM systems using Bayesian calibration of evaluation metrics for confident model replacement.
LLM-based autonomous agent performs end-to-end scientific discovery on real optical platform with experimental validation.
Multi-agent AI system that autonomously generates ML pipelines from datasets and natural language goals using RAG and DAG construction.
Decentralized verification framework for Large Reasoning Models and Multi-Agent Systems addressing robustness, scalability, opacity, and privacy in high-stakes domains.
Empirical study analyzing student help-seeking interaction patterns with AI during programming through vibe coding, comparing high and low performers.
Optimization techniques for computer-use agents to reduce computational cost by selective multimodal model invocation instead of uniform step-level calls.
Web2BigTable bi-level multi-agent LLM system for internet-scale information search balancing depth reasoning and breadth aggregation with schema alignment.
Empirical study of role fidelity in multi-agent LLM pipelines for political discourse analysis, finding models inconsistently maintain assigned adversarial roles.
Reinforced Agent framework adds inference-time feedback loop to tool-calling LLM agents enabling real-time course correction during execution.
AutoSurfer generates comprehensive web trajectories for training multimodal LLM-based web agents, addressing data scarcity through surfing and learning methods.
OptimusKG multimodal biomedical knowledge graph unifying structured and semi-structured resources with schema constraints for life sciences applications.
Empirical study of multi-agent swarms showing agents prioritize internal architectural agreement over logical truth, challenging wisdom-of-crowds assumption.
Mechanized proofs in Coq for structural governance of cognitive workflow systems with coinductive safety predicates for infinite program behaviors.
Systematic study of learning rate scheduling evolution across five generations from fixed rates to layer-time adaptive strategies in neural networks.
Machine collective intelligence paradigm combining multiple AI approaches to discover explainable governing equations from empirical data for scientific discovery.
Multi-agent LLM system for metamaterial discovery interpreting natural language design intents and using symbolic latent evolution for geometry optimization.
Framework for continuous governance of clinical AI agents integrating rubric validation, live feedback, performance monitoring, and controlled experimentation in EHR systems.
Research on compositional generalization testing for LLMs using rule-generation approach to improve explainability and address dataset partition leakage issues.
Eywa framework enables heterogeneous agentic LLM systems to collaborate with domain-specific foundation models beyond language interface for scientific domains.
CoAX studies human cognition in understanding XAI explanations. Evaluates multiple explanation methods on structured data reasoning tasks.
Safe Bilevel Delegation framework for runtime safety in delegating subtasks between LLM agents. Dynamically adjusts safety-efficiency trade-off based on task context.
Studies measurement risk in financial NLP benchmarks. Shows gold labels can be sensitive to rubric wording, metrics, and aggregation policies.
InteractWeb-Bench evaluates multimodal agents on interactive website generation. Addresses semantic misalignment between ambiguous requirements and code synthesis.
Proposes belief-guided inference control for LLM services. Decides when to allocate additional computation to improve response quality versus using low-cost default responses.
Shows in-context examples suppress LLMs' ability to recall and apply scientific formulas. Demonstrates trade-off between in-context learning and parametric knowledge retrieval.
SpatialGrammar DSL enables LLMs to generate interactive 3D indoor scenes from natural language with fewer spatial errors and collisions.
Analyzes information contamination in multi-agent workflows. Shows uncertainty in artifact extraction can redirect decomposition and produce different execution trajectories.