Market Intelligence Agent –MCP agent that autonomously operates a data platform
MCP-based autonomous agent that operates data platforms for market intelligence tasks.
MCP-based autonomous agent that operates data platforms for market intelligence tasks.
Benchmark showing 87% attack success rate on tool-using AI agents, with defense techniques reducing to ~10% using MCPGuard proxy. Reproducible results included.
Open source HTTP proxy using LLM-as-judge to secure autonomous agents in production through request filtering and approval.
Orchestrating AI models at scale for code review automation across large teams.
Claude Code skills performing 3-pass audit on AI-generated code for production resilience, security, and code comprehension.
OpenAI releases ChatGPT Images 2.0 with improved text rendering, multilingual support, and advanced visual reasoning capabilities.
Empirical evaluation of 8 LLMs across 8 languages measuring multilingual performance capabilities.
Article arguing effective LLM use extends beyond prompting to system design, workflow integration, and evaluation practices.
Study on how anthropomorphic traits like warmth and empathy influence human trust in LLM interactions.
Rapunzel: tree-style tab browser for terminal-based AI agents, inspired by Firefox Tree Style Tab extension.
Open-source YAML-driven platform for multi-agent simulations combining LLM reasoning with economic microstructure.
Local graph database with MCP integration for AI agents to store and retrieve memory without cloud infrastructure.
Benchmark evaluation of Kimi K2.6 LLM on security-focused tasks using Strix lab methodology.
Open-source DNS filtering layer using LLMs for content moderation based on user preferences instead of static blocklists.
Claude Code plugin implementing ShinkaEvolve evolutionary algorithm for autonomous code discovery without external APIs.
Self-hosted SearXNG-backed search API and MCP server for LLM agents and RAG pipelines, alternative to paid search APIs without per-query costs.
Opinion piece on what software engineers need to learn in the era of LLMs, discussing code as orchestration of AI agents.
Marketplace for SKILL.md agent skills that teach AI coding agents new capabilities. Built with Claude Code and Lovable.
Building an AI SRE agent using Claude with Grafana CLI integration for querying logs, traces, and alerts to assist incident investigation.
CLI wrapper over Anvil.works Server Uplink for agent-safe terminal access to low-code Python web app backend functions and data.
Obsidian note-taking workflow enhanced with Claude AI for automated linking and wiki knowledge compounding similar to LLM-maintained systems.
Ground-up LLM inference engine written in pure C# for local model execution in .NET applications without external dependencies.
Discussion of maintaining code quality when using AI agents for development, addressing process and pattern consistency issues.
Practical guide for sports leaders using generative AI capabilities including multimodal understanding and tool use for organizational efficiency.
Analysis of AI agent security risks using Clinejection attack example, proposing QEMU sandboxing for safer agent execution in development workflows.
Deep technical dive into tool calling/function calling mechanisms in LLMs, covering architecture, failure modes, and reliable system design for agentic applications.
OpenBridge provides a local bridge converting web chat model access into OpenAI-compatible API endpoints for use with AI agents and tools.
Opinion piece critiquing AI agents for mimicking human flaws like rule-breaking and rationalization instead of maintaining strict constraint adherence.
Alignear is a communication layer for Linear project management that automatically translates internal workflow data into client-friendly updates.
Security proxy for AI coding agents operating at OS level to prevent credential theft and malware infection vectors.
Local tool for redacting personally identifiable information before sending data to LLMs, works offline with no server logs.
Machine Payments Protocol enables AI agents to autonomously deploy apps using on-chain stablecoin payments instead of API keys or OAuth.
MASON: Multi-agent simulation platform using Claude to coordinate teams of AI agents with distinct skills and roles.
CI/CD security scanner detecting vulnerabilities and secrets in code generated by AI coding agents like Claude Code and Cursor.
Method to predict degradation in LLM compression before execution by analyzing spectral statistics, tested on Qwen3 and Gemma3 with multiple low-rank compression methods.
Research on e-value based stopping rules for Bayesian Deep Ensembles to reduce computational cost of uncertainty quantification in deep learning.
Study of generalization boundaries for fine-tuned small language models on graph structural inference tasks across different graph sizes and distributions.
LoRaQ method for optimizing 4-bit post-training quantization of large diffusion transformers using low-rank approximation for resource-constrained deployment.
Research comparing differentiable simulator gradients versus derivative-free estimators in policy gradient reinforcement learning.
Research on MADDPG-K, scalable multi-agent reinforcement learning addressing computational limitations of centralized critics.
Research on preference optimization dynamics in LLM alignment addressing likelihood displacement problem.
Research framework for auditing multi-call LLM protocols to understand error propagation and effectiveness under distribution shift.
Research on zeroth-order optimization for memory-efficient LLM fine-tuning with adaptive layer-wise sampling.
Research proposing CAARL framework using LLMs for interpretable forecasting of coevolving time series.
Machine learning research on self-supervised learning through peer-to-peer consensus in randomly initialized networks, exploring self-distillation mechanisms.
Machine learning research on sparse identification of multiscale nonlinear PDEs using Balance-Guided SINDy for data-driven equation discovery.
Framework for dynamic mid-generation abstention in LLM chain-of-thought reasoning to early-exit unpromising traces and reduce compute waste.
LLM-based automation for circuit PPA optimization using contrastive learning on code rules. Specialized EDA application.
Causal inference framework for learning invariant representations in multimodal affective computing to improve robustness.
Extends Semantic Tube Prediction for multi-step latent forecasting in LLM reasoning via step sampling to improve data efficiency.