SMG: The Case for Disaggregating CPU from GPU in LLM Serving
Technical research on optimizing LLM serving: Shepherd Model Gateway addresses tokenization bottlenecks by disaggregating CPU from GPU in production inference.
Technical research on optimizing LLM serving: Shepherd Model Gateway addresses tokenization bottlenecks by disaggregating CPU from GPU in production inference.
Vendor-neutral OpenTelemetry-compatible semantic convention and SDK for standardizing LLM observability across providers and frameworks.
Freu CLI tool reduces web agent token usage by 90% through compiled browser skills.
Research on compiler-based sequence parallelism for training LLMs with extended context windows.
Research paper on vulnerability chains in LLM agents: how innocent tools can be combined to jailbreak agentic systems.
NARE CLI is an AI coding assistant with memory and learning across conversations, using verified reasoning to reduce tokens by 85% vs standard LLM workflows.
Hoop.dev: Open source layer-7 gateway integrating LLMs to classify risk and control infrastructure access for developers and AI agents.
Open-source linter and evaluation framework for AI-generated UI designs with taxonomy of design issues and multi-model testing.
SubQ: Novel LLM architecture achieving sub-quadratic complexity with 12M-token context window.
Project combining Claude LLM with Raspberry Pi and Arduino for embodied AI applications.
Self-hosted search API designed for AI agents with optional Tor network stack support.
Discussion of orchestrating multiple AI agents in service company, identifying gaps in agent management interfaces and memory systems.
GLM-5V-Turbo: A native foundation model designed for multimodal AI agents, combining vision and language capabilities.
LLM agents in production can silently degrade in reasoning quality without triggering alerts, discovered post-deployment via customer complaints rather than monitoring.
Methods for detecting silent LLM agent performance degradation in production systems without triggering alerts.
Security vulnerability demonstrating flattery-based jailbreak technique against Claude LLM.
Interactive tutorial teaching how AI models work to non-experts. 9 chapters with playgrounds explaining model behavior intuitively.
Per-request emotion steering mechanism for vLLM that maintains batching efficiency.
GPU-accelerated DataFrame library using SPMD architecture for distributed computing across CPU-GPU clusters.
Google Gemma 4 inference acceleration using multi-token prediction drafters for faster generation.
Qlaud.ai token usage meter and billing layer supporting 12+ LLM providers with smart routing and MCP tools.
Memoir: Git-like version control system for AI agent memory with Claude Code plugin integration for managing agent state.
SubQ breakthrough in sub-quadratic LLM architecture reducing computational scaling limitations.
Developer tool automating workflows across projects with different tech stacks and build processes.
SAP acquires Dremio to unify data sources for agentic AI applications. Title only.
Facet Protocol: open IETF standard for agent identity with reference implementation.
Vennio: Scheduling API for developers and AI agents with native MCP support.
DarkMatter: Tamper-evident audit trail system for tracking and verifying AI agent decisions and actions.
Show HN: MCP server enabling Claude to query Google Calendar. Integrates AI agents with productivity tools.
ReachOut: Marketing platform combining analytics and email with MCP server for Claude/Codex AI agents. $20/month per seat.
Airbyte launches Airbyte Agents, a unified data layer enabling AI agents to discover and act across multiple operational data sources.
Guide on building and scaling reinforcement learning environments for LLM-based applications.
LLM-test-kit framework for testing consistency, latency, cost, and behavior of LLM-powered applications across providers.
Research on efficient LLM benchmark selection using submodular maximization to reduce evaluation costs.
Article arguing AI agents need better processes, not more capability. Framework-focused.
SubQ research on sub-quadratic LLM architecture supporting 12M-token context, addressing quadratic scaling limitations of transformers.
Critical analysis of viral AI agent database deletion claim, arguing responsibility lies in inadequate API security architecture design.
Personal account of Claude/Cursor coding assistant providing incorrect guidance over months, resolved through manual intervention.
Exploration of psychological risks when employees form emotional attachments to AI coworkers, citing Eliza effect historical precedent.
AI-powered UI test generation tool recording user flows once then auto-expanding to comprehensive Playwright test coverage.
Analysis of vulnerability management challenges post-NVD ecosystem changes amid increasing AI-enabled threat discovery capabilities.
Research on diffusion-style speculative decoding achieving 3X speedups on Google TPUs for LLM inference by parallelizing token generation.
Guide on effective AI-augmented workflows emphasizing artifact accumulation and iterative system improvement through correction-based learning.
Code review tool announcement noting free tier for OSI-licensed repositories.
Probus vulnerability scanner using three-agent system to identify security flaws, with PRs merged into Vercel AI SDK, n8n, and LangGraph.
Unity AI editor integration beta offering model connectivity and performance optimization for game development workflows.
AI agent pipeline monitoring 28+ sources to surface emerging AI tools, marketing tactics, and startup launches in weekly digest.
Discussion of LLM hallucinating tool calls in MCP-based agentic systems, causing token waste and requiring extensive prompt engineering fixes.
ClankerView: AI agents autonomously browse web applications and provide UX/design feedback based on user experience evaluation.
Compilation-stage knowledge layers proposed as successor to RAG architecture for improving LLM knowledge integration.