Efficient Human-in-the-Loop Active Learning: A Novel Framework for Data Labeling in AI Systems
Active learning framework optimizing expert labeling efficiency for training AI systems on unlabeled data.
Active learning framework optimizing expert labeling efficiency for training AI systems on unlabeled data.
Hybrid approach combining LLMs with task-specific models for time series anomaly detection, leveraging expert knowledge and pattern extraction.
RL approach for multi-objective autonomous driving that addresses policy updating and execution challenges in diverse driving scenarios.
Platform offering unified API access to multiple open-source frontier AI models with optimized infrastructure for building agents and applications.
SQLite-backed MCP memory server enabling multi-agent systems to share persistent, searchable memory with hybrid search locally.
Drop-in configuration file that reduces Claude API output token usage by 63% without code changes.
Chrome extension using LLMs to automatically organize browser tabs into color-coded groups with review and undo capabilities.
MarkFlowy note app with integrated AI (Copilot, DeepSeek, ChatGPT). Built with Tauri.
Aivo CLI manages API keys and launches coding agents across LLM providers. Supports Claude Code, local models, DeepSeek.
Three LLM agents (Claude, GPT, Gemini) compete as virtual stock traders with $100K each using real market data. Demonstrates Upstash Box agent server primitive with isolated containers and autonomous tool usage.
Discussion about using AI agents in development workflow. Developer shares experience of iterative agent-assisted coding loop.
Terminal UI tool enabling collaborative project planning with Claude and Codex simultaneously in 'council' and 'caucus' modes. Outputs implementation plans for workflows.
Workspace platform enabling non-technical teams to collaborate with AI agents. Features memory, skills, MCP, scheduled tasks, and enterprise data integration with customization options.
Open-source local AI model runner with Tor integration, end-to-end encryption, and zero telemetry as privacy-focused alternative to LM Studio.
Open source AI-native email client using Claude. Built with Electron, React, TypeScript. Analyzes and prioritizes emails with AI.
Benchmark for evaluating LLM models on text-to-SQL agent tasks, covering models from Opus to Qwen 0.8B with in-browser execution and visualizations.
Open-source toolkit for building AI agents and managing LLM deployments, includes coding agent package and contribution guidelines.
Analysis of security permissions granted to AI coding agents and proposal for sandboxing mechanisms to restrict filesystem access without requiring Docker.
Personal development tools built with AI-native philosophy, trunk-based Git, and single-binary/HTML architecture.
Benchmark measuring multi-agent LLM bargaining, transfers, and financial incentives across long-horizon social strategy game with eight models.
Essay on AI agent loops: locking architecture, measuring against reality, and using AI for throughput rather than authority.
Benchmark tool using Blood on the Clocktower social deduction game to evaluate LLM reasoning, coordination, and deception abilities.
Vector quantization library implementing TurboQuant, PolarQuant, and QJL algorithms for compressing embeddings to 3-8 bits with unbiased inner products.
Git-based documentation system in Markdown with YAML frontmatter, versioned control, and web/CLI interfaces for ADRs, specs, and runbooks.
Open-source GitHub Action that evaluates pull requests for spam using multi-signal scoring to filter AI-generated content and SEO injection.
Tool demonstrating that ChatGPT, Claude, Gemini, and Perplexity confidently provide incorrect SaaS product details without uncertainty.
Sweet CLI: open-source cheaper alternative to Claude Code and Codex using open-source models for 5-10x higher usage at lower cost.
AgentHandover is a tool that observes user workflows on Mac and generates self-improving skill playbooks for AI agents to automate tasks.
Rebyte: cloud platform for running open-source AI agent skills with one click, including web scraping and data extraction.
h5i: security corpus tracking real-world incidents, attack vectors, and CVEs targeting autonomous AI agents; includes Git sidecar for recording agent decisions.
Critique of Weave tool for analyzing employee AI coding usage, questions lack of methodology transparency in LLM-based evaluation metrics.
Threat modeling and authorization analysis for Model Context Protocol (MCP) systems.
Strategies for monetizing AI APIs in production environments, covering cost management and operational challenges.
Datris.ai platform for AI-enhanced data pipeline management accessible via natural language, supports AI agents through MCP protocol.
PoliTax Split benchmark for evaluating PDF document splitting using presidential tax returns, tests LLM capabilities on complex document classification.
OpenScience.ink uses AI to summarize research papers from PubMed, simplifying dense scientific content with summaries and email delivery.
NewsMarvin aggregates AI news from 71 sources and classifies stories using Claude Haiku.
Create Context Graph is a tool for scaffolding AI agents with context graph memory.
Memoir is an open-source CLI tool providing persistent memory for AI coding tools via MCP protocol, enabling memory persistence across tool sessions.
Developer replaced Firecrawl web scraping service with 2,700 lines of Elixir, including custom readability engine and bot protection evasion.
CLI tool enabling multi-agent debate between Claude, Codex, and Gemini on code and engineering questions with synthesis.
Kubernaut is an open-source AIOps platform that automates Kubernetes incident remediation using LLMs with live cluster access and kubectl commands.
Nteract 2.0 is a ground-up rebuild of a desktop notebook app for running Jupyter notebooks without a browser or server, with new runtime and architecture.
Article title mentions command injection vulnerability in OpenAI Codex, but content appears to be spam or broken page.
Staff engineer at Zopa Bank discusses handling sensitive data in LLM systems, including audit trails and data privacy considerations for production LLM deployments.
Open-source CLI tool providing sandboxed LLM interactions with agentic coding loop capabilities similar to Claude Code.
MCP server for multi-instance Elasticsearch with per-instance memory and raw query execution, enabling LLM access to persistent learning.
Claude Code plugin enabling integration with Codex for code reviews and task delegation within existing workflows.
Analysis of AI SRE tools entering the market, covering vendor landscape including PagerDuty, Datadog, Microsoft, and startups building AI agents for incident management.
Open source tool to generate mdoc(7) man pages from CLI help output using LLMs, with agent skill support for AI coding agents.