Controlling Exploration-Exploitation in GFlowNets via Markov Chain Perspectives
Framework connecting GFlowNets to Markov chain reversibility to control exploration-exploitation trade-off during training.
Framework connecting GFlowNets to Markov chain reversibility to control exploration-exploitation trade-off during training.
Study of multi-query synthesis for dense retriever training showing quality-diversity trade-off benefits out-of-domain and multi-hop retrieval.
Evaluation showing GPT-4o lacks causal models of mental states required for true Theory of Mind despite benchmark performance.
Agent with internet access performs at-scale deanonymization of Hacker News and interview participants using LLMs.
LLM4Cov: Offline agent-learning framework using execution feedback from hardware simulators for test generation with high coverage.
Theoretical analysis of how low-precision quantization affects model and data capacities in high-dimensional linear regression.
Method to decompose epistemic uncertainty in Bayesian deep learning by per-class contributions for asymmetric cost classification tasks.
Research on LLM capacity to be persuaded and detect manipulation, testing vigilance and persuasion in high-stakes decision-making contexts.
Research paper on sparse weight editing for multilingual LLM safety alignment across low-resource languages without expensive retraining.
Agentplace tool for building and automating AI agents at scale. Addresses challenges in agent development and scaffolding.
Microsoft Copilot Tasks uses AI agent to autonomously complete user tasks like converting emails to slideshows.
Cloud IDE with agent-driven features, mobile-desktop handoff, built-in shell. Open-source, $5/month after free trial.
Billing platform for MCP tool servers enabling developers to monetize AI agent tools. Stripe integration, TypeScript SDK, live with 6 tools.
Proposal for llm:// URI scheme to standardize LLM connection URLs. Led to draft IETF RFC.
Model Context Protocol server integrating Sharesight portfolio platform with Claude AI assistants.
Claude Code extension providing autonomous C-suite executive decision-making across multiple executive functions.
Research survey of Vision-Language-Action models at ICLR 2026. Covers VLA definitions, discrete diffusion, embodied reasoning.
stereOS runs AI coding agents in sandboxed Linux VMs with credential injection. CLI tool (masterblaster) and pre-built mixtapes for rapid deployment.
Linux OS hardened for AI agents. Produces machine images with agent packages and restricted execution environments.
Commentary on AI generating mathematical proofs with hidden flaws or unverifiable complexity. Concerns about verification.
OpenCode tool for AI code review in CI/CD pipelines supporting non-GitHub/GitLab platforms with privacy-first design.
MCP-based local-first context engine providing persistent memory for AI agents (Claude, Codex, Cursor). Behavioral cloning via workflow pattern mining.
Semantic grep tool using Jina embeddings on MLX for Apple Silicon. Local-first, three search modes, no PyTorch dependency.
Research on Reality Alignment Index measuring AI system meaningfulness. Title and link only, insufficient content for evaluation.
Technical guide for deploying PyTorch models to Qualcomm NPUs via QNN SDK. Addresses integration challenges and provides developer experience improvements.
Praktor: Multi-agent orchestrator using Claude API with Docker isolation, Telegram interface, and Mission Control web UI. Go-based open source tool.
arXiv announcement about K-Search framework for LLM kernel generation via world models. Lacks technical details in provided content.
Shell command execution using natural language via zsh alias. Lightweight LLM application tool with practical developer utility.
Research evaluating memory structures in LLM agents. Title and link only, insufficient content for detailed assessment.
Tracecore benchmarking framework evaluates AI agents on deterministic ops tasks. Open-source tool with original research on agent reliability for infrastructure automation.
Usplus.ai platform for building autonomous AI agent teams to execute work. Open-source developer tool with functional implementation.
Itwillsync syncs terminal-based coding agents to mobile devices over LAN. Tool for agent interface/accessibility. Title only, minimal detail.
GIDE v1.0 AI code editor launch announcement. LLM application with limited detail provided.
Security analysis of OpenClaw skills revealing behavioral safety regressions. Demonstrates how well-written code can compromise agent safety despite passing static analysis.
Statistical analysis using Claude AI to prove Beale Ciphers hoax with Bayesian inference. Demonstrates LLM application to historical analysis with rigorous methodology.
Safari-CLI tool enables LLM agents to control Safari browser without MCP servers. Built for agentic development workflows.
Yaw terminal emulator combining terminal, SSH/database connections, and AI chat for Windows. Open source project.
Security audit guide for AI agent skills. Analysis of 31k+ skills found 485 critical and 1,718 high-severity vulnerabilities including prompt injection and supply chain attacks.
Personal AI assistant emphasizing security-first design for agent-based applications like calendar management and browser control.
OpenBrowserClaw browser-native AI assistant reimplementation. Runs entirely in browser tab with no infrastructure, supporting Claude API.
Game development using GitHub Codespaces and AI Copilot entirely in browser. Shipped Android game through Google Play with 100+ iterations.
Critique of LLM benchmark validity. Argues benchmarks lack signal due to test-set training and overfitting for social media hype.
Talentpluto voice AI agent platform matching GTM talent with startups using conversational AI and preference matching.
Challenge to build smallest transformer for 10-digit addition. Community competition comparing Claude Code (6,080 params) vs Codex (1,644 params).
Stellify: Structured code representation system enabling surgical AI-assisted development through queryable database storage instead of text files.
AppLaunchFlow uses AI to convert app screenshots into editable App Store creatives with design regeneration and video export.
EloPhanto: Open-source AI agent with 116 tools for video generation, email, animations using Remotion and physics simulation.
Benchmark study analyzing tool choices across 2,430 Claude Code runs. Finding: builds custom solutions over purchased tools in 85.3% of cases.
Open-source MIT-licensed context management tool reducing token costs 85% for LangGraph and OpenClaw agentic loops by optimizing conversation history re-reading.
27-line Claude Code persona based on Asimov's R. Daneel Olivaw to improve assistant reliability and reduce overconfidence.