Browser made for AI coding agents – playwright or pencil.dev alternative [video]
Browser environment designed for AI coding agents, compared to Playwright and Pencil.dev.
Browser environment designed for AI coding agents, compared to Playwright and Pencil.dev.
Analysis of why AI agents have underperformed expectations in manufacturing applications.
Open-source transparent proxy for MCP servers enabling request tracing and latency measurement.
Open-source MCP server for Brazilian public bid data, works with Claude/Cursor/Continue clients.
Open-source infrastructure for deploying and managing fleets of AI agents in production with virtual environments, persistent memory, and isolated workspaces.
Interfaze model architecture claims superior performance on OCR, vision, STT tasks vs Gemini/Claude/GPT.
Analysis of transformer architecture convergence across 53 LLMs (2017-2025), identifying dominant design patterns like RMSNorm, RoPE, SwiGLU, and KV-sharing with trade-offs between optimization and practical constraints.
Investigation of 8 Chrome extensions with 7M+ installs that scrape AI chat content and send it to remote servers, many owned by data brokers.
Papel is a recommendation platform for academic papers with AI-powered search and question-answering on full PDFs.
Research showing LLM agent memory degrades when continuously updated via distillation, evaluated on ALFWorld and other benchmarks.
Essay on evolutionary algorithms for antenna design and implications of LLM-generated code for understanding AI outputs.
cuda-oxide: Nvidia's experimental Rust-to-CUDA compiler allowing safe Rust GPU kernel development with direct PTX compilation.
LockedCode is a security-hardened fork of OpenCode AI coding agent that adds sandboxing and access controls.
VibeServe uses AI agents to generate bespoke LLM serving systems tailored to specific model, hardware, and workload combinations.
Guide to understanding how AI benchmarks work, their limitations, and how to read benchmark results critically.
Opinion piece comparing coding professionalization to woodworking and industrialization trends.
Interactive deep learning textbook with implementations in PyTorch, JAX, TensorFlow, and MXNet, adopted by 500 universities.
CCL-Bench 1.0: Trace-based benchmark for evaluating LLM infrastructure and serving systems performance.
Discussion on managing rapid AI agent deployment in teams, balancing PoCs with production-ready RFC processes.
Discussion of LLM capabilities in software security, treating code analysis as language translation task for finding vulnerabilities.
Analysis of Copilot system instructions causing models to rush code generation, undermining thoughtful problem-solving in development.
TypeScript SDK for building AI agent skills as typed state machines.
Q1 2026 ChatGPT adoption data showing growth across demographics and geographies, excluding enterprise/education usage figures.
Hivemind tool that converts AI agent execution traces into reusable skills for team sharing.
Google report: AI-powered hacking scaled to industrial level using commercial models for attack development and exploitation.
Demo of Word2Vec nearest-neighbor search using HNSW on ESP32-S3 microcontroller with 3M vectors.
SLayer: semantic layer tool for connecting AI agents to databases, enabling data analyst chatbots and agentic applications with improved SQL management.
BrowserCode: open-source web app running Claude and Gemini agents in browser via WebAssembly with local execution and persistence.
TypeScript framework for building reactive AI agents. Limited detail provided in submission.
Open source project demonstrating LLM fine-tuning for GenZ communication style at low cost.
Technique using LLMs to hide secret text within plausible cover text via steganography.
JetBrains Junie coding agent tool for developers, AI-powered code generation and assistance.
Video demonstrating adversarial attack using Morse code to make an AI agent execute unintended high-cost transactions.
Column-oriented analytics extension for SQLite enabling efficient data analysis.
iOS app (Atrophy) designed for software engineers to self-assess potential over-reliance on LLMs in their work.
Maggy autonomous AI engineering platform with multi-agent orchestration, session memory, and code quality guardrails.
FLOX C++23 trading framework with MCP integration for AI-native development of trading systems.
Terminal UI tool for human-in-the-loop code review of AI-generated code changes.
Developer tool for calculating inference speed of local LLMs.
Analysis of running open-weight LLMs locally on edge devices and competitive advantages of edge AI deployment.
Technical writeup testing AI agent harness using Cursor and Kimi K2.5 to build apps from specifications with minimal manual coding.
Research on using LLMs to discover reinforcement learning interfaces, combining RL with language model capabilities.
Brief report on AI agents demonstrating improvements with long context models (LCM) and emerging specialized applications.
Deepfake detection tool that runs locally without sending files to cloud APIs.
Security proxy for AI coding agents at OS level. Addresses safety concerns in agentic IDEs.
Pure Rust inference engine for running machine learning models locally.
Story about Mac Mini hardware demand for Claude inference. Lacks technical detail.
Open-source API quota firewall for AI agents using Scala 3, Pekko, and PostgreSQL. Prevents budget bankruptcy via race-condition-safe ledger mechanics.
Essay on how Agile principles apply in AI-enabled software development, focusing on communication loops and expanded stakeholder models.
Framework proposing five patterns for effective AI-assisted development, drawing parallels from pair programming practices like onboarding and shared standards.