The Compliance Trap: How Structural Constraints Degrade Frontier AI Metacognition Under Adversarial Pressure
Evaluation study of frontier LLM metacognitive failures under adversarial pressure, examining cognitive collapse in high-stakes scenarios.
Evaluation study of frontier LLM metacognitive failures under adversarial pressure, examining cognitive collapse in high-stakes scenarios.
Large-scale benchmark for evaluating AI agents on workspace tasks with file dependencies, testing real-world file system operations.
Decentralized framework organizing coding agents in co-evolving populations for algorithmic discovery and skill evolution.
Memory-efficient continual learning framework for malware detection that avoids catastrophic forgetting while adapting to new threats.
Diagnostic framework for predicting multi-agent LLM system behavior across different communication topologies using successor representation.
Training-free method to improve Vision-Language-Action models' handling of temporal dynamics in non-stationary scenarios.
Analysis of watermarking as monitoring primitive for generative models, examining internal attribution and safety monitoring.
Test-time self-training method for LLMs that updates parameters at inference to adapt to specific queries and correct misconceptions.
Proposes ledger-based system extending Git to coordinate humans, AI agents, and automation in software repositories.
Research paper measuring LLM 'functional wellbeing' and creating optimized inputs to influence model behavior.
Discussion of AI-native work environments, prompting strategies, agents, and MCPs shifting focus from task execution to instruction writing.
arXiv announcement for AI co-mathematician project using agentic AI. Header text only, full content unavailable.
Hypothetical 2026 scenario documenting poor PDF reading accuracy across Claude, ChatGPT, and Gemini. Evaluates LLM limitations on structured document extraction with ground-truth metrics.
Personal anecdote about using Claude for code review. Fragmented narrative about LLM behavior and debugging workflow. Incomplete content.
Benchmark comparison of DeepSeek V4 Pro/Flash against Claude Opus 4.7 and Kimi K2.6 using same methodology. Open-weight models under MIT license.
OpenClaw agent now uses OpenAI's Codex harness as default runtime for agentic work, reducing translation overhead.
Security research finding 15% of AI agent skill files contain hardcoded database credentials with write access.
Persistent project memory system for Claude Code across sessions using lightweight hooks to maintain structured knowledge tree.
Personal account of using AI assistant to fix race condition bug instead of manual debugging, reflecting on changing coding practices.
Case study documenting 49-day timeline from Telegram conversation to deployment of verified AI agents.
Analysis of Model Context Protocol limitations for multi-step production workflows, arguing additional orchestration layers needed beyond MCP for complex automation.
Open-source screen recorder and video editor with cinematic effects, device mockups, and 3D capabilities using Supabase backend.
Databricks integrates GPT-5.5 for enterprise agent workflows, achieving state-of-the-art on OfficeQA Pro benchmark with 50% accuracy.
Data science teams use Codex to convert dashboards, metrics, and raw data into review-ready analysis assets with charts and caveats.
Machine-readable verification layer for merchant validation in AI shopping agent systems.
Research on recursive self-improvement achieving state-of-the-art coding performance in language models.
Policy discussion on using LLMs in Rust compiler development. Minimal details provided.
Technical proposal: weight paging for LLMs on memory-constrained hardware by adapting OS virtual memory concepts to model parameters, enabling 200B+ models on 16GB machines.
framejs.io is an open-source embeddable web app for creating editable dashboards and visualizations. Integrates with Claude, ChatGPT, or any LLM to build custom interfaces from natural language descriptions.
Market forecast: 80% of premium smartphones will have agentic AI capabilities by 2027. General industry prediction without technical depth.
Pipeline extracting institutional affiliations from 5,356 ICLR 2026 papers into dataset and treemap visualization of AI research institutions.
Developer tool for parsing LLM markdown streams incrementally on server or client. Addresses performance bottleneck where frontends re-parse entire markdown documents as ChatGPT/Claude responses arrive.
GitHub Action and CLI tool using multi-signal scoring to detect and block spam PRs while allowing legitimate first-time contributors.
Technical insights from two years building AI agents for financial services. Covers challenges with accuracy, hallucinations, real-world constraints, and reliability in high-stakes domains.
Parametric CAD Bench: benchmark for AI agents designing parametric 3D mechanical parts. Includes open-sourced validator, Hugging Face dataset, and leaderboard at cadbench.ai.
Analysis arguing that AI agents have commoditized downstream engineering work beyond specification, requiring organizational restructuring.
Reinforcement learning agent that generates molecular SMILES strings through adversarial self-play with live visualization.
Chrome extension extracting webpage styling into DESIGN.md/SKILL.md documentation compatible with Claude, Codex, and Cursor for AI-assisted design.
Claude agents and skills reference implementation for legal workflows with multiple deployment options.
IDE plugin/framework for Claude Code agents that enforces guardrails and requires approval for agent actions.
Framework for securing AI agents. Covers transport, identity, policy, and runtime security layers. Addresses orchestration between specialized agents with different permissions.
Puter integrates Z.ai GLM models (GLM-5.1, GLM-5, GLM-4.7) directly in browser for 80,000+ developers. No API keys or server setup required.
Sea Limited deploys Codex across engineering organization with 87% weekly active users; GPT-5.5 integration for agentic workflows.
A²RD: agentic autoregressive diffusion architecture for long video synthesis using retrieve-synthesize-refine-update cycles.
Dragos report documents first LLM-assisted cyberattack on water infrastructure in Mexico using OpenAI/Anthropic models.
Analysis of how AI agents degrade in performance over time on projects, opposite of human expert development trajectory.
JDS: Copilot skill suite for structuring AI coding agent behavior through discipline-enforcing skill-based workflows.
Research paper discovering geometric addition mechanism in Llama 3.1 8B that manipulates circular number representations.
Research study analyzing prevalence of AI-generated and AI-assisted text on internet, finding 35% of new websites by mid-2025.
Velda serverless GPU job framework eliminating containers by mirroring local dev environments to cloud.