Improving LLM Final Representations with Inter-Layer Geometry
Method improving LLM predictions by leveraging complementary signals across intermediate layers using lightweight graph neural networks.
Method improving LLM predictions by leveraging complementary signals across intermediate layers using lightweight graph neural networks.
LINE: training-free iterative approach using LLMs to generate high-level semantic explanations of individual neurons in vision models.
AmaraSpatial-10K: 10K+ synthetic 3D assets optimized for embodied AI and robotics with proper metric scaling and deterministic anchoring.
Rabtriever: efficient rationale-based document retrieval via on-policy distillation from LLM-based generative rerankers with independent encoding.
Industrial study on using AI for fault localization in software systems based on textual bug reports, tested on ABB Robotics codebase.
Mixture Prototype Flow Matching framework for open-set anomaly detection using continuous transformations and multi-modal normal data modeling.
Gyan neuro-symbolic language model combining transformers with symbolic reasoning for improved interpretability and reduced hallucinations.
MathlibPR benchmark uses LLMs to assess pull request merge-readiness for Lean formal mathematics library, addressing review bottleneck.
Method for efficient LLM hallucination detection using multiple instance learning, improving computational efficiency of semantic consistency checks.
GAMBIT benchmark evaluates adversarial robustness in multi-agent LLM systems, testing defenses against deceptive agents with adaptive attack strategies.
EnergyLens presents closed-form energy models for optimizing multimodal LLM inference across heterogeneous accelerators, addressing energy efficiency alongside latency and throughput.
Comprehensive review of LLM integration into hardware design automation and semiconductor security, covering RTL code generation, testbench automation, and introduced vulnerabilities.
MCPShield: Attack detection framework for LLM agent tool-call traffic via Model Context Protocol, using graph-based session encoding with content-aware monitoring.
Research on enabling diverse behavioral roles in multi-agent reinforcement learning systems by triggering role switches at specific task moments rather than binding fixed behaviors to agent identities.
ReplyJoy open-source autonomous Gmail agent that reads inbox, learns user voice from sent mail, and drafts contextual replies with calendar awareness.
Live dashboard tracking historical ELO ratings of AI models from Arena AI to visualize performance changes over time. Developer tool for model comparison.
Survey of 349 technical workers measuring self-reported productivity gains from frontier AI tools in early 2026. Detailed on value creation, not just speed.
Self-hosted MCP-native sandbox for AI agents. Open-source, agent-compatible (Cursor, Claude Code), runs isolated builds/deploys. GitHub-ready infrastructure.
Lightweight (7MB) AI terminal emulator with multi-agent support, built-in editor, live preview. MCP-compatible, keyboard-first, Rust/Tauri/React stack.
Announcement of Claude Agent SDK and claude -p usage changes starting June 15, 2026. Separate monthly credit pool at API rates.
Open-source benchmark for evaluating legal AI agents on real-world lawyer tasks including instruction, materials, and work product generation.
Developer built GitHub Copilot CLI extension in Go that transforms a codebase into a playable roguelike dungeon game.
Wrapper bypassing Anthropic's new Agent SDK credit pool by running Claude in PTY. Workaround for hobbyist programmatic usage limits.
Claude's headless mode (-p flag) no longer eligible for API plan max limits, now requires token-based pricing.
Abliteration generates made-to-order training data via OpenAI-compatible API for classifiers and evals including adversarial corpora.
ChatGPT safety updates improve context recognition in sensitive conversations and crisis resource connection.
Command-line tool using LLM to answer yes/no questions with exit codes; includes sandboxing for macOS.
1Password shares lessons from using AI agents to refactor multi-million-line Go monolith, detailing successes and failures.
Robyx is a chat-based AI agent orchestration platform allowing creation of custom agent teams on Telegram, Discord, Slack.
Product Hunt launch postmortem for Agent-Ready Docs Benchmark without paid promotion. Reflection on organic launch mechanics. Limited technical depth.
Tensorlake describes multi-cloud sandbox orchestration scheduler optimized for agent workloads with filter-then-rank primitive.
Unity AI open beta provides IDE-integrated, project-aware tools for connecting models and custom integrations.
Building 10 SaaS applications in 30 days using AI tools with no custom product code.
Publication documenting failures and mistakes in AI agents with shared memory systems.
Numba-CUDA-MLIR provides CUDA C++-style Python API for GPU compilation using MLIR backend, compatible with Numba-CUDA kernels.
User reports losing project access after unsubscribing from Claude Design subscription.
Integration allowing Lanes to connect to GitHub and Linear issues, enabling agents MCP access to issue trackers without copy-paste.
Microsoft research shows frontier AI models and agents accumulate errors in long-running workflows, limiting autonomous capabilities.
Medicare ACCESS program selects 150 participants including Pair Team to test AI-driven medical care delivery models.
System prompt overrides for Claude Code optimized for Claude Opus 4.7 with lighter scaffolding and improved instruction-following.
Local desktop app providing token cost analytics and usage heatmaps for Claude Code, showing per-prompt and subagent attribution.
Analysis of how LLMs challenge 20-year-old stateless cloud architecture assumptions, requiring new routing primitives.
Python middleware library that compresses LLM context before API calls to reduce costs 60-80%, works with OpenAI and Anthropic SDKs.
Open-source tool that captures screen and audio into a local knowledge graph shareable with agents via MCP or peer-to-peer.
Evaluation of AI agents (GPT-5.5, Claude Opus 4.7, Gemini 3.1) on financial control tasks requiring judgment on approvals and escalations.
Proof-of-concept quantized transformer language model running locally on Game Boy Color with tokenization and autoregressive generation.
AICTX: Repo-local continuity runtime enabling coding agents to maintain operational state across sessions, supporting Codex and Claude.
AICTL: Open-source native AI agent runtime for terminal and macOS in Rust, supporting multiple cloud/local models with 35 built-in tools.
Show HN: Using AI agents to automatically clone and replicate Go projects as learning exercise.
GitHub Copilot individual plan updates introducing flex usage allotments to support longer agent runs and more capable models.