Show HN: AgentPass – Identity layer for AI agents (passports, email, trust)
AgentPass provides cryptographic identity and authentication infrastructure for autonomous AI agents. Identity layer enabling agent autonomy.
AgentPass provides cryptographic identity and authentication infrastructure for autonomous AI agents. Identity layer enabling agent autonomy.
Relay product for efficient agent context management via ephemeral/durable classification. Token optimization for autonomous agents.
Standard for tracking human-AI creative control attribution in code generation. Developer tool for documentation and collaboration.
vLLM optimization achieving 26.2K prefill and 10.1K decode throughput on NVIDIA Blackwell for DeepSeek MoE models. Production LLM serving research.
Benchmark for evaluating LLM vision and tool-use on cursor control task based on Neuralink's Webgrid test. Measures multimodal agent capabilities.
Multi-agent autonomous workflow built Game Boy GBA emulator in 48 hours using TypeScript. Demonstrates agent orchestration and tool use.
DataClaw tool exports Claude Code conversation history to HuggingFace datasets. Data extraction tool with open source intent.
News article about Anthropic accusing Chinese AI firms of model distillation for IP theft. Industry news, not technical content.
LLM-based system for extracting clinical phenotypes from medical notes and standardizing to Human Phenotype Ontology terms for rare disease diagnosis.
Framework combining LLM-based semantic analysis of metadata with statistical causal discovery via conditional independence testing.
Diffusion models for trajectory generation in offline reinforcement learning, improving consistency between generated transitions and environment dynamics.
Benchmark for evaluating AI agents on implicit requirements (accessibility, privacy, risks) beyond explicit instructions, addressing real-world underspecified requests.
Research on improving tool descriptions for LLM agents. Addresses tool interface optimization as bottleneck for reliable agent tool use.
PreScience benchmark for forecasting scientific advances using AI. Tests if models trained on historical research can predict future collaborators, impactful directions, and emerging problems.
KairosVL integrates time series analysis with semantic reasoning using reinforcement learning for complex temporal prediction tasks combining numerical and contextual understanding.
ActionEngine framework enables GUI agents to operate programmatically via state machine memory instead of reactive step-by-step vision-language model calls, reducing costs and improving accuracy.
Research on imitation learning for AI agents to exhibit diverse human-like behaviors and adapt to contexts. Uses inner speech mechanisms for steering agent behavior in human-AI coordination tasks.
Data-centric framework that learns optimal verbalization of user interaction logs for LLM-based recommendation systems, replacing rigid template concatenation.
CausalReasoningBenchmark evaluates causal inference systems separately on identification and estimation steps across 173 real-world queries and 138 datasets.
Characterizes cross-modal bias in multimodal LLMs from algorithmic fairness perspective, examining comparative and non-comparative fairness contexts.
Examines safety guarantees for untrusted monitoring deployments where misaligned AI models oversee each other across different collusion strategies.
Proposes elicitation procedure for identifying two piecewise linear additive value functions from anonymous preference pairs without knowing decision-maker correspondence.
EmbodiedAct framework grounds LLM scientific reasoning in physical simulation with runtime perception to detect anomalies and enable interactive discovery.
Recursive Belief Vision Language Model (RBVLM) maintains hidden state for long-horizon manipulation under partial observability, reducing action repetition and inference latency.
Evaluates foundational skills required for VLM-based embodied agents in native control settings with both low and high-level task benchmarks.
PromptCD applies contrastive decoding at test time to improve LLM alignment with human preferences without additional training data or computational costs.
Introduces online algorithms with unreliable guidance (OAG) model for ML-augmented online decision making that separates predictive and algorithmic components.
ICON is an inference-time defense framework for LLM agents against indirect prompt injection attacks that mitigates threats while preserving valid agentic workflows.
Counterfactual Simulation Training (CST) improves chain-of-thought faithfulness by rewarding reasoning paths that accurately predict model outputs versus counterfactual scenarios.
Introduces BAPO, an off-policy reinforcement learning framework with verifiable rewards that improves data efficiency in LLM post-training through adaptive buffer management.
Proposes mixture of graph experts with entropy-triggered routing for multimodal recommendation systems to address modality imbalance in sparse feedback settings.
Uses RL from AI feedback (RLAIF) with LLMs to generate preference labels for multi-objective urban traffic control, reducing reliance on human annotation.
CHESS proposes context-aware token pruning for long-context LLM inference that reduces KV cache overhead while maintaining quality through step-wise relevance and semantic awareness.
PyVision-RL is an open-source reinforcement learning framework for training multimodal agentic models that prevents interaction collapse and sustains multi-turn reasoning via tool rewards.
Pipeline for automatic and interactive verification of LLM-generated mathematical solutions as alternative to answer-only evaluation, enabling formal and informal solution generation.
POMDPPlanners is an open-source Python package for evaluating POMDP planning algorithms with benchmarks, hyperparameter optimization, and parallel simulation capabilities.
Qwen-BIM proposes an LLM fine-tuned for building information modeling with domain-specific benchmark and dataset to improve performance on BIM-based design tasks.
Alignment benchmark with 904 realistic multi-turn scenarios across six categories evaluating LLM behavior under pressure including honesty, safety, and scheming.
Vision-Language Causal Graphs enable structured evaluation of causal reasoning in LVLMs, distinguishing between reasoning failures and spurious correlations.
Study examines how visual context affects human and LLM predictions of sentence acceptability compared to textual context.
HELP improves GraphRAG efficiency and accuracy by expanding hypernodes and using logical path-guided evidence localization for multi-hop reasoning.
AgentOS is a conceptual framework redefining LLMs as dynamic autonomous cognitive systems bridging token-level processing to system-level intelligence.
LogicGraph benchmark evaluates LLMs on multi-path logical reasoning with neuro-symbolic generation and verification beyond single correct proofs.
Benchmark measuring step-success probability for LLM superintelligence via test-time search using GF(2) circuit reconstruction tasks.
Novel training paradigm inspired by affective neuroscience using motivation states to activate a larger model intermittently alongside continuous base model training.
Paper addresses the initial exploration problem for lay users encountering unfamiliar Knowledge Graphs without semantic web expertise.
DEEPSYNTH benchmark evaluates LLM-based agents on complex tasks requiring multi-source information synthesis and tool use like web browsing and code execution.
CG-DMER combines contrastive and generative learning for multimodal ECG representation learning with clinical reports.
NoRD is a data-efficient Vision-Language-Action model for autonomous driving that achieves competitive performance without dense reasoning annotations.
Aletheia, an AI agent powered by Gemini 3 Deep Think, autonomously solved 6 out of 10 FirstProof mathematical problems.