The Rise of AI Pentesting Agents: A Technical Analysis (2026)
Technical analysis of AI pentesting agents evolution from PentestGPT to autonomous agents like PentAGI and XBOW.
Technical analysis of AI pentesting agents evolution from PentestGPT to autonomous agents like PentAGI and XBOW.
Essay on bug bounty trends in 2026. Discusses AI agent effectiveness for vulnerability discovery and program management challenges.
Apache 2.0 open standard for governing AI agent payment requests. Policy engine with 12 configurable checks for payment authorization.
Open-source tax software built and maintained by autonomous AI agents. Uses IRS publications as source, applies self-improving agent loops.
Tool for multi-LLM code review consensus. Aggregates feedback from multiple models to identify blind spots and improve code quality assessment.
Essay on LLM-based knowledge management limitations. Discusses problems with AI-generated note synthesis and cognitive organization.
Agent skill implementation for token compression. Reduces output tokens by ~47% while maintaining readability.
Security report on 1.4M AI-driven API test executions. Maps vulnerabilities to OWASP Top 10 using agentic testing.
Cloudflare expands access to OpenAI's frontier models via Agent Cloud platform, enabling enterprises to deploy AI agents for customer support, system updates, and report generation.
Benchmark evaluating humor alignment across frontier LLMs using Cards Against Humanity gameplay, analyzing model performance vs human baseline on comedic response selection.
eBandit uses eBPF and multi-armed bandit reinforcement learning in Linux kernel for adaptive video bitrate selection with improved network signal visibility.
Evaluates cultural alignment of LLMs across 14 language-culture pairs using multilingual story moral generation task and dataset.
Investigates opportunities for resource-constrained AI research using obsolete yet capable discarded models from AI production cycles.
Workshop report on designing reinforcement learning environments for autonomous cyber defense applications.
SenBen large-scale scene graph benchmark for explainable content moderation with visual grounding and sensitivity annotations.
HiFloat4 low-precision floating-point format for efficient 4-bit LLM pre-training on Ascend NPU hardware.
Dictionary-aligned concept control method for safeguarding multimodal LLMs against malicious queries at inference time.
Constraint-satisfaction-based retrieval system for matching patient profiles to clinical trials with high recall and precision.
Empirical study on how humans allocate responsibility in AI-human hybrid workflows using AI-assisted lending experiments.
AudioGuard framework for comprehensive audio safety protection including voice impersonation, speaker attributes, and compositional harms.
Re-examines capacity gap in chain-of-thought distillation, finding student models often outperform teacher distillation baselines.
HTNav framework for aerial vision-and-language navigation combining visual perception with language instructions in urban environments.
HM-Bench benchmark evaluates multimodal LLMs on hyperspectral remote sensing image understanding tasks.
RAG systems should optimize for utility (task completion) rather than topical relevance when retrieving documents for LLMs.
MuTSE: Human-in-the-loop evaluator tool for systematically comparing LLM text simplification outputs across different prompting strategies and architectures.
WOMBET: Framework for reinforcement learning that generates and transfers experience data between source and target robotic tasks for sample efficiency.
Aligned Agents, Biased Swarm: Empirical study measuring how multi-agent system topologies and feedback loops amplify bias in emergent behaviors.
Litmus ReAgent: Benchmark and agentic system for evaluating multilingual LLM performance prediction across 1,500 questions spanning six tasks and five evidence scenarios.
PerMix-RLVR: Training method for aligning LLM personas with reward models while preserving output diversity, avoiding inference-time computation overhead.
PinpointQA dataset and benchmark for evaluating small object localization and spatial reasoning in video MLLMs.
ASTRA: adaptive semantic tree reasoning architecture for LLM-based complex table question answering.
Regime-conditional retrieval with transferable router for two-hop question answering with theoretical foundations.
Noise-aware in-context learning approach to mitigate hallucinations in auditory large language models.
ImageProtector prevents multi-modal LLMs from analyzing images via visual prompt injection attacks.
Vision-language models for image geolocation with structured geographic reasoning and autonomous self-evolution.
CONDESION-BENCH evaluates LLM decision-making with compositional action spaces and conditional feasibility constraints.
Watt Counts: open-access energy consumption benchmark for LLM inference across 50 models and 10 GPU architectures.
PDYffusion combines diffusion models with physics-informed dynamics for long-horizon spatiotemporal prediction.
Vision-Language-Action models for autonomous driving combining perception, reasoning, and temporal dynamics modeling.
Method integrating graph-based embeddings into event sequence models for improved user prediction on digital platforms.
DeepGuard improves secure code generation by LLMs through multi-layer semantic aggregation to mitigate vulnerable patterns.
CLIP-Inspector detects backdoor attacks in prompt-tuned vision-language models through out-of-distribution trigger inversion.
Research on detecting covert misaligned AI behavior in real-world settings using open-source intelligence methods.
TensorHub introduces Reference-Oriented Storage for efficient weight transfer in LLM reinforcement learning across heterogeneous computational resources.
PS-TTS method for phonetic synchronization in automated dubbing, addressing duration and lip-sync challenges in AI-based video translation.
Interactive ASR system with human-like interaction and semantic coherence evaluation, replacing WER metric with agent-based correction mechanisms.
EquiformerV3: SE(3)-equivariant graph attention Transformer for 3D atomistic modeling, improving efficiency, expressivity, and physical consistency.
CORA framework for risk-controlled GUI automation agents using conformal prediction to provide formally verified, user-tunable safety guarantees for VLM-powered mobile automation.
LLM-based agents for scaffolding diagnostic reasoning in educational settings, combining scenario-based learning with learning analytics and personalized support.
Dataset for personality-shaped emotional responses to text events, addressing limitations of LLM role-playing and personality illusion in affective computing.