Institutional Red-Teaming: Deployment Rules, Not Just Models, Causally Shape Multi-Agent AI Safety
Institutional red-teaming methodology for testing deployment rules in multi-agent AI systems with IABench-CA benchmark spanning 228 contexts.
Institutional red-teaming methodology for testing deployment rules in multi-agent AI systems with IABench-CA benchmark spanning 228 contexts.
Investigates whether model-free RL agents can identify price manipulation more effectively than model-based approaches in asset markets.
Analyzes memory poisoning attacks on LLM agents with long-term memory that access emails, calendars, and code repositories.
TriRoute jointly optimizes attention resolution, expert selection, and KV-cache allocation in language models to decouple quality from inference cost.
Mixture-of-experts approach for long-term forecasting addressing dataset-level distribution shifts via regime modeling.
Horizon-scanning study identifying security and privacy challenges in agentic AI systems.
Framework for optimizing diffusion model sampling via dynamic preference learning on schedules and guidance.
Open foundation model for wearable motion sensing with systematic study of pretraining and scaling.
LLM-guided approach for time-series forecasting in industrial processes using semantic metadata.
LLM-based adaptation method for specialist models in industrial processes without retraining.
Dynamic computation framework combining adaptive architecture with few-step distillation for video generation.
Study on how specification-grounded test generation improves LLM code quality and edge case handling.
Retrieval-augmented generation system for public health question answering reducing LLM hallucinations through corpus-grounded responses.
Graph matching method using diffusion-enabled optimal transport for comparing graphs with sparse or noisy node features and structure.
Library providing block/residual/sieve resampling and conformal prediction methods for time series uncertainty quantification and bootstrap confidence intervals.
Review of Vision Language Action models enabling robots to follow natural language instructions for manipulation and aerial tasks.
Evaluation framework for LLM-based software engineering agents addressing fragmentation and developer alignment in autonomous development contributions.
Continual learning framework for adaptive control of modular soft robots with deformable and reconfigurable structures.
LLM-based tool for automated detection and repair of YAML configuration errors in smart home automation platforms.
Standards perspective on agentic AI and large AI models for autonomous 6G network management with runtime software evolution capabilities.
Knowledge distillation approach for compressing state-of-the-art deep learning time series models for resource-constrained deployment.
Study of signals predicting correctness in text-to-SQL generation using self-consistency and execution-based confidence metrics on BIRD and Spider benchmarks.
LLM pipeline workflow for discovering detector rules across 68 physiological datasets for contactless health monitoring platform design.
Security framework for detecting malicious behaviors in LLM-based multi-agent systems through activation-based detection of semantic attacks.
E-commerce ad headline generation using reinforcement learning policy gradient methods on masked language models with self-critical training.
Gradient-based method for speech-to-text alignment compatible with CTC, transducers, attention-based encoders, and speech LLMs.
Study using rule-based expert baseline to evaluate reinforcement learning agents in imperfect-information card game Gin Rummy across 100+ trained agents.
Visual robot navigation policy using frozen multimodal LLM with low-rank adaptation for waypoint navigation without custom encoders or large training datasets.
Framework for explaining deep learning image classifier decisions at scale using local-to-global relevance analysis on large datasets.
Theoretical framework modeling AI-augmented computation as interaction between probabilistic Turing machines and stochastic oracles.
Parameter-efficient fine-tuning method using spatially-aware low-rank adaptation for vision foundation models to reduce computational costs.
Multi-factor scoring system for comprehensive evaluation of LLM responses across accuracy, consistency, and readability.
Survey of LLM and generative AI security applications, covering dual-use risks, malware generation, and defensive strategies.
FRAMe: end-to-end LLM flight planning system using RAG memory and multi-modal coach agents for eVTOL aircraft.
Hybrid least squares/gradient descent optimization method for accelerating MIONet training.
WAM-TTT: test-time training framework for adapting robot foundation models using human video demonstrations.
Gimitest: open-source framework for comprehensive testing of reinforcement learning policies across environments and algorithms.
AnchorPrune method for efficient visual token pruning in vision-language models balancing relevance and diversity.
Systematic analysis of training dynamics and solutions in deep feedforward ReLU networks.
Progressive crystallization framework converts agent exploration into deterministic, cost-effective production workflows through lifecycle stages.
arXiv paper on proprioceptive grounding mechanism (GeoProp) for vision-based robotic manipulation policies.
arXiv paper applying tree-of-thoughts reasoning to improve text-to-image in-context learning with multimodal LLMs.
arXiv paper on entropy pacing for multi-task reinforcement learning with agentic LLMs to handle varied exploration dynamics.
arXiv paper on pre-deployment safety evaluation of LLMs by simulating realistic deployment with de-identified conversations.
arXiv quantitative analysis of AI industry restructuring 2026-2030 examining memory constraints, open models, inference economics, and compute markets.
arXiv paper on scalable graph generation via diffusion models using graphon theory for dense graphs.
arXiv paper on unsupervised time series clustering using Mamba-based multiview contrastive learning framework.
arXiv paper on multi-fidelity Bayesian optimization framework for genetic algorithm hyperparameter tuning with neural surrogates.
arXiv paper studying data extraction attacks in federated learning using correlation encoding and segmented aggregation.
arXiv paper on uncertainty estimation in hypergraph neural networks using stochastic differential equations.