aiAuthZ: Off-Host, Identity-Bound Authorization for AI Agents
aiAuthZ security framework for identity-bound authorization in AI agents, evaluating LLM vulnerability to tool-call forgery attacks across 15 models.
aiAuthZ security framework for identity-bound authorization in AI agents, evaluating LLM vulnerability to tool-call forgery attacks across 15 models.
Self-Review Reinforcement Learning method for LLMs using cross-episode memory and policy distillation to handle sparse/delayed environmental feedback.
Study showing LLM conformity to peer responses is largely confounded by repeated answers independent of speaker presence in benchmark tasks.
Analysis of LLM yes-no bias in binary judgments, isolating effects of answer order and wording from moral judgment shifts using psychometric methods.
Evaluation study comparing prompt robustness between objective and subjective LLM tasks, showing different sensitivity patterns across model families.
Training-free approach using generative image models for 3D primitive shape abstraction without task-specific fine-tuning.
Analysis of structural concentration in AI bias research community, examining whose fairness definitions and debiasing frameworks are produced.
ResonatorLM architecture for efficient long-context language modeling using causal resonant field mixing as alternative to transformers.
Continual learning research challenging retention-centered assumptions and prioritizing real-time adaptation in non-stationary environments.
BaFCo benchmark dataset for Bangla form comprehension using multimodal LLMs, addressing low-resource language document understanding gaps.
EvalLoop methodology for evaluation-driven iterative improvement of LLM-based business systems, focusing on diagnosis and fixing rather than static model selection.
Empirical taxonomy of code mutations in AI-generated pull requests for performance optimization, analyzing 33k agent PRs.
RPAM: metric for measuring semantic associations and biases in language models with predictive validity for downstream task performance.
Survey study examining human trust in AI-generated legal advice versus human lawyers when advice is legally correct but socially controversial.
SCOReD: chain-of-thought distillation optimization for recommendation systems using student-aware training to avoid verbose reasoning traces.
Survey of execution-layer security research for AI coding agents covering isolation, access control, TOCTOU vulnerabilities, and protocol threats.
Security vulnerability analysis of Model Context Protocol implementations showing tool metadata concealment attacks across AI coding agent servers.
Counterfactual supervision method to train LLMs when to use external search versus parametric knowledge for improved task performance.
Data-dependent evaluation methods for budgeted submodular maximization algorithms in machine learning optimization problems.
Legato 2 pipeline for optical music recognition processing sheet music sequentially to extract symbolic notation and semantic knowledge.
FORGE framework enables robots to generalize tool use across novel objects via keypoint trajectory reasoning for functional transfer learning.
SegAnswer: multimodal LLM method using pixel-level segmentation before answering to improve visual reasoning accuracy in image understanding tasks.
AI-based screening system combining image classification and vessel segmentation for detecting retinopathy of prematurity in infants.
Vulnerability analysis of aligned LLMs comparing same-lineage models to isolate safety behavior from architectural differences in code review terminology.
In-context learning approach for ranking antibody candidates by binding affinity using contextual information from labeled antigen-specific comparisons.
Unsupervised anomaly detection method for identifying information operations users via behavioral and language patterns on social media.
Natural gradient descent variant for differentially private training that accounts for loss curvature to improve optimization efficiency under privacy constraints.
Analytical framework for LLM serving optimization using floor-first residual-driven triage to estimate resource bottlenecks before grid search.
Harrison.Rad 1.5 multimodal foundation model for automated radiology report generation from images, clinical history, and prior studies.
Simple coreset selection method using medoids for efficient few-shot knowledge distillation that surpasses random baseline sample selection.
Selective gradient computation framework reducing training cost by excluding low-loss samples from backward pass with unbiased gradient estimation.
Benchmark and methods for policy-adaptive image safety guardrails that generalize to policy changes without retraining.
Contextual multimodal document retrieval benchmark and methods that preserve textual/visual content while resolving multi-page aggregation queries.
Distributed multi-MCP architecture for vendor-agnostic SDN-based automation and autonomous control of multi-layer IPoDWDM networks with E2E service lifecycle automation.
Cost-efficient influencer matching system using three-stage cascade of small open-weight models instead of frontier LLM prompting for semantic matching on Thai marketing criteria.
Evaluation of LLM-generated metadata for RDF dataset search, comparing rewriting and agentic graph-based generation methods for retrieval effectiveness and faithfulness.
Agentic AI architecture using MCP protocol for autonomous control and lifecycle automation of multi-vendor IPoDWDM networks with closed-loop control validation.
Training-free method to improve grounding confidence in multimodal LLMs by detecting hallucinated spatial/temporal predictions using multi-token localized attention.
PluraMath extends mathematical reasoning evaluation to 99 languages beyond English and Chinese, addressing dataset bias in LLM benchmarks.
Prompt Coach is agentic tutor providing Socratic guidance for developers learning prompt engineering within IDE workflow.
SocaSim is LLM-based multi-agent simulation framework modeling Putnam's Social Capital Theory for studying collective action.
Studies incidental learning loss in AI-assisted software development and proposes design strategies for agents to support developer education.
RoME uses mixture of low-rank experts to achieve robustness against multiple adversarial perturbations with reduced trade-offs.
LLM-guided approach to correct measurement credibility in industrial soft sensing and process inference for robust predictions.
Training-free acceleration method for diffusion and flow matching models via x-prediction without retraining or distillation.
Predicts Java method energy usage incorporating execution time to enable early energy optimization in software development.
Empirical study of fine-tuning and evaluation metrics for neural decompilation of Dart AOT binaries using code-generation models.
Analyzes property-driven synthetic data generation engineering for data-scarce domains like breast cancer treatment planning systems.
Study on LLM agents for deliberative collaboration under partial observability with joint decision-making benchmark and multi-agent coordination.
LongCrafter synthesizes long-context supervised fine-tuning data with evidence graphs to improve LLM understanding across diverse tasks.